Evaluation, saliency, and intervention experiments
The headline accuracy looks strong, but subgroup metrics, Grad-CAM, and inference-time interventions all point to the same conclusion: the model leans on the background.
Subgroup accuracy
| Subgroup | Count | Accuracy | Avg confidence | Notes |
|---|---|---|---|---|
Waterbird on land (conflict) waterbird-land | 642 | 59.5% | 90.1% | worst group · shortcut conflict |
Landbird on water (conflict) landbird-water | 2255 | 73.6% | 87.6% | shortcut conflict |
Waterbird on water (majority) waterbird-water | 642 | 93.3% | 96.9% | majority group |
Landbird on land (majority) landbird-land | 2255 | 98.6% | 98.4% | majority group |
Grad-CAM saliency
For every subgroup, a large share of the model's attention falls on the background. The conflict groups are where that reliance turns into misclassifications.

Majority case: waterbird on water. Saliency overlaps the bird; background and foreground both look 'consistent'.

Majority case — model is confident and attention spans bird + water.

Another majority example, same pattern.

Conflict group success: a waterbird on land where the model still found the bird.

Shortcut failure: a waterbird placed on land. The model misclassifies it as landbird — saliency leaks onto the land background.

Another waterbird-on-land failure — strong evidence the background is steering the prediction.

Easy majority case for landbirds.

Landbird on land — accurate prediction.

Landbird majority case.

Conflict success: landbird placed on water but still classified correctly.

Shortcut failure: landbird on water becomes 'waterbird'. Saliency drifts toward the water background.

Another landbird-on-water failure — the background pattern dominates.
Intervention experiments
Editing the image at inference time isolates the background's causal role. Masking the foreground collapses accuracy and flips many predictions; changing only the background leaves accuracy intact or even improves it.
| Condition | Subgroup | N | Accuracy | Flip rate | Conf. drop | FG saliency | BG saliency |
|---|---|---|---|---|---|---|---|
| Original | Landbird on land (majority) | 352 | 98.3% | 0.0% | 0.0% | 54.6% | 45.4% |
| Original | Landbird on water (conflict) | 383 | 70.0% | 0.0% | 0.0% | 63.6% | 36.4% |
| Original | Waterbird on land (conflict) | 139 | 50.4% | 0.0% | 0.0% | 55.3% | 44.7% |
| Original | Waterbird on water (majority) | 126 | 93.7% | 0.0% | 0.0% | 61.7% | 38.3% |
| Background blur | Landbird on land (majority) | 352 | 97.7% | 1.1% | 1.9% | 61.9% | 38.1% |
| Background blur | Landbird on water (conflict) | 383 | 79.1% | 12.8% | -1.2% | 71.3% | 28.7% |
| Background blur | Waterbird on land (conflict) | 139 | 53.2% | 11.5% | 0.6% | 64.9% | 35.1% |
| Background blur | Waterbird on water (majority) | 126 | 90.5% | 6.3% | 2.0% | 70.2% | 29.8% |
| Background mask | Landbird on land (majority) | 352 | 96.9% | 1.4% | 3.0% | 71.2% | 28.8% |
| Background mask | Landbird on water (conflict) | 383 | 86.2% | 17.8% | -4.4% | 72.3% | 27.7% |
| Background mask | Waterbird on land (conflict) | 139 | 59.7% | 12.2% | 0.0% | 73.2% | 26.8% |
| Background mask | Waterbird on water (majority) | 126 | 84.1% | 11.1% | 4.1% | 73.4% | 26.6% |
| Background patch shuffle | Landbird on land (majority) | 352 | 98.3% | 1.7% | 1.0% | 57.1% | 42.9% |
| Background patch shuffle | Landbird on water (conflict) | 383 | 83.6% | 19.8% | -3.3% | 61.2% | 38.8% |
| Background patch shuffle | Waterbird on land (conflict) | 139 | 48.9% | 15.8% | 1.1% | 58.8% | 41.2% |
| Background patch shuffle | Waterbird on water (majority) | 126 | 81.0% | 12.7% | 7.8% | 61.2% | 38.8% |
| Foreground mask | Landbird on land (majority) | 352 | 97.2% | 4.0% | 4.0% | 34.9% | 65.1% |
| Foreground mask | Landbird on water (conflict) | 383 | 20.6% | 55.6% | 6.0% | 29.8% | 70.2% |
| Foreground mask | Waterbird on land (conflict) | 139 | 7.2% | 47.5% | -1.5% | 33.8% | 66.2% |
| Foreground mask | Waterbird on water (majority) | 126 | 81.7% | 19.8% | 11.0% | 33.0% | 67.0% |






Conclusion
The 24.4% gap between overall and worst-group accuracy, the high background saliency in every subgroup, and the intervention results together show that the CNN has partially learned the shortcut background → bird type. It still uses the bird, but it relies on the background enough to hurt minority-subgroup generalisation. The main limitations are that the foreground proxy is a fixed centre crop rather than a true segmentation mask, and that results come from a single ResNet18 run.