Project 18 · Advanced Deep Learning
Saliency-based Analysis of Shortcut Learning in CNNs
A ResNet18 trained on the Waterbirds dataset reaches 84% accuracy overall but only 60% on its worst subgroup. We use Grad-CAM and inference-time interventions to show that the gap comes from the model relying on the background rather than the bird.

Overall test accuracy
83.9%
Balanced test split
Worst-group accuracy
59.5%
Waterbird on land
Overall − worst-group gap
24.4%
Hidden by the headline number
Accuracy after background mask
86.0%
Removing the background helps
Pipeline
Train, evaluate by subgroup, run Grad-CAM, score foreground vs. background attention, intervene at inference time, then compare.
- 01Load Waterbirdsgrodino/waterbirds · 4 subgroups→
- 02Train ResNet18Select best by worst-group acc.→
- 03Subgroup evalAccuracy · P · R · F1 · confusion→
- 04Grad-CAMSaliency on layer4[-1]→
- 05Bias scoreBackground saliency / total→
- 06InterventionsBlur · mask · shuffle→
- 07CompareΔ accuracy · Δ flips · Δ saliency
Method
Dataset bias, training setup, and how the saliency and intervention experiments are defined.
Open →
Results
Subgroup accuracy, confusion matrix, Grad-CAM gallery, and the four intervention experiments.
Open →
Live demo
Upload a bird image and see the prediction, Grad-CAM, and background-bias score.
Open →