Strongest robustness gains
Rationale-based data augmentation, especially Sufficiency, most clearly improves robustness when shortcut and foreground evidence can be separated.
Interactive paper companion
1 German Research Center for Artificial Intelligence (DFKI) GmbH
2 RPTU University Kaiserslautern-Landau, Department of Computer Science
Explanation-Guided Learning promises models that are not only accurate, but also rely on task-relevant evidence. This work tests that promise across four complementary axes: shortcut-shift robustness, foreground versus shortcut reliance, explanation plausibility, and explanation fidelity. The interactive results below expose where these properties align—and where optimizing a plausible explanation does not actually change the model's decision strategy.
Paper overview
We compare attribution-constraining objectives with rationale-based data augmentation on a controlled cFMNIST benchmark containing localized patch, background, and global color shortcuts. The central result is that explanation-level improvements and behavioral improvements are strongly objective-specific and cannot be used as substitutes for one another.
Rationale-based data augmentation, especially Sufficiency, most clearly improves robustness when shortcut and foreground evidence can be separated.
Attribution-constraining objectives most strongly increase foreground-aligned Importance Mass, but yield weaker and less consistent behavioral correction.
Robustness, reliance, plausibility, and perturbation-based fidelity can diverge; none of them should be inferred from rationale agreement alone.
The cFMNIST dataset combines Fashion-MNIST foreground objects with CIFAR-style image backgrounds and controlled shortcut cues. The examples below illustrate how different shortcut regimes encode class-correlated information through patches, backgrounds, or color transformations, while preserving the foreground class label.
Supplementary Results
This section shows curated qualitative examples for inspecting how explanation maps change across objectives, λ values, and explanation methods. Examples are selected only when the compared model seeds share the same prediction for the input, so some model, dataset, class, and case-type combinations may have fewer than five examples or no matching example at all. The displayed examples are randomly selected individual cases and are intended for illustration only, not as aggregate evidence.
Select one objective. Columns show explanation methods.
Toggle which methods are shown as columns.
Select one explanation method. Columns show EGL objectives.
Toggle which objectives are shown as columns.
Supplementary Results
The quantitative panels provide the full experimental results as supplementary material for readers who want to inspect the findings in more detail. Cell values report the median over multiple random seeds. The shared controls select the model architecture and the rule used to choose λ⋆, while each panel shows the results for one evaluation axis. VisFIS* denotes a combination of configurations selected from separate objective-specific tuning runs, not a jointly optimized multi-objective model. Details on the experimental setup, evaluation protocol, and main findings are provided in the paper.
RQ1
Accuracy changes relative to the unguided baseline. The panels compare in-distribution accuracy with shortcut-randomized test accuracy under the selected model and λ* criterion.
RQ2
Diagnostic accuracy on foreground-only and shortcut-only inputs. Points show whether a selected model relies more strongly on task-relevant foreground evidence or on the isolated shortcut cue.
RQ3
Plausibility of post-hoc explanations with respect to the foreground rationale. The heatmaps compare how different objectives affect explanation alignment across methods and shortcut regimes.
RQ4
Perturbation-based explanation fidelity measured by DDS. The heatmaps show whether post-hoc explanations better reflect the model's predicted-class behavior under feature perturbation.
Paper and citation
The manuscript contains the full methodology, experimental protocol, discussion, and references. This website provides the complete interactive qualitative and quantitative result space.
David Dembinsky, Adriano Lucieri, and Andreas Dengel. Right for the Wrong Reasons? Evaluating Plausibility and Behavioral Change in Explanation-Guided Learning. Manuscript, 2026.
@misc{dembinsky2026right,
title = {Right for the Wrong Reasons? Evaluating Plausibility and Behavioral Change in Explanation-Guided Learning},
author = {Dembinsky, David and Lucieri, Adriano and Dengel, Andreas},
year = {2026},
note = {Manuscript}
}