Interactive paper companion

Right for the Wrong Reasons?

Evaluating Plausibility and Behavioral Change in Explanation-Guided Learning

David Dembinsky1,2 · Adriano Lucieri1 · Andreas Dengel1,2

1 German Research Center for Artificial Intelligence (DFKI) GmbH
2 RPTU University Kaiserslautern-Landau, Department of Computer Science

Explanation-Guided Learning promises models that are not only accurate, but also rely on task-relevant evidence. This work tests that promise across four complementary axes: shortcut-shift robustness, foreground versus shortcut reliance, explanation plausibility, and explanation fidelity. The interactive results below expose where these properties align—and where optimizing a plausible explanation does not actually change the model's decision strategy.

Paper overview

Plausible explanations are not enough

We compare attribution-constraining objectives with rationale-based data augmentation on a controlled cFMNIST benchmark containing localized patch, background, and global color shortcuts. The central result is that explanation-level improvements and behavioral improvements are strongly objective-specific and cannot be used as substitutes for one another.

+4.8 pp

Strongest robustness gains

Rationale-based data augmentation, especially Sufficiency, most clearly improves robustness when shortcut and foreground evidence can be separated.

+14 pp

Strongest plausibility gains

Attribution-constraining objectives most strongly increase foreground-aligned Importance Mass, but yield weaker and less consistent behavioral correction.

4 axes

No single success criterion

Robustness, reliance, plausibility, and perturbation-based fidelity can diverge; none of them should be inferred from rationale agreement alone.

Dataset

The cFMNIST dataset combines Fashion-MNIST foreground objects with CIFAR-style image backgrounds and controlled shortcut cues. The examples below illustrate how different shortcut regimes encode class-correlated information through patches, backgrounds, or color transformations, while preserving the foreground class label.

Shortcut regime
Foreground class

Canonical shortcut

Non-canonical shortcut

Supplementary Results

Qualitative Results

This section shows curated qualitative examples for inspecting how explanation maps change across objectives, λ values, and explanation methods. Examples are selected only when the compared model seeds share the same prediction for the input, so some model, dataset, class, and case-type combinations may have fewer than five examples or no matching example at all. The displayed examples are randomly selected individual cases and are intended for illustration only, not as aggregate evidence.

Example selection

Model
Dataset setup
Foreground class
Prediction case
Example

Comparison mode

Comparison view
EGL objective

Select one objective. Columns show explanation methods.

Shown explanation methods

Toggle which methods are shown as columns.

Supplementary Results

Quantitative Results

The quantitative panels provide the full experimental results as supplementary material for readers who want to inspect the findings in more detail. Cell values report the median over multiple random seeds. The shared controls select the model architecture and the rule used to choose λ⋆, while each panel shows the results for one evaluation axis. VisFIS* denotes a combination of configurations selected from separate objective-specific tuning runs, not a jointly optimized multi-objective model. Details on the experimental setup, evaluation protocol, and main findings are provided in the paper.

RQ1

Generalization Accuracy

Accuracy changes relative to the unguided baseline. The panels compare in-distribution accuracy with shortcut-randomized test accuracy under the selected model and λ* criterion.

Cell labels

RQ2

Shortcut vs Foreground Diagnostic

Diagnostic accuracy on foreground-only and shortcut-only inputs. Points show whether a selected model relies more strongly on task-relevant foreground evidence or on the isolated shortcut cue.

RQ3

Plausibility

Plausibility of post-hoc explanations with respect to the foreground rationale. The heatmaps compare how different objectives affect explanation alignment across methods and shortcut regimes.

Metric
Cell labels

RQ4

Fidelity

Perturbation-based explanation fidelity measured by DDS. The heatmaps show whether post-hoc explanations better reflect the model's predicted-class behavior under feature perturbation.

DDS perturbation
Cell labels

Paper and citation

Cite this work

The manuscript contains the full methodology, experimental protocol, discussion, and references. This website provides the complete interactive qualitative and quantitative result space.

David Dembinsky, Adriano Lucieri, and Andreas Dengel. Right for the Wrong Reasons? Evaluating Plausibility and Behavioral Change in Explanation-Guided Learning. Manuscript, 2026.

@misc{dembinsky2026right,
  title  = {Right for the Wrong Reasons? Evaluating Plausibility and Behavioral Change in Explanation-Guided Learning},
  author = {Dembinsky, David and Lucieri, Adriano and Dengel, Andreas},
  year   = {2026},
  note   = {Manuscript}
}