Banner Banner

A review of vision transformer explainability: Methods, evaluation, and future research directions

Nora Koreuber
Claudia Winklmayr
Jim Berend
Wojciech Samek
Dagmar Kainmüller

September 02, 2026

Vision transformers have emerged as a powerful alternative to Convolutional Neural Networks in computer vision, increasingly serving as their successor in many applications. However, their black-box nature hinders their deployment in critical domains, where understanding model decisions is essen tial. Understanding how models work internally is also crucial for detecting biases and model debugging. This need has led to a growing number of explainability methods developed for or adapted to vision transformers in recent years, making it challenging to gain a clear understanding of the current state of the art. This comprehensive review explores the landscape of vision transformer explainability methods. We provide a struc tured overview of 81 methods, categorized in a novel fine-grained taxonomy, and illustrate its application through two exemplary use case scenarios for method selection. We further perform a comparative analysis of how these methods are evaluated in terms of models, datasets, baselines, and metrics. We observe a shift from early attention- and gradient-based attribution visualizations toward more holistic techniques, a growing use of concepts and embeddings as explanation targets, a gap in textual explanations of pure vision models, and the lack of integration into interactive explanation systems. Our findings reveal a highly fragmented evaluation landscape, with only partly overlapping use of metrics, over 130 distinct baseline methods, and a scarcity of human-centered and actionable evaluation.