Banner Banner

Theory Meets Practice #1: Opening the Black Box with LRP: A Hands-On Guide to Explainable AI for LLMs

Icon

January 07, 2026 Icon 18:00 - 21:00

Icon

TU Berlin (Room tba)

Icon

Reduan Achtibat

Opening the Black Box with LRP: A Hands-On Guide to Explainable AI for LLMs

©BLISS

Abstract: While Large Language Models have demonstrated unprecedented capabilities in reasoning and retrieval, their internal decision-making processes remain largely opaque. As they grow in complexity, treating these models as 'black boxes' is no longer sufficient for building reliable, safe, and transparent systems.

​Hosted by the Explainable AI group at Fraunhofer HHI, this workshop offers a practical, hands-on deep dive into the inner workings of Transformer models. The workshop will demonstrate how to apply advanced attribution and analysis methods to demystify how LLMs process information, store knowledge, and generate predictions.

 

 

​In this session, the workshop will cover:

  • Faithful Attribution: Going beyond input-level explanations to trace the model's reasoning from input tokens through latent concepts to the final prediction with Layer-wise Relevance Propagation.

  • Anatomy of In-Context Learning: Identifying how specific attention heads retrieve information from context versus storing factual knowledge.

  • Control and Source Tracking: Intervening in latent representations to steer model generation and trace information sources.

Logistics: Please bring your laptop to code along! The organizers provide the environment via Google Colab.

​Join to bridge the gap between model performance and model understanding.

​​Who is this event for?

​​This workshop is open to everyone curious about Explainability - including students, PhD candidates, and professionals! A basic understanding of Python and LLMs is all you need to participate.