Skip to main content
Ctrl K

MIME

Mechanistic Interpretability Methodology & Ecosystem Exploration

This project explores the landscape of research and tooling in Mechanistic Interpretability (MI) to assess its integration potential for explainable AI frameworks, including DIANNA. While traditional post-hoc XAI tools excel at input-level feature attribution, MI provides deeper insight by inspecting a model's internal representations, activations, and circuits.

This exploratory research will review state-of-the-art MI tools (e.g., TransformerLens, nnsight) to evaluate their capabilities, limitations, and operational gaps. Simultaneously, the project establishes a connection with academic experts in interpretability to gather guidance, validate findings, and co-design a representative test case.

By benchmarking these tools against a pilot evaluation scenario, the project will identify missing capabilities in current ecosystems, evaluate whether extending DIANNA with MI functionalities is viable, and build the foundational knowledge necessary to design a concrete, large-scale research project for the following year.