Three Desiderata for Faithfulness in Machine Learning Explanations: The Case for Causal Abstraction
Mette Friis Andersen, Maria Heuss, and Ana Lucic
NeurIPS Workshop on Mechanistic Interpretability, 2025
Faithfulness is a broadly agreed-upon desideratum for explanations of machine learning model predictions. While many different methods have been adopted by the community, there is no agreed-upon definition of faithfulness. Here, we propose three desiderata for faithfulness beyond the standard intuition of accurately representing the reasoning process of the model, related to (1) enabling reverseengineering of specific behaviors, (2) capturing interventionist causal relations, and (3) achieving an appropriate model decomposition. We argue that causal abstraction satisfies these, and provides a framework for evaluating faithfulness claims in the community.