1 Introduction
Machine learning and deep learning combined have become popular solutions to a wide range of logical and creative problems in the real world. However, understanding the learned decision process of these models is difficult due to the opaqueness of the model. Explainable AI (XAI) methods were introduced to understand the underlying decision process of black box models(Molnar et al. 2020). These XAI methods are primarily categorized based on their scope and applicability(Molnar 2022).
The scope of XAI methods can be limited to explaining the influence of features on one observation or over the entire dataset. Based on the scope, XAI methods are categorized as global interpretability methods and local interpretability methods(Molnar et al. 2020). Global interpretability methods are used to show the influence of one or more features on the black box model’s decision process, while local methods are used to show the model’s decision process or feature influence for a specific observation of interest. Common examples for global methods are Partial Dependency Plots(Friedman 2001; Greenwell et al. 2017) and Accumulated Local Effect plots(Apley and Zhu 2020). For local methods, Local Interpretable Model Explanations (LIME) (Ribeiro et al. 2016), SHAP(Lundberg and Lee 2017), Anchors(Ribeiro et al. 2018) and Counterfactuals(Wachter et al. 2018) are common examples.
XAI methods are also categorized by whether the method is capable of explaining only a set family of machine learning models (e.g. tree based models, neural networks) or any black box model regardless of underlying architecture. In this categorization, XAI methods are divided into two categories: model-specific and model-agnostic. Model specific methods use underlying information about the model’s architecture to provide explanations while the model agnostic methods rely on the prediction mechanism of the model to approximate the internal workings of the model. Most XAI methods are model agnostic with a few model specific methods available specifically to deep learning models such as saliency maps(Simonyan et al. 2013) and learned feature visualisation(Olah et al. 2017) where information on the internal gradients are available.
In addition to the growth of XAI, in the recent years with the growth of Large Language Models, a new field of research surrounding specifically deep learning models have come to the spotlight. Mechanistic interpretability is focused on reverse engineering the inner workings of the model to understand the learned facts and knowledge inside. The growth of mechnastic interpretability while being rapid with multiple innovations made within a short period of time, still lacks clear, easy to understand visualisations for the general public. This is primarily due to the main focus been aimed at developing methods for model developer’s eye to debug and identify gaps in the model’s understanding.
1.1 Motivation
As decision boundaries becoming increasingly non linear around observations that model incorrectly classifies (James et al. 2013), generating individual explanations to understand the decision boundary around given observation becomes essential. From the plethora of available local interpretability methods, our focus will be on the methods of LIME, SHAP, Counterfactuals, and Anchors as they are model agnostic and therefore can be applied to a wide variety of models (Molnar 2022). These four methods have a similar structure that showcase different aspects of the model such as understanding the influence of features on a given prediction and the model boundary associated around the local neighborhood of the prediction.
When applying these XAI methods to real-world problems, the biggest hurdle is understanding which facet of the black box model is explained through each of these XAI methods. Different XAI methods can highlight various aspects of a model’s behavior, such as feature importance or decision pathways. This diversity, while powerful, also presents a challenge: practitioners must discern which method is most appropriate for their specific context and what particular aspect of the model’s decision-making process it elucidates. The lack of a unified framework often leaves users uncertain about how to interpret the results provided by different XAI techniques (Lundberg and Lee 2017).
Visual representations of local XAI explanations are needed to interpret the model boundary in the local neighborhood and to understand any disagreement that may occur between XAI methods. However, to observe the XAI methods visually we need geometric representations of each XAI method which can translate the abstract outputs of XAI methods into visual formats (objects) easily interpreted by humans. For example a geometric representation of an interval can be a bar, a line, or an error bar (Wickham 2010).
By mapping geometric representations of explanations into the data space (the data space \mathcal{X} is the p-dimensional region ,or vector space, where the set of observed p-dimensional data resides) we can leverage spatial intuition to better understand the model’s behavior. However, with large datasets the number of dimensions in the data and the resulting explanations grows drastically. Therefore high-dimensional visualisation methods are needed to generate insights on the complex high dimensional model boundaries.
The primary method for visualising high-dimensional data is using linear projections, for example, principal component analysis (Shlens 2014), or more generally sequences of linear projections as provided by a grand tour (Cook et al. 1995). A tour provides a continuous sequence of low dimensional linear projections in a given time frame as an animation by interpolating between the pairs of bases using geodesic interpolation. The sequence of linear projections for grand tour is generated by randomly selecting a basis from all possible projections to cover the entire spectrum of linear projections while a guided tour generates the sequence using an index function that seeks to find structures of interest. Aside from these methods the manual tour provides a sequence of projections that aim to travel between all axis parallel projections such that for each dimension of the data, the data is projected to a target dimensionality to cut out the specific dimension.
Using high dimensional visualisation techniques, the first research project is aimed at demystifying XAI in order to bridge the gap between explanations of XAI values and the data-model space using visual representations.
The second project is motivated by a concerning educational and industrial trend where model performance is increasingly equated with architectural complexity. In academic settings, students learn to associate state-of-the-art results with complex neural architectures, developing an intuitive liking toward complexity that persists into professional practice. This often leads engineers to default to convoluted, over-parameterised solutions without first considering simpler, more interpretable alternatives. The result is a field that consistently risks over-engineering solutions to problems that may have more elegant and efficient answers.
Therefore, a central motivation for the second project is to challenge the assumption that increasingly complex models are always necessary for strong performance. In practice, not all tasks require deep and intricate architectures; a well-initialised smaller model may perform just as well as, or even better than, a larger counterpart, while offering significant advantages in interpretability.
This pursuit of parsimony and interpretability leads directly to the challenges posed by today’s leading architectures in large language models: modern transformers(Vaswani et al. 2017). Transformers have scaled to billions of parameters, enabling them to encode vast amounts of world knowledge and achieve remarkable performance. However, their sheer scale renders them opaque, making it incredibly difficult to understand precisely how they represent facts or perform logical reasoning. Building on the previous project’s goal of demystifying models through visualization, the third project is aimed at studying much simpler, transformers to uncover the basic computational circuits that underlie reasoning and knowledge representation, thereby providing a path toward simplifying and understanding their larger counterparts.