5  kumquat: A simpler lime

The existing lime package in R was adapted from the original Java implementation to handle a variety of data types beyond tabular data. However, it abstracts away critical intermediate components. Because the model object, perturbation data, and applied weights are not readily accessible to the end user, tracing the exact calculations behind the resulting LIME explanations is difficult.

To resolve this, we built a streamlined implementation of LIME that explicitly exposes these underlying elements. By providing direct access to the perturbation dataset and the local linear model, users can see exactly what is happening behind the scenes and effectively troubleshoot any anomalies arising from the XAI method itself.

This reimplementation is named kumquat (as kumquats are a smaller, lime-like fruit). It captures the fundamental mechanics of LIME for tabular data, but makes several deliberate technical departures from the original paper. The perturbation distribution is limited to a small, localized area around the observation, rather than extending broadly. In addition, observation weights are set to 1, giving equal importance to all generated perturbations. Finally, the local model is a Generalized Linear Model (logistic regression for classification tasks) fitted without regularization.

Unlike the original LIME implementation, which attempts to simplify the local model, kumquat prioritizes absolute fidelity to the original black-box model within the local region. Avoiding regularization ensures that feature importance scores are calculated for every considered variable, regardless of how small that importance might be.

The returned kumquat object provides the end user with a comprehensive diagnostic payload: the perturbation distribution, predictions from both the black-box and local linear models, the GLM model object, and the coefficients representing feature importances. This allows users to accurately visualize the exact data that generated the explanation.

Rather than submitting a pull request to append these features to the existing lime package, developing kumquat independently ensures that current lime users are not disrupted by structural changes to the package outputs. More importantly, it allows us to maintain a clean, highly transparent, and targeted implementation specifically designed for researchers auditing tabular data.

The package kumquat is currently available on CRAN as a published software. The website for the package can be found here