We formalize and investigate explainability of machine learning classification in the particular setting where the explanation is to be provided by a part (substructure) of an instance. This is the common setting of string classification task in NLP and in graph-based setting appearing frequently in computational biology and chemistry. The point of this work is not to build explainers but to provide a formal framework that allows to study the properties of explanations in general. To this end, we consider a classifier f on some space of instance X, where Z is structured such that X is subset of a usually much larger set of substructure Z. As an example consider the problem of determining toxicity for a space X of molecules. Then Z is a space of fragments of molecules that also contains the complete molecules as a (small) subset. Inclusion of fragments in molecules and other fragments naturally induced a partial order structure on Z. An explanation for the classification f(x) of molecule x is a fragment y=e(x) that 'forces' the classification. The explainer, e, is simply a function that assigns e(x) in Z to every instance x. An explainer is said to be consistent if e(x)=e(x') implies f(x)=f(x'). An explanation is sufficient if the fact that e(x) explains the classification f(x) of x, also forces the same classification of every other instance that contains e(x), i.e., if e(x) <= x' implies f(x)=f(x').
We will consider conditions for both the existence and no-existence of sufficient explainers and observe that this depends both on the structure of the (sub)instance space Z and the classifier f. The condition can be relaxed to requiring the sufficiency condition only for a single class of instances $f(x)=a$. An example are graph classes defined by forbidden subgraph. It can be tested efficiently on empirical data sets -- in particular the labeled training data -- whether or not sufficient explanations can exist, resulting in very general constraints on explainability. It is also possible to obtain such sufficient explainers in practise.
(Joint work with Klaus Weinbauer, Sagar Malhotra, Thomas Gärtner, and Lukas Böhm)
Speaker
Peter StadlerExternal Professor