Explanation methods and systems for generative artificial intelligence model decisions
By introducing a response-perception attention mechanism and interpreter into the GenAI model, the shortcomings in the decision interpretation of GenAI model in the prior art are solved, and higher interpretation fidelity and interpretability are achieved.
Patent Information
- Application Number
- CN202510237871.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-03-03
AI Technical Summary
Existing generative artificial intelligence model interpretation methods cannot effectively explain the decision-making process of the GenAI model, especially in taking into account the model's feedback on the decision process and dealing with the decision-making prior heterogeneity.
A GenAI decision interpretation model including an interpreter and a response-perceptual attention mechanism is designed. By obtaining input data and responses, using the response-perceptual attention mechanism to introduce an interpreter to enhance internal features, and generate explanations through dot product calculations.
Improves the fidelity and interpretability of explanations, and can more accurately explain the decision-making process of the GenAI model, providing a concise and faithful explanation.
Smart Images

Figure CN119740614B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of generative artificial intelligence technology, and in particular to a method and system for explaining generative artificial intelligence model decisions. Background Art
[0002] GenAI (Generative Artificial Intelligence) models are essentially a type of machine learning model that can generate content in the form of text, images, or other media based on user prompts. Currently, many leading technology companies are investing heavily in the development of GenAI models to enhance and expand the functionality of their commercial products and services. However, GenAI models are opaque "black box" systems, and this opacity may generate inaccurate or fabricated content, leading to adverse or even serious consequences. In response to this issue, explanations need to be provided for GenAI's decisions.
[0003] At present, there are two main types of explanation methods for artificial intelligence models: the first type is the explanation method based on traditional machine learning, which uses the traditional machine learning framework to explain the model, such as the propagation-based explanation method and the perturbation-based explanation method. The second type is the information-based explanation method, which uses the concept of information theory to select and generate feature subsets in the model decision process to provide explanations for model decisions, such as learning explanations, variational information bottlenecks, etc.
[0004] However, while existing propagation-based and perturbation-based explanation methods provide a variety of intuitive ways to generate post-hoc explanations, their trade-off between fidelity and interpretability, and the inherent limitations of these methods may make them incompatible with the needs of generative AI models and various scenarios (e.g., reviews, news, and stock markets). Information-based explanation methods handle this trade-off more finely by separating the explanation and approximation processes. However, existing information-based explanation methods directly use the "input-label" learning loop to estimate the attribution score, ignoring the feedback of GenAI on the decision-making process, resulting in poor fidelity of the resulting explanations. From the above analysis, it can be seen that although generative AI models belong to AI models, there are certain differences between generative AI models and ordinary AI models, and the existing explanation methods proposed for AI models are not very suitable for generative AI models. Summary of the invention
[0005] 1. Technical issues to be resolved
[0006] In view of the shortcomings of the prior art, the present invention provides a method and system for interpreting generative artificial intelligence model decisions, filling the gap in the prior art that there is no interpretation method for generative artificial intelligence models.
[0007] (II) Technical solution
[0008] To achieve the above objectives, the present invention is implemented through the following technical solutions:
[0009] In a first aspect, the present invention provides a method for explaining decision making of a generative artificial intelligence model, comprising:
[0010] Get input data and response;
[0011] Input data and responses are processed through the pre-built GenAI decision explanation model to obtain explanations;
[0012] The pre-built GenAI decision explanation model includes an explainer and a response-aware attention mechanism;
[0013] The response-aware attention mechanism is used to process the response, and during the processing, an enhanced internal feature vector obtained by processing the input data by an interpreter is introduced to obtain an external attribution score;
[0014] The interpreter is used to process the input data and the external attribution score to obtain a continuous relaxed random vector, and obtain an interpretation after performing a dot product calculation on the continuous relaxed random vector and the embedding vector of the input data.
[0015] Preferably, the interpreter comprises a first feature embedding layer, a first embedding enhancement layer, an embedding element-level convolution layer, a first normalization layer, a first mean calculation layer, a second normalization layer and a sparse feature layer;
[0016] The response-aware attention mechanism includes a second feature embedding layer, a second embedding enhancement layer, a second mean calculation layer, and a third normalization layer;
[0017] The input data is processed by a first feature embedding layer, a first embedding enhancement layer, an embedding element-level convolution layer, and a first normalization layer in the interpreter to obtain an internal feature attribution score;
[0018] The response is passed through a second feature embedding layer and a second embedding enhancement layer in a response-aware attention mechanism to obtain an enhanced external feature vector, wherein the enhanced external feature vector is aligned with the enhanced internal feature vector output by the first embedding enhancement layer;
[0019] The enhanced internal feature vector and the enhanced external feature vector are subjected to dot product calculation to obtain a dot product; the dot product passes through the second mean calculation layer and the third normalization layer to obtain an external attribution score;
[0020] The internal feature attribution score and the external attribution score are passed through the first mean calculation layer and the second normalization layer to obtain a fused attribution score; after the fused attribution score passes through the sparse feature layer, a continuous relaxed random vector is obtained, and the continuous relaxed random vector and the input data embedding vector output by the first feature embedding layer are interpreted after the dot product calculation.
[0021] Preferably, the loss function of the pre-built GenAI decision explanation model during the training optimization process is include:
[0022]
[0023] in, Represents the cross entropy loss, which is used to measure the cross entropy between the decision prediction of the GenAI decision explanation model and the GenAI model decision; represents the KL divergence loss, which is used to measure the KL divergence between the continuous relaxed random vector and the prior vector;
[0024] The decision of the GenAI decision explanation model is obtained by: processing the explanation through an approximator to obtain a decision prediction;
[0025] The method for obtaining the prior vector includes: processing the data set level input and response of the task through the response-guided GenAI decision prior estimator to obtain the prior vectors of all instances in the task.
[0026] Preferably, the response-guided GenAI decision prior estimator includes a word embedding layer, a similarity calculation layer, a mean calculation layer and a normalization layer;
[0027] In the response-guided GenAI decision prior estimator, generated by the word embedding layer
[0028] The word embedding vectors of the input and response of the task at the dataset level; the average cosine similarity between each feature in the input document and each word in the response is calculated through the similarity calculation layer to estimate the candidate priors for each feature in the input; then the co-occurrence probability is counted, and the average prior of all features contained in the task document is calculated through the mean calculation layer to obtain the task-level prior; the task-level prior is processed through the normalization layer to obtain the prior vectors of all instances in the task.
[0029] Preferably, when the input data is text data, the first feature embedding layer uses the GloVe word embedding method to capture the semantics of each word and obtain the embedding vector of each word.
[0030] Preferably, the embedding element-wise convolutional layer uses a CNN approach to map each enhanced word embedding vector to an attribution score.
[0031] Preferably, the sparse feature layer uses the Gumbel-softmax technique to sparse the fused attribution score, specifically including:
[0032] Construct a set of specific random vectors as a differentiable approximation of the argmax operation, where the set of specific random vectors is defined as:
[0033]
[0034] The calculation method is:
[0035]
[0036] in, is a random perturbation, is a temperature parameter, when Approaching 0, the probability distribution obtained after softmax converges to a discrete distribution;
[0037] Perform k independent subset samplings to perform argmax operations and further construct a continuous relaxed random vector :
[0038] .
[0039] In a second aspect, the present invention provides a generative artificial intelligence model decision explanation system, comprising:
[0040] Input layer, used to obtain input data and responses;
[0041] The model processing and output layer is used to process the input data and responses through the pre-built GenAI decision explanation model to obtain explanations;
[0042] The pre-built GenAI decision explanation model includes an explainer and a response-aware attention mechanism;
[0043] The response-aware attention mechanism is used to process the response, and during the processing, an enhanced internal feature vector obtained by processing the input data by an interpreter is introduced to obtain an external attribution score;
[0044] The interpreter is used to process the input data and the external attribution score to obtain a continuous relaxed random vector, and obtain an interpretation after performing a dot product calculation on the continuous relaxed random vector and the embedding vector of the input data.
[0045] In a third aspect, the present invention provides a computer-readable storage medium storing a computer program for interpreting generative artificial intelligence model decisions, wherein the computer program enables a computer to execute the method for interpreting generative artificial intelligence model decisions as described above.
[0046] In a fourth aspect, the present invention provides an electronic device, comprising:
[0047] One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the programs include an explanation method for executing the generative artificial intelligence model decision as described above.
[0048] (III) Beneficial effects
[0049] The present invention provides a method and system for explaining generative artificial intelligence model decisions. Compared with the prior art, it has the following beneficial effects:
[0050] The present invention first obtains input data and responses; then processes the input data and responses through a pre-built GenAI decision explanation model to obtain an explanation. The pre-built GenAI decision explanation model includes an interpreter and a response-aware attention mechanism. The response-aware attention mechanism is used to process the response, and the enhanced internal feature vector obtained after the interpreter processes the input data is introduced during the processing to obtain an external attribution score; the interpreter is used to process the input data and the external attribution score to obtain a continuous relaxed random vector, and the continuous relaxed random vector is subjected to a dot product calculation with the embedded vector of the input data to obtain an explanation. Aiming at the characteristics of the GenAI model, the present invention designs a novel response-aware attribution attention mechanism, takes into account the response of the GenAI model to the decision-making process, improves the fidelity of the explanation, and provides a concise and faithful explanation for the decision of the GenAI model. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0052] Figure 1 It is a framework diagram of the GenAI decision explanation model in an embodiment of the present invention;
[0053] Figure 2This is an overall framework diagram of the GenAI decision explanation model in an embodiment of the present invention when processing classification tasks. DETAILED DESCRIPTION
[0054] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are clearly and completely described. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0055] The embodiments of the present application provide a method and system for explaining the decisions of a generative artificial intelligence model, thereby filling the gap in the prior art in the lack of an explanation method for a generative artificial intelligence model, and specifically targeting the characteristics of the GenAI model to provide a concise and faithful explanation for the decisions of the GenAI model.
[0056] The technical solution in the embodiment of the present application is to solve the above technical problems, and the overall idea is as follows:
[0057] At present, the common interpretation methods of artificial intelligence models mainly include propagation-based, perturbation-based and information-based methods. Among them, the propagation-based interpretation method back-propagates the contribution of neurons or gradients in the neural network to each input feature, such as integrated gradients (IG) and deep learning important features (DeepLIFT). The perturbation-based interpretation method generates perturbations of each instance to build interpretable local approximate proxy models (such as linear models). However, due to their nature of simultaneous interpretation and approximation, both propagation-based and perturbation-based interpretation methods encounter a difficult trade-off between fidelity (i.e., imitating model results) and human interpretability (the trade-off between fidelity and interpretability, trade-off). The information-based interpretation method handles this trade-off more finely by separating the interpretation and approximation processes.
[0058] However, existing information-based explanation methods directly use the "input-label" learning loop to estimate the attribution score without considering the feedback (response) of the GenAI to the decision process, which makes it challenging to achieve faithful feature attribution (ignoring the response of the GenAI model to the decision process). Second, the implicit assumptions on the attribution prior (such as uniform distribution) are naturally incompatible with typical practical situations, where the GenAI prior is heterogeneous across different tasks and even different instances (the problem of heterogeneity of GenAI model decision priors).
[0059] In order to solve the above problems, the embodiments of the present invention specifically propose a method and system for explaining the decisions of generative artificial intelligence models, which maximizes the compression of input features (interpretability) while generating a local explanation of the maximum amount of information about the GenAI decision (fidelity). In order to solve the problem of "ignoring the response of the GenAI model to the decision-making process", the embodiments of the present invention design a response-aware attribution attention mechanism. In order to solve the "heterogeneity problem of GenAI model decision priors", the embodiments of the present invention design a response-guided GenAI decision prior estimator, and use the information bottleneck principle to guide model learning, generate a "good" bottleneck (explanation), and improve the interpretability and fidelity of the explanation.
[0060] In order to better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0061] The embodiment of the present invention provides a method for explaining generative artificial intelligence model decisions, including:
[0062] Get input data and response;
[0063] Input data and responses are processed through the pre-built GenAI decision explanation model to obtain explanations;
[0064] The pre-built GenAI decision explanation model is as follows: Figure 1 As shown, it includes an interpreter and an attention mechanism for response perception;
[0065] The interpreter includes a first feature embedding layer, a first embedding enhancement layer, an embedding element-level convolution layer, a first normalization layer, a first mean calculation layer, a second normalization layer, and a sparse feature layer;
[0066] The response-aware attention mechanism includes a second feature embedding layer, a second embedding enhancement layer, a second mean calculation layer, and a third normalization layer;
[0067] The input data is processed by a first feature embedding layer, a first embedding enhancement layer, an embedding element-level convolution layer, and a first normalization layer in the interpreter to obtain an internal feature attribution score;
[0068] The response is passed through a second feature embedding layer and a second embedding enhancement layer in a response-aware attention mechanism to obtain an enhanced external feature vector, wherein the enhanced external feature vector is aligned with the enhanced internal feature vector output by the first embedding enhancement layer;
[0069] The enhanced internal feature vector and the enhanced external feature vector are subjected to dot product calculation to obtain a dot product; the dot product passes through the second mean calculation layer and the third normalization layer to obtain an external attribution score;
[0070] The internal feature attribution score and the external attribution score are passed through the first mean calculation layer and the second normalization layer to obtain a fused attribution score; after the fused attribution score passes through the sparse feature layer, a continuous relaxed random vector is obtained, and the continuous relaxed random vector and the input data embedding vector output by the first feature embedding layer are interpreted after the dot product calculation.
[0071] The embodiment of the present invention specifically targets the characteristics of the GenAI model and designs a novel response-aware attribution attention mechanism, which takes into account the response of the GenAI model to the decision-making process, improves the fidelity of the explanation, and provides a concise and faithful explanation for the decision of the GenAI model.
[0072] In the specific implementation process, the pre-built GenAI decision explanation model needs to use the response-guided GenAI decision prior estimator and approximator during the training process. The overall structure of the GenAI decision explanation model is as follows: Figure 2 shown.
[0073] Among them, the purpose of the explainer is to estimate the attribution score of each feature and generate an explanation based on the feature subset of each instance. However, there are two problems in generating explanations: ① Many words rely on adjacent words to convey accurate and complete meanings. Considering each word independently may not be enough to capture all semantic features, resulting in incomplete explanations; however, considering the entire context of the word may inadvertently capture redundant semantics for interpretation, resulting in insufficient interpretability. Therefore, a feature embedding layer is designed to solve this problem. ② Semantics are usually represented by embedding vectors, while the attribution score is a scalar. Therefore, an embedding element-level convolutional layer is designed to obtain the attribution score. The explainer is mainly divided into two parts: feature attribution and generating explanations. Among them: feature attribution mainly includes three parts: internal attribution score based on input, external attribution score based on response, and fusion attribution score based on internal and external. Generating explanations mainly includes two parts: sparse features and generating explanations.
[0074] The response-aware attention mechanism is used to capture the external features of the decision response and estimate the external feature attribution score as a complement to the internal feature attribution.
[0075] The approximator predicts the GenAI model decision based on the explanations generated by the explainer. When the task is a classification task, the predicted decision is the category label of the input text.
[0076] Taking the classification task as an example, the input data is text data, and the normalization layer is normalized by the softmax function. Figure 1 , Figure 2The various layers in the explainer, response-aware attention mechanism, approximator, and response-guided GenAI decision prior estimator are described in detail:
[0077] First feature embedding layer: The input data is text data. The input data first enters the feature embedding layer and uses the Glove word embedding method to capture the semantics of each word and obtain the embedding vector of each word:
[0078]
[0079] in: is the input embedding vector, ( ) is the first i The feature vector of each word, Indicates that the input document contains words, and the embedding dimension is The GloVe word embedding method is used because it is a context-independent semantic capture method, which can explain the above-mentioned problem① with the following embedding enhancement layer.
[0080] The first embedding enhancement layer: This layer is a word-level convolutional layer using the CNN method. This layer is set to have A 1D convolutional layer with convolution kernels, each of which has a size of , is a hyperparameter, indicating the number of adjacent words in the context of each word. The step size of the convolutional layer is 1 and the padding size is The use of CNN convolution method can consider contextual semantics based on the Glove method, improve problem ①, and enhance the semantic representation of the embedded vector:
[0081]
[0082] in , For the i The enhanced embedding vectors of the words.
[0083]
[0084] Get the enhanced embedding vector of the input document .
[0085] Embedding element-wise convolutional layer: Use the CNN method to map each enhanced word embedding vector to an attribution score. This layer is set as a 1D convolutional layer with 1 convolution kernel and a size of , with a step size of 1 and no padding. After passing through the softmax layer, the internal attribution score is obtained:
[0086]
[0087] in: It is i The internal feature attribution score of each word. Indicates i The enhanced embedding vectors of the words.
[0088] Second feature embedding layer: Response refers to the process of GenAI self-rationalization of decisions. For example, when a decision is made, GenAI is further asked why it makes such a decision, why it makes such a decision, how it makes such a decision, etc. Similar to the acquisition of internal attribution scores, the embodiment of the present invention designs a feature embedding layer to obtain external features based on the response document input.
[0089]
[0090] in: is the response embedding vector, ( ) is the first The feature vector of each word. Indicates that the input document contains words, and the embedding dimension is .
[0091] Second embedding enhancement layer: Get the enhanced response feature word embedding vector:
[0092]
[0093] Second mean calculation layer: To capture the relevance between the input document and the response document and obtain the response-based external attribution score, the average dot product (similarity) between the embedding vector of each word in the input document and the embedding vector of each word in the response document is calculated.
[0094] Third softmax layer: The output of the second mean calculation layer passes through the second softmax layer to obtain the external attribution score.
[0095]
[0096] in: It is i External attribution fraction, m is the number of words in the response document.
[0097] First mean calculation layer: After obtaining the internal attribution score and the external attribution score, the first mean calculation layer calculates the mean of the two scores.
[0098] Second softmax layer: The output of the first mean calculation layer passes through the first softmax layer and normalizes the mean of the two scores through the softmax function to obtain the final fusion attribution score.
[0099]
[0100]
[0101] in, It is i The final fused attribution score for the word, is the fused attribution score vector.
[0102] Sparse feature layer: Based on the principle of information bottleneck, in order to generate explanations that are both concise and rich in information, the embodiment of the present invention uses the K-hot method to further sparse features. can be interpreted as a discrete distribution, and in order to further sparse features and generate explanations with k features, s can be sampled k times independently. However, this subset sampling method is inherently discrete and non-differentiable, which makes it incompatible with gradient-based training schemes in variational deep learning. Therefore, the Gumbel-softmax trick is used to approximate non-differentiable categorical sampling with continuously differentiable samples.
[0103] The first step is to construct a set of specific random vectors as a differentiable approximation of the argmax operation:
[0104] The specific set of random vectors is defined as:
[0105]
[0106] The calculation method is:
[0107]
[0108] in: is a random perturbation, is a temperature parameter, when If it approaches 0, the probability distribution obtained after softmax converges to a discrete distribution.
[0109] Then proceed k The argmax operation is performed on the independent subset sampling, and a continuous relaxed random vector is further constructed. to approximate the k-hot vector, also known as the "bottleneck":
[0110]
[0111] It should be noted that, in the specific implementation process, feature sparseness can also be achieved through other methods, such as reparameterization trick, continuous relaxation, etc.
[0112] Generate Explanation: Compute the dot product of the bottleneck and the input vector, generating k Explanation of the features :
[0113]
[0114] In the specific implementation process, the GenAI decision explanation model needs to be pre-trained and optimized. During the training and optimization work, an approximator and a response-guided GenAI decision prior estimator are required.
[0115] The purpose of the approximator is to predict the decision of the GenAI model and obtain a decision prediction. In the embodiment of the present invention (classification task), the approximator predicts the decision of the GenAI model based on the explanation generated by the interpreter, that is, the category label of the input text.
[0116] The approximator receives the explanation provided by the interpreter To obtain the semantic information of the explanation, the approximator uses a Bi-LSTM (semantic representation layer) to process these words. The reason for using Bi-LSTM is that Bi-LSTM can take into account the contextual information of the word and generate rich semantic representations for each key word in the explanation.
[0117] Hierarchical processing: In order to obtain the semantic information of the explanation more comprehensively and approximate the model decision, the word-level Bi-LSTM layer and Sentence-level Bi-LSTM are designed.
[0118] The word-level Bi-LSTM processes the key words selected by the explainer and generates semantic representations for each word. This level of Bi-LSTM can capture the local context information of the word and understand the meaning of the word in a specific context. And through the forward and backward propagation of Bi-LSTM, the model can encode the temporal dependencies in the word sequence, which is crucial for understanding the interaction between words and the overall structure of the sentence.
[0119] Then, based on the Word-level Bi-LSTM, it is further processed by Sentence-level Bi-LSTM. It aggregates the semantic representation of each sentence to generate a global semantic representation of the entire sentence. When processing articles or documents composed of multiple sentences, sentence-level Bi-LSTM can integrate the semantic information of each sentence to provide a unified semantic representation for the entire document.
[0120] Prediction layer: Finally, a fully connected layer (prediction layer) is used to obtain the predicted probability distribution. In the classification task of this instance, the probability of the category label is obtained. The decision prediction of the GenAI decision explanation model is obtained. .
[0121] The loss function in the pre-built GenAI decision explanation model of the embodiment of the present invention takes into account the cross entropy loss and KL divergence loss.
[0122]
[0123] in, Represents decision prediction Decision making with GenAI models The cross entropy of is minimized during training. The purpose is to guide decision-making predictions Approximate the real decision of the GenAI model, optimize the parameters, and further obtain the optimal bottleneck. represents the KL divergence between the continuous relaxed random vector and the prior vector. In an embodiment of the present invention, a response-guided GenAI decision prior estimator is designed. It uses the input of the dataset level and the response of the response to estimate the prior vector of the GenAI decision.
[0124] In the response-guided GenAI decision prior estimator, the dataset-level input documents and response documents are first passed through the word embedding layer to generate dataset-level word embedding vectors.
[0125] The candidate priors for each feature in the input are then estimated by computing the average cosine similarity between each feature in the input document and each word in the response:
[0126]
[0127] For the The candidate priors for features, .
[0128] Then count the co-occurrence probability and the average prior of all features included in the task documents; calculate the task-level prior:
[0129]
[0130] It is a feature The task-level priors of The task contains features The total number of all instances of .
[0131] Finally, the prior vector of the instance is estimated through softmax.
[0132]
[0133] Computational bottleneck and estimated GenAI decision priors , deoptimize to obtain the optimal bottleneck and generate a concise and informative representation.
[0134] In the training and optimization process of the GenAI decision explanation model, the explainer and the approximator are put together and jointly learned based on the variational information bottleneck objective:
[0135]
[0136] Where b represents the batch size, H represents the number of hidden layers in the GenAI decision explanation model, and h represents the hth hidden layer. C represents the total number of words in the experiment, and c represents the cth word. , is the index, indicating the words for summation; is a weight parameter, Represents two probability distributions , The KL divergence between Indicates h The predicted log probability of the cth word output by the hidden layer; represents the true label; Represents the model prediction label.
[0137] z(b) yes b From the variational distribution under batch p ( z|x ) is a sample of z obtained by sampling;
[0138] It should be noted that the information bottleneck principle is an information theory principle that aims to extract relevant information contained in one random variable (vector) about another random variable (vector). It provides "a surprisingly rich framework for discussing various problems in signal processing and learning." In the context of supervised learning, the principle essentially seeks to identify an optimal, highly compressed input mapping that retains the maximum information about the output (label). The compressed representation is also called a "bottleneck."
[0139] The final mapping (bottleneck) is obtained by optimizing the following objectives:
[0140]
[0141] in: is the mutual information term, indicating: Explanation Decision making with GenAI models The mutual information between them should be as large as possible; is a compressed term, indicating that the input With explanation The mutual information between them should be as small as possible to achieve feature compression; is a weight parameter used to balance the trade-off between information preservation and information compression. Using the optimal bottleneck objective function to optimize the GenAI decision explanation model can achieve: maximize the information representation under the maximum compression feature (that is, obtain a concise, informative and faithful explanation). The optimal bottleneck solution is defined as: k-hot vector , , and the optimal explanation is generated from the optimal bottleneck .
[0142] Then the optimal bottleneck objective function is computationally intractable. However, the recent development of variational deep learning provides a feasible method to approximate the mutual information and then determine the optimal bottleneck, namely the variational information bottleneck. It introduces variational bounds in the deep learning process and uses deep neural networks for nonlinear mapping to obtain a variational lower bound for the information bottleneck objective:
[0143]
[0144] Where G represents a constant term, For a given hour, The conditional probability distribution of .
[0145] In the process of training and optimizing the GenAI decision explanation model, back propagation calculation The gradient of , updates the network parameters (weights and bias); if If it does not decrease in 5 rounds, RGDL is optimal and the training ends.
[0146] The effectiveness of the interpretation method of the generative artificial intelligence model decision-making in the embodiment of the present invention is verified by a specific example below:
[0147] Example 1, Image Caption Generation Task (ICG):
[0148] In this example, 20,000 images were collected from the COCO image captioning dataset, of which 10 themes (people, animals, vehicles, electronics, food, indoor, furniture, kitchen, outdoor, and sports) each had about 2,000 images. In this task, the results of the fidelity of the explanation method of the generative artificial intelligence model decision of the embodiment of the present invention and some existing solutions are shown in Table 1.
[0149] Table 1 Fidelity of various explanation methods in the image caption generation task
[0150]
[0151] Note: CIDEr stands for Continuous Interpretability Diagnostic for Evaluation of Recurrent Neural Networks. CIDEr is a metric used to evaluate the performance of natural language generation models. SPICE (Semantic Propositional Image Caption Evaluation) is a metric used to evaluate the quality of image descriptions, mainly used to measure the semantic similarity between automatically generated image descriptions and human descriptions.
[0152] In the table, RGDL represents the interpretation method of the generative artificial intelligence model decision of an embodiment of the present invention.
[0153] DeepLIFT (Deep Learning Important Features) is a method for explaining deep learning model predictions that aims to quantify the contribution of input features to changes in model output. The implementation process of this method is as follows: 1) Reference point selection: First, a reference point or baseline input is selected, usually a background distribution or zero input. 2) Difference attribution: Calculate the impact of the difference between the actual input and the reference point on the model output. This step involves calculating the difference in model output under the actual input and the reference point input, and attributing it to the change in input features. 3) Attribution score: By comparing the activation difference of the model under the actual input and the reference point input, DeepLIFT assigns an attribution score to each input feature, indicating the degree to which the feature contributes to the model prediction.
[0154] IG (Integrated Gradients) is a gradient-based attribution method for explaining the predictions of deep learning models. IG is implemented through the following steps: 1) Path Integration: Starting from a baseline (such as the zero vector or the mean of the data), the gradient of the model output relative to the input is integrated along the path from the baseline to the actual input. This step involves accumulating the gradients at each point on the path. 2) Attribution Score: By calculating the gradient contribution of each input feature on the path, IG assigns an attribution score to each feature, indicating the degree to which the feature contributes to the model prediction. 3) Satisfaction Properties: IG satisfies some important attribution properties, such as completeness (the sum of the attribution scores of all features is equal to the output change) and uniqueness (the attribution score does not depend on the choice of path).
[0155] LIME fits a simple model in the local area around the target sample to approximate the behavior of the complex model. This simple model is interpretable (such as a linear model, a decision tree model) and can capture the decision logic of the complex model in the local area.
[0156] L2X (Local to Global Explanations) is a method that extends local explanations to global explanations. It is achieved through the following steps: 1) Local approximation: L2X uses simplified models (such as linear models) within the local neighborhood of the model to approximate the behavior of complex models. 2) Global synthesis: By synthesizing multiple local explanations, L2X constructs a global explanation to reveal the overall behavior and decision-making process of the model. 3) Model independence: L2X does not depend on the specific type of the model and can be applied to various machine learning models, including deep learning models.
[0157] VIBI (Variational Information Bottleneck) is an explanation method based on the deep variational information bottleneck method. It is implemented through the following steps: 1) Information bottleneck principle: VIBI uses the information bottleneck principle to identify the part of the input features that is informative to the output. 2) Variational inference: Through variational inference techniques, VIBI estimates the mutual information between input features and outputs to quantify the importance of features. 3) Optimizing compressed representation: VIBI aims to optimize the compressed representation of input data while retaining information useful for prediction tasks.
[0158] SelfExp (Self-Explaining Neural Networks) is a self-explaining neural network method that is implemented through the following steps: 1) Built-in explanations: SelfExp directly builds explanations during the training process of the neural network, making the model itself explanatory. 2) Feature visualization: By visualizing the intermediate layer features of the network, SelfExp provides insights into how the model learns and makes predictions. 3) Model transparency: SelfExp aims to improve the transparency of the model, making the model's decision-making process more intuitive and easy to understand.
[0159] Example 2: News Topic Classification Task (NTC):
[0160] In this example, 20,000 news articles were collected from the AG News dataset, with 5,000 articles for each of the four topics (world, sports, business, and science / technology). In this task, the results of the fidelity of the explanation method of the generative artificial intelligence model decision of the embodiment of the present invention and some existing solutions are shown in Table 2.
[0161] Table 2 Fidelity of each explanation method in the news topic classification task
[0162]
[0163] The interpretability performance of each explanation method in the image title generation task and news topic classification task is shown in Table 3.
[0164] Table 3 Interpretability of various explanation methods in the image title generation task and news topic classification task
[0165]
[0166] Through the above verification process, it can be seen that the fidelity and interpretability of the interpretation method of the generative artificial intelligence model decision in the embodiment of the present invention are significantly better than the existing methods.
[0167] The embodiment of the present invention also provides a generative artificial intelligence model decision explanation system, including:
[0168] Input module, used to obtain input data and responses;
[0169] The model processing module is used to process the input data and responses through the pre-built GenAI decision explanation model to obtain explanations;
[0170] The pre-built GenAI decision explanation model includes an explainer and a response-aware attention mechanism;
[0171] The response-aware attention mechanism is used to process the response, and during the processing, an enhanced internal feature vector obtained by processing the input data by an interpreter is introduced to obtain an external attribution score;
[0172] The interpreter is used to process the input data and the external attribution score to obtain a continuous relaxed random vector, and obtain an interpretation after performing a dot product calculation on the continuous relaxed random vector and the embedding vector of the input data.
[0173] It is understandable that the interpretation system for generative artificial intelligence model decisions provided in the embodiments of the present invention corresponds to the interpretation method for generative artificial intelligence model decisions mentioned above, and the explanations, examples, beneficial effects, etc. of the relevant contents can refer to the corresponding contents in the interpretation method for generative artificial intelligence model decisions, which will not be repeated here.
[0174] An embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program for interpreting generative artificial intelligence model decisions, wherein the computer program enables a computer to execute the method for interpreting generative artificial intelligence model decisions as described above.
[0175] An embodiment of the present invention also provides an electronic device, comprising: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the program includes an interpretation method for executing the generative artificial intelligence model decision as described above.
[0176] In summary, compared with the prior art, the present invention has the following beneficial effects:
[0177] The existing explanation methods ignore the response of the GenAI model to the decision-making process and have difficulty in dealing with the heterogeneity of the GenAI model's decision priors. The embodiment of the present invention first designs a response-aware attribution attention mechanism and a response-guided GenAI decision prior estimator based on the GenAI decision response, and proposes a method for explaining the decisions of a generative artificial intelligence model. This method takes into account the response of the GenAI model to the decision-making process through a novel response-aware attribution attention mechanism and a response-guided GenAI decision prior estimator, estimates the prior probability of the GenAI model's decision, and improves the accuracy and fidelity of the explanation.
[0178] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0179] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for explaining decisions of a generative artificial intelligence model, characterized in that: include: Get input data and response; Input data and responses are processed through the pre-built GenAI decision explanation model to obtain explanations; The pre-built GenAI decision explanation model includes an explainer and a response-aware attention mechanism; The response-aware attention mechanism is used to process the response, and during the processing, an enhanced internal feature vector obtained by processing the input data by an interpreter is introduced to obtain an external attribution score; The interpreter is used to process the input data and the external attribution score to obtain a continuous relaxed random vector, and obtain an interpretation after performing a dot product calculation on the continuous relaxed random vector and the embedding vector of the input data; The interpreter includes a first feature embedding layer, a first embedding enhancement layer, an embedding element-level convolution layer, a first normalization layer, a first mean calculation layer, a second normalization layer, and a sparse feature layer; The response-aware attention mechanism includes a second feature embedding layer, a second embedding enhancement layer, a second mean calculation layer, and a third normalization layer; The input data is processed by a first feature embedding layer, a first embedding enhancement layer, an embedding element-level convolution layer, and a first normalization layer in the interpreter to obtain an internal feature attribution score; The response is passed through a second feature embedding layer and a second embedding enhancement layer in a response-aware attention mechanism to obtain an enhanced external feature vector, wherein the enhanced external feature vector is aligned with the enhanced internal feature vector output by the first embedding enhancement layer; The enhanced internal feature vector and the enhanced external feature vector are subjected to dot product calculation to obtain a dot product; the dot product passes through the second mean calculation layer and the third normalization layer to obtain an external attribution score; The internal feature attribution score and the external attribution score are passed through the first mean calculation layer and the second normalization layer to obtain a fused attribution score; after the fused attribution score passes through the sparse feature layer, a continuous relaxed random vector is obtained, and the continuous relaxed random vector and the input data embedding vector output by the first feature embedding layer are interpreted after the dot product calculation.
2. The method for explaining generative artificial intelligence model decisions as claimed in claim 1, characterized in that: The loss function of the pre-built GenAI decision explanation model during training optimization include: in, represents the cross entropy loss, which is used to measure the cross entropy between the decision prediction of the GenAI decision explanation model and the GenAI model decision; represents the KL divergence loss, which is used to measure the KL divergence between the continuous relaxed random vector and the prior vector; The decision of the GenAI decision explanation model is obtained by: processing the explanation through an approximator to obtain a decision prediction; The method for obtaining the prior vector includes: processing the data set level input and response of the task through the response-guided GenAI decision prior estimator to obtain the prior vectors of all instances in the task.
3. The method for explaining generative artificial intelligence model decisions as claimed in claim 2, characterized in that: The response-guided GenAI decision prior estimator includes a word embedding layer, a similarity calculation layer, a mean calculation layer, and a normalization layer; In the response-guided GenAI decision prior estimator, generated by the word embedding layer The word embedding vectors of the input and response at the dataset level of the task; the candidate priors for each feature in the input are estimated by calculating the average cosine similarity between each feature in the input document and each word in the response through the similarity calculation layer; Then, the co-occurrence probability is counted, and the average prior of all features contained in the task document is calculated through the mean calculation layer to obtain the task-level prior; the task-level prior is processed through the normalization layer to obtain the prior vector of all instances in the task.
4. The method for explaining generative artificial intelligence model decisions according to claim 1, characterized in that: When the input data is text data, the first feature embedding layer uses the GloVe word embedding method to capture the semantics of each word and obtain the embedding vector of each word.
5. The method for explaining generative artificial intelligence model decisions as claimed in claim 1, characterized in that: The embedding element-wise convolutional layer uses a CNN approach to map each enhanced word embedding vector to an attribution score.
6. The method for explaining generative artificial intelligence model decisions as claimed in claim 1, characterized in that: The sparse feature layer uses the Gumbel-softmax technique to sparse the fused attribution scores, specifically including: Construct a set of specific random vectors as a differentiable approximation of the argmax operation, where the set of specific random vectors is defined as: The calculation method is: in, Indicates i A random perturbation, represents the temperature parameter, Indicates i A random vector; Indicates n A random vector; Indicates i The final fused attribution score for the word; Perform k independent subset sampling to perform argmax operation, generate k continuous relaxation vectors for each random vector, take the maximum value element by element, get a continuous relaxation random vector, and further construct a continuous relaxation random vector group : 。 7. A generative artificial intelligence model decision explanation system, characterized in that: include: Input layer, used to obtain input data and responses; The model processing and output layer is used to process the input data and responses through the pre-built GenAI decision explanation model to obtain explanations; The pre-built GenAI decision explanation model includes an explainer and a response-aware attention mechanism; The response-aware attention mechanism is used to process the response, and during the processing, an enhanced internal feature vector obtained by processing the input data by an interpreter is introduced to obtain an external attribution score; The interpreter is used to process the input data and the external attribution score to obtain a continuous relaxed random vector, and obtain an interpretation after performing a dot product calculation on the continuous relaxed random vector and the embedding vector of the input data; The interpreter includes a first feature embedding layer, a first embedding enhancement layer, an embedding element-level convolution layer, a first normalization layer, a first mean calculation layer, a second normalization layer, and a sparse feature layer; The response-aware attention mechanism includes a second feature embedding layer, a second embedding enhancement layer, a second mean calculation layer, and a third normalization layer; The input data is processed by a first feature embedding layer, a first embedding enhancement layer, an embedding element-level convolution layer, and a first normalization layer in the interpreter to obtain an internal feature attribution score; The response is passed through a second feature embedding layer and a second embedding enhancement layer in a response-aware attention mechanism to obtain an enhanced external feature vector, wherein the enhanced external feature vector is aligned with the enhanced internal feature vector output by the first embedding enhancement layer; The enhanced internal feature vector and the enhanced external feature vector are subjected to dot product calculation to obtain a dot product; the dot product passes through the second mean calculation layer and the third normalization layer to obtain an external attribution score; The internal feature attribution score and the external attribution score are passed through the first mean calculation layer and the second normalization layer to obtain a fused attribution score; after the fused attribution score passes through the sparse feature layer, a continuous relaxed random vector is obtained, and the continuous relaxed random vector and the input data embedding vector output by the first feature embedding layer are interpreted after the dot product calculation.
8. A computer-readable storage medium, characterized in that: It stores a computer program for explaining generative artificial intelligence model decisions, wherein the computer program enables a computer to execute a method for explaining generative artificial intelligence model decisions as described in any one of claims 1 to 6.
9. An electronic device, characterized in that: include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the programs include an interpretation method for executing the generative artificial intelligence model decision as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Visual interpretation method and system for deep neural network model.
CN112861933A
Graph neural network interpretation method, system and equipment based on evolution integration
CN117669742A