An algorithm for evaluating the alignment degree between the fine-grained decision-making logic of an artificial intelligence black-box model and human cognition

By analyzing the interaction distribution and computational interaction utility of the black box model, the problem of difficulty in evaluating the alignment of the decision logic of the black box model with human cognition in the prior art is solved, and effective reflection and evaluation of the potential representation defects of the model are achieved.

CN119150077BActive Publication Date: 2025-06-17QUANXIN QUANYI INTELLIGENT TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411386656.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2025-06-17
Estimated Expiration
2044-09-30

AI Technical Summary

Technical Problem

The prior art is difficult to evaluate the alignment of the fine decision logic of the artificial intelligence black box model with human cognition, and cannot effectively reflect the potential representation defects of the model.

Method used

By analyzing the interactive distribution modeled by the model, using interaction utility calculation and division methods, the alignment of the model decision logic with human cognition is evaluated. The specific steps include selecting the black box model, inputting sample decomposition, calculating interactive utility, dividing reliable and unreliable components, and evaluating indicator statistics.

Benefits of technology

A reliable evaluation of the alignment of the fine decision logic of the black box model with human cognition is realized, reflecting the potential representation defects of the model and improving the credibility of the model in high-risk scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119150077B_ABST
    Figure CN119150077B_ABST
Patent Text Reader

Abstract

This application relates to the field of machine learning technology, and discloses a method and system for evaluating the alignment degree between the fine decision-making logic of an artificial intelligence black-box model and human cognition. The method and system can evaluate the alignment degree between the model decision-making logic and human cognition by analyzing the interaction distribution modeled by the model. The implementation of the method and system includes the following steps: providing input samples; using the black-box model to predict the input samples to obtain the prediction results of the model; based on the output of the black-box model, modeling the interaction between the input units of the samples, calculating the interaction intensity of the combinations formed between the input units, and expressing the black-box model output as the "interaction utility" between the input unit combinations; using interaction-based evaluation metrics to evaluate the alignment degree between the model decision-making logic and human cognition. The advantage of the present invention is that it provides a quantitative method for evaluating the alignment degree between the fine decision-making logic of an artificial intelligence black-box model and human cognition. Compared with previous studies, it can better characterize the differences between the model decision-making logic and human cognition and reflect the potential representation defects of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and particularly to a method and a system for evaluating the alignment degree between the fine decision logic of an artificial intelligence black-box model and human cognition. Background Art

[0002] At present, deep learning has been widely applied in various fields and demonstrated powerful performance, but it is difficult for users and researchers to analyze the black-box nature of artificial intelligence models. In the existing technologies, the interaction-based explanation method can analyze an artificial intelligence model to obtain sparse and concise explanations. However, the obtained explanations still cannot directly evaluate the difference between the decision logic of the black-box model and human cognition, nor reflect the potential representation defects of the model. And evaluating the alignment degree between the decision logic of the model and human cognition is exactly a necessary condition for humans to trust the black-box model and apply the black-box model in high-risk scenarios.

[0003] Therefore, how to reliably evaluate the alignment degree between the fine decision logic of the black-box model and human cognition is an urgent problem to be solved in the field of interpretability. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and a system for evaluating the alignment degree between the fine decision logic of an artificial intelligence black-box model and human cognition, which can evaluate the alignment degree between the decision logic of the model and human cognition by analyzing the interaction distribution modeled by the model.

[0005] The present invention discloses a method for evaluating the alignment degree between the fine decision logic of an artificial intelligence black-box model and human cognition, which is characterized in that the method comprises the following steps:

[0006] (1) Select a black-box artificial intelligence model;

[0007] Select a black-box artificial intelligence model to be analyzed, and the black-box artificial intelligence model includes an artificial intelligence model pre-trained on a certain data set;

[0008] (2) Select input samples and perform recognition;

[0009] Select input samples for calculating interaction utility, and perform recognition on the input samples, so as to decompose the input samples into n input units, and combine the n input units to obtain combinations of 2 n combinations of the input units; the input samples are selected from the following group: tabular data, pictures, texts, voices or combinations thereof;

[0010] (3) Calculate and combine "interaction utility";

[0011] Combine the 2 in step (2) nThe combinations of the input units are respectively input into the black-box artificial intelligence model, and the output of the black-box artificial intelligence model is obtained; based on the output of the black-box artificial intelligence model, the interactions between the input units are modeled, so as to obtain the "interaction utility" of the black-box artificial intelligence model for each combination of the input units; and the output of the black-box artificial intelligence model on a certain combination of input units is interpreted as the combination of its "interaction utilities" on the combination of input units.

[0012] (4) Divide the "interaction utility" into a "reliable component of the interaction utility" and an "unreliable component of the interaction utility".

[0013] According to whether each input unit is cognitively relevant to the output of the black-box artificial intelligence model, all input units are divided into "relevant input units", "irrelevant input units" and "mutually exclusive units" by using an algorithm or user-defined.

[0014] Divide the "interaction utilities" of all combinations of input units in step (3) into a "reliable component of the interaction utility" and an "unreliable component of the interaction utility". Among them, the "reliable component of the interaction utility" represents the interaction utility consistent with human cognition, and the "unreliable component of the interaction utility" represents the interaction utility inconsistent with human cognition. The interaction utility can be divided into "AND interaction utility" and "OR interaction utility", which respectively represent the "AND relationship" between the input units modeled by the model and the "OR relationship" between the input units modeled by it. Preferably, the "reliable component of the AND interaction utility" can be calculated as the interaction utility of the combination that includes "relevant input units" and does not include "mutually exclusive units"; the "unreliable component of the AND interaction utility" can be calculated as the interaction utility of the combination that does not include "relevant input units" or includes "mutually exclusive units". Preferably, a "reliable component of the OR interaction utility" can be calculated as the interaction utility of the combination that includes "relevant input units" and does not include "mutually exclusive units"; the "unreliable component of the OR interaction utility" can be calculated as the interaction utility of the combination that does not include "relevant input units" or includes "mutually exclusive units". Another algorithm is to evenly distribute the "OR interaction utility" to each input unit included in this interaction, so that the "reliable component of the OR interaction utility" can be calculated as the interaction utility component evenly distributed to the "relevant input units" in this "OR interaction utility"; the "unreliable component of the OR interaction utility" can be calculated as the interaction utility component evenly distributed to the "irrelevant input units" and "mutually exclusive units" in this "OR interaction utility".

[0015] (5) Use interaction-based evaluation metrics to evaluate the alignment degree between the model decision logic and human cognition.

[0016] According to the proposed interaction-based evaluation metrics, the proportions of the "reliable components of interaction utility" and the "unreliable components of interaction utility" are statistically calculated to evaluate the alignment degree between the fine-grained decision-making logic of the model and human cognition.

[0017] Among them, the order of steps (1) and (2) can be arbitrarily replaced or carried out simultaneously.

[0018] In a preferred example, the output of the black-box artificial intelligence model on the combination of input units is defined as follows: It is defined that if the input units in the set N\S in the input sample x are occluded without occluding the input units in the set S, the sample x S . is obtained. The output obtained by using x S as the model input is denoted as v(x S ); if the input sample is input into the black-box model without occlusion, the obtained output is denoted as v(x N ); if all input units in the input sample are completely occluded and then input into the black-box model, the obtained output is denoted as Preferably, in natural language processing applications, the occlusion operation can be achieved by replacing the word embedding vectors of each token covered on each input unit included in S with a certain specific reference value vector.

[0019] In a preferred example, if the black-box artificial intelligence model is a classification model, the output v(x S ) of the model on a certain input unit set S can be expressed as the logit value corresponding to the dimension of the true label, that is, If the black-box artificial intelligence model is a language generation model, the model output can be expressed as the logit value corresponding to the dimension of the next token predicted by the model with the highest probability of generation, that is, where y max represents the dimension corresponding to the next token predicted by the model with the highest probability of generation; the model output v(x S ) can also be expressed as the sum of the logits values corresponding to the input dimensions of each token with the highest probability (or pre-determined target token) in the subsequent T tokens generated. Here, the subsequent T target tokens generated are denoted as y1, y2,..., y T , then v(x S ) can be calculated as

[0020]

[0021] Similarly, v(x S ) can also be calculated as

[0022]

[0023] In a preferred example, step (3) further includes the following steps: interpreting the output v(x S ) of the black-box artificial intelligence model on a certain input unit set S as a combination of its "interaction utilities" on the input unit combination. Preferably, the "interaction utilities" include "AND interaction utility" and "OR interaction utility".

[0024] In a preferred example, the "AND interaction utility" represents "the additional utility generated by the combination for the output of a black-box model when and only when none of the units in the combination of input units are occluded";

[0025] In a preferred example, for the black-box artificial intelligence model, the combination of input units defines I and (S|v and ,x) as the "AND interaction utility" corresponding to the combination S of input units for the artificial intelligence model, which can be calculated by the following formula:

[0026]

[0027] Here, v and (x L ) represents the output component determined by the "AND interaction utility" decomposed from v(x L ), corresponding to the formula v(x L ) = v and (x L ) + v or (x L ), and v or (x L ) represents the output component determined by the "OR interaction utility" decomposed from v(x L ).

[0028] In a preferred example, the "OR interaction utility" represents "the additional utility generated by the combination for a set of black-box model outputs when at least one of the units in the combination of input units is not occluded";

[0029] In a preferred example, for the black-box artificial intelligence model, the combination of input units defines I or (S|v or ,x) as the "OR interaction utility" corresponding to the combination S of input units for the artificial intelligence model, which can be calculated by the following formula.

[0030]

[0031] Here, vor (x L ) represents the output component determined by the "OR interaction effect" decomposed from v(x L ), corresponding to the formula v(x L ) = v and (x L ) + v or (x L ), and v and (x L ) represents the output component determined by the "AND interaction effect" decomposed from v(x L ).

[0032] The output of the neural network can be expressed as the sum of different interactions. In a preferred example, the output

[0033] v(x S ) of the black box artificial intelligence model on the combination S of input units can be interpreted as a combination of "AND interaction effect" and "OR interaction effect", that is, v(x S ) =

[0034] v and (x S ) + v or (x S ). The output v(x S ) of the black box artificial intelligence model on the combination S of input units can also be interpreted as a combination of "AND interaction effect", that is, v(x S ) = v and (x S ). It should be understood that, without ambiguity, the "interaction effect" below can represent the combination of all "AND interaction effects" and "OR interaction effects".

[0035] In a preferred example, the step (4) further includes the following steps: dividing all input units into "relevant input units", "irrelevant input units" and "mutually exclusive units". Define the interaction effect of the combination that includes "relevant input units" and does not include "mutually exclusive units" as the "reliable component of the interaction effect", and define the interaction effect of the combination that includes "irrelevant input units" or "mutually exclusive units" as the "unreliable component of the interaction effect".

[0036] In a preferred example, according to whether each input unit is relevant to the output of the black box artificial intelligence model in human cognition, the set N of all input units is divided into the set R of "relevant input units", the set T of "irrelevant input units" and the set M of "mutually exclusive units" by using an algorithm or user-defined, Among them, the "relevant input unit" refers to an input unit that is strongly correlated with the model output in human cognition or has a direct causal relationship; the "irrelevant input unit" refers to an input unit that is weakly correlated with the model output in human cognition or has no direct causal relationship; the "mutually exclusive unit" refers to an input unit that should not affect the model output in human cognitive logic or is mutually exclusive with the model output in cognitive logic. Some input samples may contain only one or more of the "relevant input unit", "irrelevant input unit" and "mutually exclusive unit". Preferably, for text input, the input unit can be the word embedding vector of a token, or all the word embedding vectors corresponding to a word, phrase, sentence, or paragraph; for image input, the input unit can be a pixel or an image region.

[0037] In a preferred example, in the task of evaluating the black-box artificial intelligence model, the user can label the input units in the vocabulary as "relevant input units", "irrelevant input units" and "mutually exclusive units" from the perspective of human cognition; the user can also use an algorithm to automatically determine or label the sets of "relevant input units", "irrelevant input units" and "mutually exclusive units"; the user can use an attribution algorithm including the Shapley value to measure the importance of all input units to the model output, and automatically divide all input units into "relevant input units", "irrelevant input units" and "mutually exclusive units" according to the importance scores; the user can also use a complex artificial intelligence algorithm to measure the correlation degree between all input units and the model output at the semantic logic level, and automatically divide all input units into "relevant input units", "irrelevant input units" and "mutually exclusive units" according to the correlation scores.

[0038] In a preferred example, the "interaction utility" of all combinations of input units in step (4) is further divided into a "reliable component of the interaction utility" and an "unreliable component of the interaction utility". Among them,

[0039] The "reliable component of the interaction utility" is the interaction utility consistent with human cognition, which can be defined as the interaction utility of combinations that contain "relevant input units" and do not contain "mutually exclusive units"; the "unreliable component of the interaction utility" is the interaction utility inconsistent with human cognition, which can be defined as the interaction utility of combinations that contain "irrelevant input units" or "mutually exclusive units".

[0040] In a preferred example, the interaction utility can be divided into an "AND interaction utility" and an "OR interaction utility", which respectively represent the "AND relationship" between the input units modeled by the model and the "OR relationship" between the input units modeled by it. The "AND interaction utility" I and (S|v and ,x) can be further divided into the "reliable component of the AND interaction utility" and "the unreliable component of the interaction utility" It can be calculated by the following formula:

[0041]

[0042] Among them, the set N of all input units is divided into the set R of "relevant input units", the set T of "irrelevant input units", and the set M of "mutually exclusive units".

[0043] In a preferred example, the interaction utility can be divided into "AND interaction utility" and "OR interaction utility", which respectively represent the "AND relationship" between the input units modeled by the model and the "OR relationship" between the input units modeled by it. The "OR interaction utility" I or (S|v or ,x) can be further divided into "the reliable component of the OR interaction utility" and "the unreliable component of the OR interaction utility" Preferably, the division of an "OR interaction utility" is similar to the division of an "AND interaction utility" and can be calculated by the following formula:

[0044]

[0045]

[0046] Among them, the set N of all input units is divided into the set R of "relevant input units", the set T of "irrelevant input units", and the set M of "mutually exclusive units".

[0047] Another division of the "OR interaction utility" considers the proportion of the number |S∩R| of "relevant input units" in the combination S to the number |S| of input units in the combination S and can be calculated by the following formula:

[0048]

[0049] In a preferred example, step (5) further includes the following steps: using the proposed evaluation index based on interaction utility, counting the proportion of the "reliable component of the interaction utility" and the "unreliable component of the interaction utility", and evaluating the alignment degree between the fine-grained decision logic of the model and human cognition. The user can quantify the alignment degree between the fine-grained decision logic of the model and human cognition by the proportion of the "reliable component of the interaction utility" in the interaction utility.

[0050] In a preferred example, significant interaction is defined as follows. Given a preset threshold τ, significant interaction is defined as the set of interactions Ω and and Ω or , which can be calculated by the following formula:

[0051]

[0052] Among them, the threshold τ can be set to the percentage output by the black-box artificial intelligence model. Preferably, the threshold τ can also be set to the numerical value of the k-th interaction utility absolute value after sorting the absolute values of all interaction utilities from large to small. Preferably, in an input sample containing 10 input units, there are a total of 2 "AND interactions"; or simultaneously contain 2 10 "AND interactions" and 2 10 "OR interactions", and the threshold τ distinguishes a certain proportion of "significant AND interactions" and "significant OR interactions". The threshold τ can also be set according to user experience, such as setting it to the 100th largest interaction utility absolute value among the absolute values of all interaction utilities. 10

[0053] In a preferred example, the sum of the intensities of all significant interaction utilities is:

[0054]

[0055] In a preferred example, the sum of the intensities of all interaction utilities is:

[0056]

[0057] In a preferred example, the order of the interaction is defined as the number of input variables in the set S, that is, order(S) = |S|. For each order of interaction o, the sum of the reliable components of the interaction utility is:

[0058]

[0059] In a preferred example, for each order of interaction o, the intensity of the reliable component of the interaction utility is:

[0060]

[0061] In a preferred example, the total intensity of the reliable components of the interaction utility is:

[0062]

[0063] In a preferred example, for each order of interaction o, the sum of the unreliable components of the interaction utility is:

[0064]

[0065] In a preferred example, for each order of interaction o, the intensity of the unreliable component of the interaction utility is:

[0066]

[0067] ​In a preferred example, the total intensity of the unreliable components of the interaction utility is:

[0068]

[0069] In a preferred example, the following index represents the proportion of the intensity of the "reliable component of the interaction utility" in the intensity of the significant interaction utility.

[0070]

[0071] Among them, the proportion of the intensity of the "reliable component of the interaction utility" in the intensity of the interaction utility of the significant interaction The higher it is, the higher the degree of alignment between the fine decision-making logic of the model and human cognition. For a model with a higher degree of alignment with human cognition, the index should be higher.

[0072] In a preferred example, the following index represents the proportion of the intensity of the "reliable component of the interaction utility" in the intensity of the interaction utility of all interactions.

[0073]

[0074] Among them, the proportion of the intensity of the "reliable component of the interaction utility" in the intensity of the interaction utility of all interactions The higher it is, the higher the degree of alignment between the fine decision-making logic of the model and human cognition. For a model with a higher degree of alignment with human cognition, the index should be higher.

[0075] In a preferred example, the following index and respectively represent the proportion of the "reliable component of the interaction utility" in the sum of the positive utilities and the sum of the negative utilities of the significant interactions of each order.

[0076] Or

[0077] Or

[0078] Among them, ∈ represents a very small positive real number to prevent the division by zero operation of the fraction. For each order of interaction o, calculate the sum of the positive interaction utilities of the significant interactions and the sum of the negative interaction utilities of the significant interactions Calculate the positive interaction utility of the "reliable component of the interaction utility" and the negative interaction utility of the "reliable component of the interaction utility" Index and The higher it is, that is, for each order of interaction o, the higher the proportion of the "reliable component of the interaction utility" in the positive and negative utilities of the significant interaction, the higher the degree of alignment between the fine decision-making logic of the model and human cognition. For a model with a higher degree of alignment with human cognition, the proportion in the sum of the positive and negative utilities of each order of significant interaction or index should be higher.

[0079] In a preferred example, the following index represents the proportion of positive and negative cancellation in the utility of each order of significant interaction.

[0080]

[0081] Among them, the index The higher it is, the higher the proportion of the effective interaction remaining after the positive and negative cancellation of the significant interaction utility for each order of interaction o. This index is helpful for understanding whether low-order interaction or high-order interaction plays a major role in the model output. A well-trained model tends to model low-order interactions more.

[0082] In a preferred example, the following index O reliable represents the weighted average order of "reliable interactions" in the significant interactions.

[0083] or

[0084]

[0085] Among them, the weighted average order of "reliable interactions" O reliable The lower it is, the more easily the influence of the fine decision-making logic of the model is affected by low-order interactions; on the contrary, the higher the index O reliable The higher it is, the more easily the influence of the fine decision-making logic of the model is affected by high-order interactions. For a model with a higher degree of alignment with human cognition, the weighted average order of "reliable interactions" O in the significant interactions reliable ∈[0,|N|] should be lower, indicating that the model's modeling tendency is to model lower-order "reliable interactions".

[0086] In a preferred example, the differences between the model decision-making logic and human cognition can be represented by visualizing the statistical charts of the utility of each order of significant interaction, the reliable component of the interaction utility, and the unreliable component of the interaction utility, reflecting the potential representational defects of the model. Specifically, for each order of interaction o, the sum of the positive interaction utilities of the significant interaction Salient + (o), the sum of the negative interaction utilities of the significant interaction Salient - (o), the sum of the positive utilities of the reliable interaction component Reliable +(o), the sum of the negative utilities of the reliable interaction components Reliable - (o), the sum of the positive utilities of the unreliable interaction components Unreliable + (o): = Salient + (o) - Reliable + (o), the sum of the negative utilities of the unreliable interaction components Unreliable - (o): = Salient - (o) - Reliable - (o).

[0087] For each order of interaction o, when the indicators and are higher, that is, the higher the proportion of the reliable component of the interaction utility in the significant interaction, the higher the degree of alignment between the fine-grained decision-making logic of the model and human cognition.

[0088] In summary, statistical charts of the significant interaction utilities of each order, the reliable component of the interaction utility, and the unreliable component of the interaction utility, as well as the indicators and can be used to characterize the differences between the model decision-making logic and human cognition and reflect the potential representational defects of the model.

[0089] The second aspect of the present invention discloses a system for evaluating the degree of alignment between the fine-grained decision-making logic of an artificial intelligence black-box model and human cognition. The system includes:

[0090] (1) An input module, which is configured to receive a pre-trained black-box artificial intelligence model and a set of data to be analyzed.

[0091] (2) A calculation module, which is configured to calculate the "interaction utility" between the input units of the data to be analyzed modeled by the model based on the model and data in the input module. Divide the "interaction utility" into a "reliable component of the interaction utility" and an "unreliable component of the interaction utility". Use the proposed interaction-based evaluation indicators to count the proportions of the "reliable component of the interaction utility" and the "unreliable component of the interaction utility".

[0092] (3) An output module, which is configured to evaluate the degree of alignment between the fine-grained decision-making logic of the model and human cognition based on the interaction-utility-based evaluation indicators and the statistics of the significant interaction utilities of each order, the reliable component of the interaction utility, and the unreliable component of the interaction utility.

[0093] The advantages of the present invention are:

[0094] 1) The present invention realizes a reliable evaluation of the alignment degree between the fine decision-making logic of the black-box model and human cognition by reasonably classifying the input components and introducing the definition and calculation of reliable components, thereby reflecting the potential representation defects of the model.

[0095] 2) The present invention introduces multiple criteria to evaluate the alignment degree between the fine decision-making logic of the black-box model and human cognition, ensuring that the method used in the present invention is applicable to most black-box models.

[0096] A large number of technical features are recorded in the description of this application, distributed in various technical solutions. If all possible combinations of technical features (i.e., technical solutions) of this application are listed, the description will be too long. To avoid this problem, each technical feature disclosed in the above-mentioned invention content of this application, each technical feature disclosed in the following embodiments and examples, and each technical feature disclosed in the drawings can be freely combined with each other to form various new technical solutions (these technical solutions are all regarded as having been recorded in this specification), unless the combination of such technical features is technically infeasible. For example, in one example, features A + B + C are disclosed, and in another example, features A + B + D + E are disclosed, and features C and D are equivalent technical means that play the same role. Only one of them can be used technically and it is impossible to use both at the same time. Feature E can be combined with feature C technically. Then, the solution of A + B + C + D should not be regarded as having been recorded due to technical infeasibility, while the solution of A + B + C + E should be regarded as having been recorded. BRIEF DESCRIPTION OF THE DRAWINGS

[0097] Figure 1 is a schematic flowchart of evaluating the alignment degree between the fine decision-making logic of the artificial intelligence black-box model and human cognition according to the first embodiment of the present invention;

[0098] Figure 2 is a schematic diagram showing the absolute utilities of "AND interaction utility" and "OR interaction utility" obtained according to the present invention arranged in descending order;

[0099] Figure 3a is an evaluation index based on interaction utility obtained according to the present invention;

[0100] Figure 3b is a statistical chart of the significant interaction utilities of each order, the reliable components of the interaction utility, and the unreliable components of the interaction utility; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0101] After careful and in-depth research, the inventor of the present invention has developed for the first time a method and system for realizing the evaluation of the alignment degree between the fine decision-making logic of the artificial intelligence black-box model and human cognition.

[0102] It should be noted that in the invention document of this patent, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one" does not exclude the presence of additional identical elements in the process, method, article or device comprising that element. In the invention document of this patent, if it is mentioned that an act is performed according to a certain element, it means that the act is performed at least according to that element, including two cases: the act is performed only according to that element, and the act is performed according to that element and other elements. Expressions such as multiple, many times, various, etc. include 2, 2 times, 2 kinds, as well as more than 2, more than 2 times, more than 2 kinds.

[0103] The present invention includes the following steps:

[0104] (1) Select a black-box artificial intelligence model

[0105] Select a black-box artificial intelligence model to be analyzed, and the black-box artificial intelligence model includes an artificial intelligence model pre-trained on a certain data set.

[0106] (2) Select an input sample and perform identification

[0107] Select an input sample for interactive utility calculation, and perform identification on the input sample, so as to decompose the input sample into n input units, and combine the n input units to obtain 2 n combinations of the input units; the input sample is selected from the following group: tabular data, pictures, texts, voices or combinations thereof;

[0108] (3) Calculate the "interactive utility" and perform combination;

[0109] Input the 2 n combinations of input units in step (2) into the black-box artificial intelligence model respectively, and obtain the output of the black-box artificial intelligence model; based on the output of the black-box artificial intelligence model, model the interaction between the input units, so as to obtain the "interactive utility" of the black-box artificial intelligence model for each combination of the input units; and interpret the output of the black-box artificial intelligence model on a certain combination of input units as the combination of its "interactive utility" on the combination of input units;

[0110] (4) Divide the "interaction utility" into a "reliable component of the interaction utility" and an "unreliable component of the interaction utility".

[0111] All input units are divided into "relevant input units", "irrelevant input units" and "mutually exclusive units" algorithmically or user-defined according to whether each input unit is cognitively relevant to the output of the black-box artificial intelligence model.

[0112] Divide the "interaction utility" of all combinations of input units in step (3) into a "reliable component of the interaction utility" and an "unreliable component of the interaction utility". Among them, the "reliable component of the interaction utility" is the interaction utility consistent with human cognition, which can be expressed as the interaction utility of a combination that includes "relevant input units" and does not include "mutually exclusive units"; the "unreliable component of the interaction utility" is the interaction utility inconsistent with human cognition, which can be expressed as the interaction utility of a combination that includes "irrelevant input units" and "mutually exclusive units". Preferably, the interaction utility can be divided into "AND interaction utility" and "OR interaction utility", which respectively represent the "AND relationship" between the input units modeled by the model and the "OR relationship" between the input units modeled by it. Preferably, the "reliable component of the AND interaction utility" can be calculated as the interaction utility of a combination that includes "relevant input units" and does not include "mutually exclusive units"; the "unreliable component of the AND interaction utility" can be calculated as the interaction utility of a combination that does not include "relevant input units" or includes "mutually exclusive units". Preferably, a "reliable component of the OR interaction utility" can be calculated as the interaction utility of a combination that includes "relevant input units" and does not include "mutually exclusive units"; the "unreliable component of the OR interaction utility" can be calculated as the interaction utility of a combination that does not include "relevant input units" or includes "mutually exclusive units". Another algorithm is to evenly distribute the "OR interaction utility" to each input unit included in this interaction. Thus, the "reliable component of the OR interaction utility" can be calculated as the interaction utility component evenly distributed to the "relevant input units" in this "OR interaction utility"; the "unreliable component of the OR interaction utility" can be calculated as the interaction utility component evenly distributed to the "irrelevant input units" and "mutually exclusive units" in this "OR interaction utility".

[0113] (5) Use the interaction-based evaluation metrics to evaluate the alignment degree between the model decision logic and human cognition;

[0114] According to the proposed interaction-based evaluation metrics, count the proportions of the "reliable component of the interaction utility" and the "unreliable component of the interaction utility", and evaluate the alignment degree between the fine decision logic of the model and human cognition.

[0115] Embodiment

[0116] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe the implementation manners of this application in detail with reference to the accompanying drawings.

[0117] The first embodiment of the present invention relates to a method for evaluating the alignment degree between the fine decision-making logic of an artificial intelligence black-box model and human cognition. The process is as Figure 1 shown, and this method includes the following steps:

[0118] In step 101: Based on a certain data set, train a black-box artificial intelligence model as the model to be analyzed. Optionally, a model can be a deep neural network.

[0119] After that, enter step 102, and this step can be further divided into the following two sub-steps:

[0120] (a) Collect the data to be analyzed. Optionally, if the model in step (1) can be used for natural language generation, the data to be analyzed can be a text.

[0121] (b) Preprocess the data to be analyzed. Optionally, the data to be analyzed can be tokenized, encoded into embedding vectors, etc. to adapt it to the input of the above model.

[0122] After that, enter step 103, and this step can be further divided into the following five sub-steps:

[0123] (a) Calculate the output value of the black-box model under any occlusion of the above data to be analyzed;

[0124] (b) Calculate the "AND interaction utility" in the data to be analyzed modeled by the black-box model;

[0125] (c) Calculate the "OR interaction utility" in the data to be analyzed modeled by the black-box model;

[0126] (d) Based on the "AND interaction utility" and "OR interaction utility", interpret the output of the black-box model as a combination of the "AND interaction utility" and "OR interaction utility" between input unit combinations;

[0127] (e) Based on the combination of the "AND interaction utility" and "OR interaction utility", further optimize the combination of the "AND interaction utility" and "OR interaction utility" of the model to make the interaction sparser and more concise.

[0128] In the above sub-step (a), the black-box artificial intelligence model is denoted as v, and the data sample x to be analyzed contains n input units, denoted as the set N = {1, 2,..., n}, where the number of input units is generally 5 to 20. Optionally, each token in the input data can be regarded as an input unit, or the input data can be divided into multiple phrases or short sentences, and the word embedding vectors of all tokens covered by each phrase or short sentence can be regarded as an input unit. For extremely long input samples, all tokens covered by multiple sentences or multiple paragraphs and sections can be counted as the same input unit, and different input units correspond to different paragraphs and sections in the input sample.

[0129] Any subset of the input units is called a combination among the input units, and v(x S ) represents the output of the black-box model when given a certain combination S among the input units and when the input units in N\S are occluded. In particular, if the input sample is input into the model without occlusion, the obtained output is denoted as v(x N ), and if the output obtained after completely occluding all input units in the input sample is denoted as

[0130] When calculating the output v(x S ) of the black-box model on a certain input unit combination S, it is necessary to keep the input values of the samples to be analyzed on the input units included in S, and replace all input units in the complement N\S of S with the reference value. In particular, for natural language processing tasks, assume that the reference value is b, and assume that each input unit contains only one token. Then, the respective input units of the occluded sample x S are defined by the following formula:

[0131]

[0132] Here, (x S ) i = x i means that the word embedding vector corresponding to the i-th input unit (i.e., the i-th token) of the occluded sample remains unchanged, which is still the word embedding vector of the i-th input unit (the i-th token) in the original input sample; (x S ) i = b i means that the word embedding vector corresponding to the i-th input unit (i.e., the i-th token) of the occluded sample is replaced with the word embedding vector reference value.

[0133] In this way, v(x S ) can be calculated as substituting the sample x SThe output value obtained from the input model. Optionally, the baseline value can be set to the mean of the sample set, a random value, zero, etc., or it can be obtained through learning. For natural language processing tasks, the baseline value can be set to a special word embedding vector, and the occlusion of a single token can be achieved by replacing the word embedding vector of this token with a special word embedding vector.

[0134] If the i-th input unit includes multiple tokens (for example, the i-th input unit is a word, phrase, short sentence, etc.), then the occlusion of the i-th input unit needs to simultaneously replace the word embedding vectors corresponding to all the tokens it contains. Further, the output v(x S ) of the black box model calculated based on the baseline value b can be specifically denoted as v(x S |b). It should be understood that, without ambiguity, v(x S |b) can be abbreviated as v(x S ).

[0135] We can use "AND interaction utility" and "OR interaction utility" to explain the output of the model. Given the output v of the black box model and an input sample x, when calculating the output v(x S ) on a certain combination S of input units, the model output v(x S ) is decomposed into two terms, that is, v(x S ) = v and (x S ) + v or (x S ). In this way, we can use AND interaction to explain v and (x S ), and use OR interaction to explain v or (x S ).

[0136] In the above sub-step (b), by calculating the "AND interaction utility", the contribution value of the AND interaction between any combination of input units in the sample to be analyzed to the output v and (x) corresponding to the AND interaction of each model is evaluated. The "AND interaction utility" represents the additional utility generated by the units in the combination when they are all triggered (i.e., not occluded) for the output corresponding to the AND interaction of a black box model. It should be understood that in a black box model, input units often do not contribute to the model output individually, but interact with each other to affect the model output.

[0137] Therefore, for define I and (S|v and ,x) as the "AND interaction utility" corresponding to the combination S of input units of an artificial intelligence black box model, such that when all input variables in S are not occluded, Iand (S|v and , x) is triggered and accumulated into the output v corresponding to the interaction of the black-box model and (x). Specifically, I and (S|v and , x) can be calculated by the following formula.

[0138]

[0139] The above "interaction utility" satisfies that for any given combination of input units of the sample to be analyzed The output v corresponding to the interaction of an artificial intelligence model under this combination of input units and (x T ) can be split into the sum of all triggered interactions I and (S|v and , x), as shown in the following formula.

[0140]

[0141] Where, e T represents whether the input variable is occluded. When i ∈ T, (e T ) i = 1, indicating that the input variable i is not occluded (or is triggered), otherwise, (e T ) i = 0 indicates that the input variable i is occluded; ∧ i∈S (e T ) i represents whether all the units in the input unit combination S are not occluded. The result obtained by performing the "AND" operation on the elements after the ∧ symbol, that is, ∧ i∈S (e T ) i = ∏ i∈S (e T ) i .

[0142] Furthermore, the above calculation formula of the interaction can be equivalently expressed as the following form:

[0143] Suppose and exhaust all interactions of the input unit combinations, then there exists such that I and = Av and .

[0144] In the above sub-step (c), by calculating the "OR interaction utility" to evaluate the OR relationship between any input unit combinations of the sample to be analyzed for the output v of the OR interaction of the black-box model orThe contribution value of (x). "Or interaction utility" means the additional utility generated by the or interaction output of the black box model when at least one of the units in the combination is triggered (i.e., not occluded). It should be understood that in a black box model, input units often do not contribute to the model output individually, but interact with each other to affect the model output.

[0145] Therefore, for Define I or (S|v or ,x) as the "or interaction utility" corresponding to the combination S of input units, such that when at least one of all the input variables in S is not occluded, I or (S|v or ,x) is triggered and accumulated into the or interaction output v or (x) of the black box model. Specifically, I or (S|v or ,x) can be calculated by the following formula.

[0146]

[0147] The above "or interaction utility" satisfies that for any given combination of input units of the sample to be analyzed The or interaction output v or (x T ) of the black box model under this combination of input units can be split into the sum of all triggered or interactions I or (S|v or ,x), as shown in the following formula.

[0148]

[0149]

[0150] Where e T represents whether the input variable is not occluded. When i ∈ T, (e T ) i = 1, which means the input variable i is not occluded. Otherwise, (e T ) i = 0, which means the input variable i is occluded at this time. ∨ i∈S (e T ) i represents whether at least one of the units in the input unit combination S is not occluded, and is the result obtained by performing an "or" operation on the elements after the ∨ symbol, that is, V i∈S (e T ) i = 1 - ∏ i∈S [1 - (e T ) i .

[0151] Furthermore, the above-mentioned or interactive calculation formula can be equivalently expressed in the following form. Assuming that exhausts the interactions over all input unit combinations, then there exists such that I or = Bv or .

[0152] In an embodiment of the present invention, using the "AND interaction utility" and "OR interaction utility" calculation formulas, the values of the "AND interaction utility" and "OR interaction utility" of each model are calculated respectively. As Figure 2 shown, it respectively shows a schematic diagram of the distribution of the absolute values of the "AND interaction utility" and "OR interaction utility" of different input unit combinations obtained by the present invention and the traditional method, arranged in descending order.

[0153] In the above sub-step (d), based on the "AND interaction utility" and "OR interaction utility", the output of the black-box model is interpreted as a combination of the "AND interaction utility" and "OR interaction utility" between input unit combinations.

[0154] In sub-step (d), based on the "AND interaction utility" and "OR interaction utility" in step (3), the output v(x T ) of the black-box model under any input unit combination T is interpreted as a combination of the "AND interaction utility" and "OR interaction utility".

[0155]

[0156] In sub-step (e), based on the combination of the "AND interaction utility" and "OR interaction utility" of each model obtained in sub-step (d), in order to further obtain sparse interactions, the present invention hopes to minimize the sum of the L1-norms of all interactions, that is

[0157]

[0158] where ‖·‖1 represents the L1-norm of a vector.

[0159] Optionally, when optimizing the above expression, it can be defined such that then the above optimization problem can be transformed into the following form.

[0160]

[0161] s.t. |q S | < τ S

[0162] where τS is a preset threshold value. In this embodiment,

[0163] After that, it enters step 104, and this step can be further divided into the following sub-steps:

[0164] (a) Divide all input units into "relevant input units", "irrelevant input units" and "mutually exclusive units";

[0165] (b) Divide the "interaction utility" of all combinations of input units in step 103 into "reliable components of interaction utility" and "unreliable components of interaction utility".

[0166] In the above sub-step (a), according to whether each input unit is relevant to the output of the black-box artificial intelligence model in human cognition, all input units are divided into "relevant input units", "irrelevant input units" and "mutually exclusive units" by using an algorithm or user-defined.

[0167] According to whether each input unit is relevant to the output of the black-box artificial intelligence model in human cognition, the user can customize / empirically divide the set N of all input units into the set R of "relevant input units", the set T of "irrelevant input units" and the set M of "mutually exclusive units". Among them, "relevant input units" refer to input units that are strongly relevant to the model output in human cognition or input units that have a direct causal relationship with the model output, "irrelevant input units" refer to input units that are weakly relevant to the model output in human cognition or input units that have no direct causal relationship with the model output, and "mutually exclusive units" refer to input units that should not affect the model output in human cognitive logic.

[0168] In the above sub-step (b), the "interaction utility" of all combinations of input units in step 103 is divided into "reliable components of interaction utility" and "unreliable components of interaction utility". Among them, the "reliable components of interaction utility" are the interaction utilities consistent with human cognition, which can be expressed as the interaction utilities of combinations that include "relevant input units" and do not include "mutually exclusive units"; the "unreliable components of interaction utility" are the interaction utilities inconsistent or in conflict with human cognition, which can be expressed as the interaction utilities of combinations that include "irrelevant input units" and "mutually exclusive units".

[0169] In a preferred example, the "interaction utility" I and (S|v and ,x) can be further divided into "reliable components of interaction utility" and "unreliable components of interaction utility" which can be calculated by the following formula:

[0170]

[0171] Among them, the set N of all input units is divided into the set R of "relevant input units", the set T of "irrelevant input units", and the set M of "mutually exclusive units".

[0172] In a preferred example, the "OR interaction utility" I or (S|v or ,x) can be further divided into the "reliable component of the OR interaction utility" and the "unreliable component of the OR interaction utility" Preferably, the division of an "OR interaction utility" is similar to the division of an "AND interaction utility" and can be calculated by the following formula:

[0173]

[0174] Among them, the set N of all input units is divided into the set R of "relevant input units", the set T of "irrelevant input units", and the set M of "mutually exclusive units". Another division of the "OR interaction utility" considers the ratio of the number |S∩R| of "relevant input units" in the combination S to the number |S| of input units in the combination S and can be calculated by the following formula:

[0175]

[0176] After that, step 105 is entered, and this step can be further divided into the following sub-steps:

[0177] (a) Calculate the evaluation index based on interaction, and count the ratio of the "reliable component of the interaction utility" and the "unreliable component of the interaction utility";

[0178] (b) Evaluate the alignment degree between the fine decision logic of the model and human cognition through the statistics of the significant interaction utilities of each order, the reliable component of the interaction utility, and the unreliable component of the interaction utility.

[0179] In the above sub-step (a), calculate the following evaluation index based on the interaction utility, and count the ratio of the "reliable component of the interaction utility" and the "unreliable component of the interaction utility". As Figure 3a shown, the user can quantify the alignment degree between the fine decision logic of the model and human cognition by the ratio of the "reliable component of the interaction utility" in the interaction utility.

[0180] The significant interaction is defined as follows. Given a preset threshold τ, the significant interaction is defined as the set Ω of interactions greater than the threshold τ and and Ω or , and can be calculated by the following formula:

[0181]

[0182]

[0183] Among them, the threshold τ is set to the numerical value of the 100th absolute value of the interaction utility after sorting all the absolute values of the interaction utilities from large to small.

[0184] The sum of the intensities of all significant interaction utilities is:

[0185]

[0186] The sum of the intensities of all interaction utilities is:

[0187]

[0188] The order of the interaction is defined as the number of input variables in the set S, that is, order(S) = |S|. For each order of interaction o, the sum of the reliable components of the interaction utility is:

[0189]

[0190] For each order of interaction o, the intensity of the reliable component of the interaction utility is:

[0191]

[0192] The total intensity of the reliable components of the interaction utility is:

[0193]

[0194] For each order of interaction o, the sum of the unreliable components of the interaction utility is:

[0195]

[0196] For each order of interaction o, the intensity of the unreliable component of the interaction utility is:

[0197]

[0198] The total intensity of the unreliable components of the interaction utility is:

[0199]

[0200] The following index represents the proportion of the intensity of the "reliable component of the interaction utility" in the intensity of the significant interaction utilities.

[0201]

[0202] Among them, the proportion of the intensity of the "reliable component of the interaction utility" in the intensity of the interaction utilities of the significant interactions The higher it is, the higher the alignment degree between the fine-grained decision-making logic of the model and human cognition.

[0203] The following index represents the proportion of the intensity of the "reliable component of interaction utility" in the intensity of interaction utility of all interactions.

[0204]

[0205] Among them, the proportion of the intensity of the "reliable component of interaction utility" in the intensity of interaction utility of all interactions The higher it is, the higher the alignment degree between the fine-grained decision-making logic of the model and human cognition.

[0206] The following index and respectively represent the proportion of the "reliable component of interaction utility" in the sum of positive utility and the sum of negative utility of significant interactions of each order.

[0207] or

[0208] or

[0209] Among them, ∈ represents a very small positive real number to prevent the division by zero operation of fractions. For each order of interaction o, calculate the sum of positive interaction utilities of significant interactions and the sum of negative interaction utilities of significant interactions Calculate the positive interaction utility of the "reliable component of interaction utility" and the negative interaction utility of the "reliable component of interaction utility" Index and The higher they are, that is, for each order of interaction o, the higher the proportion of the "reliable component of interaction utility" in the positive and negative utilities of significant interactions, the higher the alignment degree between the fine-grained decision-making logic of the model and human cognition.

[0210] The following index represents the proportion of positive and negative cancellation in the significant interaction utilities of each order.

[0211]

[0212] Among them, the index The higher it is, the higher the proportion of the effective interaction remaining after the positive and negative cancellation of significant interaction utilities for each order of interaction o. This index is helpful for understanding whether low-order interactions or high-order interactions play a major role in the model output. A well-trained model often tends to model low-order interactions.

[0213] The following index O reliableRepresents the weighted average order of "reliable interactions" in significant interactions.

[0214] or

[0215]

[0216] where the weighted average order O of "reliable interactions" reliable The lower it is, the more easily the influence of the model's fine decision-making logic is affected by low-order interactions; conversely, the higher the index O reliable is, the more easily the influence of the model's fine decision-making logic is affected by high-order interactions.

[0217] Therefore, as Figure 3a shown, evaluation metrics based on interaction utility can be used to statistically analyze the proportions of "reliable components of interaction utility" and "unreliable components of interaction utility", characterize the differences between the model's decision-making logic and human cognition, and reflect the potential representational deficiencies of the model.

[0218] In the above sub-step (b), as Figure 3b shown, statistical graphs of the significant interaction utilities, reliable components of interaction utility, and unreliable components of interaction utility at each order are respectively displayed, representing the differences between the model's decision-making logic and human cognition and reflecting the potential representational deficiencies of the model.

[0219] For each order of interaction o, the sum of the positive interaction utilities of significant interactions Salient + (o), the sum of the negative interaction utilities of significant interactions Salient - (o), the sum of the positive utilities of reliable interaction components Reliable + (o), the sum of the negative utilities of reliable interaction components Reliable - (o), the sum of the positive utilities of unreliable interaction components Unreliable + (o): = Salient + (o) - Reliable + (o), the sum of the negative utilities of unreliable interaction components Unreliable - (o): = Salient - (o) - Reliable - (o).

[0220] For each order of interaction o, when the indices and are higher, that is, the proportion of the reliable components of interaction utility in significant interactions is higher, it indicates a higher degree of alignment between the model's fine decision-making logic and human cognition.

[0221] Therefore, as Figure 3bAs shown, statistical charts of various-order significant interaction effects, reliable components of interaction effects, and unreliable components of interaction effects can be used to depict the differences between the decision-making logic of the model and human cognition, and to reflect the potential representational defects of the model.

[0222] Note that when calculating the indicators mentioned in all the above preferred examples, a small positive real number ∈ can be added to the denominator to prevent the situation of "dividing by zero".

[0223] It should be noted that those skilled in the art should understand that the implementation functions of the various modules shown in the above implementation manners of the system for evaluating the alignment degree between the fine decision-making logic of the artificial intelligence black-box model and human cognition can be understood with reference to the relevant descriptions of the method for evaluating the alignment degree between the fine decision-making logic of the artificial intelligence black-box model and human cognition. The functions of the various modules shown in the above implementation manners of the system for evaluating the alignment degree between the fine decision-making logic of the artificial intelligence black-box model and human cognition can be implemented by a program (executable instructions) running on a processor, or can also be implemented by specific logic circuits. If the system for evaluating the alignment degree between the fine decision-making logic of the artificial intelligence black-box model and human cognition in the embodiments of the present invention is implemented in the form of software function modules and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solutions of the embodiments of the present invention, in essence, or the parts that contribute to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods of the various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read Only Memory), magnetic disks, or optical discs, etc., which can store program codes. In this way, the embodiments of the present invention are not limited to any specific combination of hardware and software.

[0224] All the documents mentioned in the present invention are considered to be integrally included in the disclosure content of the present invention so that they can be used as a basis for modification when necessary. In addition, it should be understood that the above is only a preferred embodiment of this specification and is not used to limit the protection scope of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of one or more embodiments of this specification should be included in the protection scope of one or more embodiments of this specification.

Claims

1. A method for evaluating the degree of alignment between the fine decision logic of an artificial intelligence black box model and human cognition, characterized in that: The method comprises the following steps: (1) Provide a black box artificial intelligence model; Providing a black box artificial intelligence model to be analyzed, wherein the black box artificial intelligence model includes an artificial intelligence model pre-trained based on a certain data set; (2) Select input samples and perform recognition; Select an input sample for interactive utility calculation, identify the input sample, decompose the input sample into n input units, and combine the n input units to obtain 2 n The input sample is selected from the following group: table data, picture, text, voice or a combination thereof; (3) Calculate and combine "interaction utility"; In step (2), n The combination of input units is respectively input into the black box artificial intelligence model, and the output of the black box artificial intelligence model is obtained; based on the output of the black box artificial intelligence model, the interaction utility between the input units is modeled, so as to obtain the "interaction utility" of the black box artificial intelligence model for each combination of the input units; and the output of the black box artificial intelligence model on a certain input unit combination is interpreted as the combination of its "interaction utility" on the input unit combination; the "interaction utility" is used to represent the contribution value of the interaction between any input unit combination to the corresponding output of each model; (4) Roughly divide the "interaction utility"; According to whether each input unit is relevant to the output of the black box artificial intelligence model in human cognition, all input units are divided into "relevant input units" and "non-relevant input units"; (5) refining the relevant input units and performing further calculations; The relevant input units are divided twice, and the non-relevant input units are further divided into "irrelevant input units" and "mutually exclusive units", and based on the "relevant input units", "irrelevant input units" and "mutually exclusive units", the "interaction utility" of the combination of all input units in step (3) is divided into "reliable component of interaction utility" and "unreliable component of interaction utility"; the reliable component of interaction utility represents interaction utility consistent with human cognition, and the unreliable component of interaction utility represents interaction utility inconsistent with human cognition; the "irrelevant input unit" represents an input unit that is weakly correlated with the model output in human cognition, or an input unit that has no direct causal relationship; the "mutually exclusive unit" represents an input unit that does not affect the model output in human cognitive logic, or an input unit that is mutually exclusive with the model output in cognitive logic; (6) Use interaction-based evaluation indicators to evaluate the degree of alignment between the model’s decision logic and human cognition; According to the preset interaction-based evaluation indicators, the ratio of "reliable component of interaction utility" to "unreliable component of interaction utility" is calculated to evaluate the degree of alignment between the model's refined decision-making logic and human cognition; The order of step (1) and step (2) can be replaced at will or performed simultaneously.

2. The method according to claim 1, characterized in that The interaction utility is divided into "AND interaction utility" and "OR interaction utility", which respectively represent the "AND relationship" between the input units modeled by the model and the "OR relationship" between the input units modeled by it.

3. The method according to claim 2, characterized in that The reliable component of the "and interaction utility" is calculated as the interaction utility of the combination that includes "related input units" and does not include "mutually exclusive units"; the unreliable component of the "and interaction utility" is calculated as the interaction utility of the combination that does not include "related input units" or includes "mutually exclusive units".

4. The method according to claim 3, characterized in that The reliable component of the "interaction utility" and the unreliable component of "interaction utility" The specific calculation is obtained by the following formula: The set N of all input units is divided into a set R of "related input units", a set T of "unrelated input units" and a set M of "mutually exclusive units"; "S" is the combination of input units, "x" is the input sample, and "I and " is the "and interaction utility" corresponding to the combination S of input units of the artificial intelligence model; "v and " is the output component determined by "and interaction utility".

5. The method according to claim 2, characterized in that: The step (4) also The method comprises the following steps: evenly distributing the "or interaction utility" to each input unit included in the interaction, thereby obtaining the "or interaction utility reliable component", wherein the "or interaction utility reliable component" is calculated as the interaction utility component evenly distributed to the "related input unit" in the "or interaction utility"; and the "or interaction utility unreliable component" is calculated as the interaction utility component evenly distributed to the "irrelevant input unit" and the "mutually exclusive unit" in the "or interaction utility", which is obtained by the following formula: Where |S∩R| is the number of "related input units" in the combination S considered for the partition of "or interaction utility", |S| is the number of input units in the combination S; "S" is the combination of input units, "x" is the input sample, "I or " is the "or interaction utility" corresponding to the combination S of input units of the artificial intelligence model; "v or " is the output component determined by "or interaction utility".

6. The method according to claim 5, characterized in that The reliable component of the or interaction utility is calculated by the following formula: Among them, the set N of all input units is divided into a set R of "related input units", a set T of "unrelated input units" and a set M of "mutually exclusive units".

7. The method according to claim 6, characterized in that The evaluation index includes: calculating the strength of the reliable component and the unreliable component of the interaction utility or their proportion in the interaction utility and comparing them as indicators; visualizing the statistical graphs of significant interaction utilities of each order, the reliable component of the interaction utility, and the unreliable component of the interaction utility and using them as indicators; calculating the proportion of the reliable component of the interaction utility in significant interactions as an indicator; the significant interaction is defined as the interaction set Ω greater than a threshold τ and and Ω or , which is calculated by the following formula: The threshold τ is set to the value of the absolute value of the Xth interaction utility after the absolute values ​​of all interaction utilities are sorted from large to small, where X is set according to user experience; Ω and is the set of significant interactions among the interactions, Ω or is the set of significant interactions in or among interactions; and " is the "and interaction utility" corresponding to the combination S of input units of the artificial intelligence model; "v and " is the output component determined by "and interaction utility".

8. The method according to claim 7, characterized in that The proportion of the reliable component of the interaction utility in the significant interaction Calculate using the following formula: