Algorithm for evaluating degree of alignment between fine-grained decision-making logic of artificial intelligence black-box model and human cognition

By analyzing the interaction utility of the black-box model and dividing it into reliable and unreliable components, the degree of alignment between the model and human cognition is evaluated, which solves the problem of insufficient trust in existing technologies and improves the trustworthiness of the model in high-risk scenarios.

WO2026066024A1PCT designated stage Publication Date: 2026-04-02SHANGHAI JIAOTONG UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing technologies struggle to reliably assess the alignment between the sophisticated decision-making logic of black-box models and human cognition, resulting in insufficient trust when applying black-box models in high-risk scenarios.

Method used

By analyzing the interaction distribution of the black-box artificial intelligence model, input samples are selected and decomposed into input units. The interaction utility is calculated and divided into reliable and unreliable components. Interaction-based evaluation metrics are used to assess the alignment between the model's decision-making logic and human cognition.

Benefits of technology

It enables reliable evaluation of the fine-grained decision-making logic of black-box models and human cognition, reflects potential representational defects of the models, and improves trust in high-risk scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025088584_02042026_PF_FP_ABST
    Figure CN2025088584_02042026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of machine learning, and discloses a method and system for evaluating the degree of alignment between a fine-grained decision-making logic of an artificial intelligence black-box model and human cognition. The method and system can evaluate the degree of alignment between a decision-making logic of a model and human cognition by analyzing the distribution of interactions modeled by the model. The implementation of the method and system comprises the following steps: providing an input sample; using a black-box model to predict the input sample to obtain a prediction result of the model; modeling the interactions between input units of the sample on the basis of the output of the black-box model, calculating the interaction strength of combinations formed between the input units, and expressing the output of the black-box model as an "interaction effect" between the combinations of the input units; and using interaction-based evaluation metrics to evaluate the degree of alignment between a decision-making logic of the model and human cognition. The advantage of the present invention is that: a quantification method for evaluating the degree of alignment between a fine-grained decision-making logic of an artificial intelligence black-box model and human cognition is provided. Compared with previous research, the present invention can better characterize the difference between a decision-making logic of a model and human cognition, thereby reflecting potential representation defects of the model.
Need to check novelty before this filing date? Find Prior Art

Description

An algorithm for evaluating the alignment degree of fine decision logic of an artificial intelligence black box model and human cognition TECHNICAL FIELD

[0001] The present application relates to the technical field of machine learning, in particular to a method and system for evaluating the alignment degree of fine decision logic of an artificial intelligence black box model and human cognition. BACKGROUND

[0002] At present, deep learning has been widely used in various fields and has shown strong performance, but it is difficult for users and researchers to analyze the black box nature of artificial intelligence models. In the existing technology, the interactive explanation method can analyze an artificial intelligence model to obtain sparse and simple explanations. However, the obtained explanations still cannot directly evaluate the difference between the decision logic of the black box model and human cognition, reflecting the potential representation defects of the model. Evaluating the alignment degree of the decision logic of the model and human cognition is a necessary condition for human to trust the black box model and apply the black box model in high-risk scenarios.

[0003] Therefore, how to reliably evaluate the alignment degree of fine decision logic of a black box model and human cognition is a problem to be solved in the field of explainability. SUMMARY

[0004] The purpose of the present application is to provide a method and system for evaluating the alignment degree of fine decision logic of an artificial intelligence black box model and human cognition, which can evaluate the alignment degree of the decision logic of the model and human cognition by analyzing the interaction distribution modeled by the model.

[0005] The present application discloses a method for evaluating the alignment degree of fine decision logic of an artificial intelligence black box model and human cognition, characterized in that the method comprises the following steps:

[0006] (1) selecting a black box artificial intelligence model;

[0007] Select a black box artificial intelligence model to be analyzed, wherein the black box artificial intelligence model comprises a pre-trained artificial intelligence model based on a certain data set;

[0008] (2) selecting an input sample and performing identification;

[0009] Select an input sample for interactive utility calculation, and identify the input sample, so as to decompose the input sample into n input units, and combine the n input units to obtain 2 n input unit combinations; the input sample is selected from the group consisting of table data, pictures, text, speech or combinations thereof;

[0010] (3) calculating the "interactive utility" and combining;

[0011] 2 in step (2) n The black-box AI model is input into a combination of input units, and the output of the black-box AI model is obtained. Based on the output of the black-box AI model, the interaction between the input units is modeled to obtain the "interaction utility" of the black-box AI model for each combination of input units. The output of the black-box AI model on a certain combination of input units is interpreted as a combination of its "interaction utility" on that combination of input units.

[0012] (4) Divide “interaction utility” into “reliable components of interaction utility” and “unreliable components of interaction utility”;

[0013] Based on whether each input unit is related to the output of the black-box artificial intelligence model in human cognition, all input units are divided into "relevant input units", "irrelevant input units" and "mutually exclusive units" using algorithms or user-defined methods.

[0014] The "interaction utility" of all combinations of input units in step (3) is divided into "reliable components of interaction utility" and "unreliable components of interaction utility". The "reliable components of interaction utility" represent interaction utility consistent with human cognition, while the "unreliable components of interaction utility" represent interaction utility inconsistent with human cognition. Interaction utility can be divided into "AND interaction utility" and "OR interaction utility", representing the "AND relationship" between the input units modeled by the model and the "OR relationship" between the input units modeled by the model, respectively. Preferably, the "reliable components of AND interaction utility" can be calculated as the interaction utility of combinations containing "related input units" but not "mutually exclusive units"; the "unreliable components of AND interaction utility" can be calculated as the interaction utility of combinations not containing "related input units" or containing "mutually exclusive units". Preferably, a "reliable component of OR interaction utility" can be calculated as the interaction utility of combinations containing "related input units" but not "mutually exclusive units"; the "unreliable components of OR interaction utility" can be calculated as the interaction utility of combinations not containing "related input units" or containing "mutually exclusive units". Another algorithm is to distribute the "OR interaction utility" equally among the input units included in the interaction, so that the "reliable component of OR interaction utility" can be calculated as the interaction utility component evenly distributed among the "relevant input units" in this "OR interaction utility"; the "unreliable component of OR interaction utility" can be calculated as the interaction utility component evenly distributed among the "unrelated input units" and "mutually exclusive units" in this "OR interaction utility".

[0015] (5) Use interaction-based evaluation metrics to evaluate the alignment between the model's decision-making logic and human cognition;

[0016] According to the proposed interaction-based evaluation index, the proportion of "reliable component of interaction utility" and "unreliable component of interaction utility" is counted to evaluate the alignment degree of the model's fine-grained decision logic and human cognition.

[0017] Wherein, the order of step (1) and step (2) can be replaced or performed at the same time.

[0018] In a preferred embodiment, the output of the black-box artificial intelligence model on the combination of the input units is defined as follows: if the input units in the set N\S in the input sample x are shielded and the input units in the set S are not shielded, a sample x S is obtained, and the output obtained by inputting x S into the model is represented as v(x S ); if the input sample is inputted into the black-box model without shielding, the output obtained is represented as v(x N ); if all the input units in the input sample are completely shielded and then inputted into the black-box model, the output obtained is represented as v(x Preferably, in natural language processing applications, the shielding operation can be realized by replacing the embedding vectors of each token covered on each input unit contained in S with a certain specific reference value vector.

[0019] In a preferred embodiment, if the black-box artificial intelligence model is a classification model, the output of the model on a certain input unit set S v(x S ) can be represented as the logit value corresponding to the real label dimension, that is, If the black-box artificial intelligence model is a language generation model, the model output can be represented as the logit value of the dimension corresponding to the next token with the maximum probability generated by the model, that is, where y max represents the dimension corresponding to the next token with the maximum probability generated by the model; the model output v(x S ) can also be represented as the sum of the logits values of the input dimensions corresponding to each token with the maximum probability (or the target token determined in advance) in the subsequent T tokens generated. Here, the subsequent T target tokens generated are represented as y1, y2, …, y T , then v(x S ) can be calculated as

[0020] Similarly, v(x S ) can also be calculated as

[0021] In a preferred embodiment, the step (3) further comprises the step of: interpreting the output v(x S ) of the black-box artificial intelligence model on a certain input unit set S as a combination of the "interaction utilities" of the black-box artificial intelligence model on the input unit combination.

[0022] In a preferred embodiment, the "and interaction utility" means "the additional utility of a combination of input units for a black-box model output when and only when none of the units in the combination is occluded";

[0023] In a preferred embodiment, for the black-box artificial intelligence model, a combination of input units Definition I and (S|v and ,x) is the "and interaction utility" of the artificial intelligence model on the input unit combination S, which can be calculated by the following formula:

[0024] Here v and (x L ) represents the output component determined by the "and interaction utility" decomposed from v(x L ), corresponding to the formula v(x L ) = v and (x L ) + v or (x L ), and v or (x L ) represents the output component determined by the "or interaction utility" decomposed from v(x L ).

[0025] In a preferred embodiment, the "or interaction utility" means "the additional utility of a combination of input units for a set of black-box model outputs when and only when at least one of the units in the combination is not occluded";

[0026] In a preferred embodiment, for the black-box artificial intelligence model, a combination of input units Definition I or (S|v or ,x) is the "or interaction utility" of the artificial intelligence model on the input unit combination S, which can be calculated by the following formula.

[0027] Here v or (x L ) represents the output component determined by the "or interaction utility" decomposed from v(x L ).L ) = v and (x L )+v or (x L ), and v and (x L ) represents from v(x L The output component determined by "interaction utility" is decomposed from the product.

[0028] The output of a neural network can be represented as a sum of different interactions. In a preferred embodiment, the output v(x) of the black-box AI model on the combination of input units S can be expressed as... S ) can be interpreted as a combination of "AND interaction utility" and "OR interaction utility", i.e., v(x) S )=

[0029] v and (x S )+v or (x S The output v(x) of the black-box artificial intelligence model on the combination of input units S can also be expressed as... S ) can be interpreted as a combination of "interaction utility", i.e., v(x) S ) = v and (x S It should be understood that, without ambiguity, "interaction utility" in the following text can refer to all combinations of "with interaction utility" and "or interaction utility".

[0030] In a preferred embodiment, step (4) further includes the following steps: dividing all input units into “relevant input units”, “irrelevant input units”, and “mutually exclusive units”. The interaction utility of a combination that includes “relevant input units” but does not include “mutually exclusive units” is defined as a “reliable component of interaction utility”, and the interaction utility of a combination that includes either “irrelevant input units” or “mutually exclusive units” is defined as an “unreliable component of interaction utility”.

[0031] In a preferred embodiment, based on whether each input unit is cognitively relevant to the output of the black-box AI model, the set N of all input units is divided into a set R of "relevant input units", a set T of "irrelevant input units", and a set M of "mutually exclusive units" using an algorithm or user-defined method. wherein, “relevant input units” means input units that are strongly related to the model output in human cognition, or input units that have direct causal relationship; “irrelevant input units” means input units that are weakly related to the model output in human cognition, or input units that have no direct causal relationship; “mutually exclusive units” means input units that should not affect the model output in human cognitive logic, or input units that are logically repulsive to the model output in human cognitive logic. Certain input samples can only contain one or more of “relevant input units”, “irrelevant input units” and “mutually exclusive units”. Preferably, for text input, the input units can be token word embedding vectors, or all word embedding vectors corresponding to words, phrases, sentences, paragraphs; for image input, the input units can be pixels, image regions.

[0032] [According to Rule 26, corrected on 30.04.2025] For example, given a legal case involving multiple people “On a certain day, the defendant Zhao 2 lied to the victim Zhao 1 that he could help Zhao 1 eliminate his record of theft and record of assault. After gaining Zhao 1’s trust, he cheated Zhao 1 out of money.”, and evaluate the fine decision-making logic of the legal large model on the judgment result of Zhao 2 “crime of intentional injury”. In this legal case, the set of all input units N = {chat, lie, eliminate, theft, cumulative record, cheat, money}. Among them, the set of “relevant input units” R = {"lie", "cheat", "money"} is the input unit that has a direct causal relationship with the judgment result of Zhao 2 “fraud”. The set of “irrelevant input units” T = {"eliminate"} is the input unit that has no direct causal relationship with the judgment result of Zhao 2 “fraud”. The set of “mutually exclusive units” M = {"cumulative record", "theft"} is the content that Zhao 2 cheats Zhao 1, not the behavior of Zhao 2, which is logically mutually exclusive with the judgment result of Zhao 2 “crime of intentional injury” in human cognition.

[0033] In a preferred embodiment, in the task of evaluating the black-box artificial intelligence model, the user can label the input units in the vocabulary as "relevant input units", "irrelevant input units" and "mutually exclusive units" from the perspective of human cognition; the user can also automatically determine or label the set of "relevant input units", "irrelevant input units" and "mutually exclusive units" using an algorithm; the user can measure the importance of all input units to the model output using an attribution algorithm including the Shapley value, and automatically divide all input units into "relevant input units", "irrelevant input units" and "mutually exclusive units" according to the importance score; the user can also measure the semantic logical level of correlation between all input units and the model output using a complex artificial intelligence algorithm, and automatically divide all input units into "relevant input units", "irrelevant input units" and "mutually exclusive units" according to the correlation score.

[0034] In a preferred embodiment, the "interaction utility" of the combination of all input units in step (4) is further divided into "reliable components of interaction utility" and "unreliable components of interaction utility". Among them, the "reliable components of interaction utility" are the interaction utilities consistent with human cognition, which can be defined as the interaction utilities of combinations containing "relevant input units" and not containing "mutually exclusive units"; the "unreliable components of interaction utility" are the interaction utilities inconsistent with human cognition, which can be defined as the interaction utilities of combinations containing "irrelevant input units" or "mutually exclusive units".

[0035] In a preferred embodiment, the interaction utility can be divided into "and interaction utility" and "or interaction utility", representing the "and relationship" between the input units modeled by the model and the "or relationship" between the input units modeled by the model, respectively. The "and interaction utility" I and (S|v and ,x) can be further divided into "reliable components of and interaction utility" and "unreliable components of and interaction utility" which can be calculated by the following formula:

[0036] Among them, the set of all input units N is divided into the set of "relevant input units" R, the set of "irrelevant input units" T and the set of "mutually exclusive units" M.

[0037] In a preferred embodiment, the interaction utility can be divided into "and interaction utility" and "or interaction utility", representing the "and relationship" between the input units modeled by the model and the "or relationship" between the input units modeled by the model, respectively. The "or interaction utility" I or (S|v or ,x) can be further divided into "reliable components of or interaction utility" and "unreliable component of or interaction utility" Preferably, the classification of "or interaction utility" is similar to the classification of "and interaction utility", which can be calculated by the following formula:

[0038] Wherein, the set N of all input units is divided into the set R of "relevant input units", the set T of "irrelevant input units" and the set M of "mutually exclusive units".

[0039] Another classification of "or interaction utility" considers the proportion of the number of "relevant input units" in the combination S |S∩R| to the number of input units in the combination S |S|. It can be calculated by the following formula:

[0040] In a preferred embodiment, the step (5) further comprises the following steps: using the proposed interaction utility-based evaluation index, the proportion of "reliable component of interaction utility" and "unreliable component of interaction utility" is calculated to evaluate the degree of alignment between the fine-grained decision logic of the model and human cognition. Users can quantify the degree of alignment between the fine-grained decision logic of the model and human cognition with the proportion of "reliable component of interaction utility" in the interaction utility.

[0041] In a preferred embodiment, the significant interaction is defined as follows. Given a preset threshold τ, the significant interaction is defined as the interaction set Ω and and Ω or , which can be calculated by the following formula:

[0042] Wherein, the threshold τ can be set as a percentage of the output of the black-box artificial intelligence model. Preferably, the threshold τ can also be set as the numerical size of the kth interaction utility absolute value after sorting all interaction utility absolute values from large to small. Preferably, in an input sample containing 10 input units, there are 2 10 "and interaction utilities"; or containing 2 10 "and interaction utilities" and 2 10 "or interaction utilities", the threshold τ distinguishes a certain proportion of "significant and interaction" and "significant or interaction". The threshold τ can also be set according to user experience, such as setting it as the 100th largest interaction utility absolute value among all interaction utility absolute values.

[0043] In a preferred embodiment, the sum of the strengths of all significant interaction utilities is:

[0044] In a preferred embodiment, the sum of the strengths of all interaction utilities is:

[0045] In a preferred embodiment, the order of an interaction is defined as the number of input variables in the set S, i.e., order(S) = |S|. For each order of interaction o, the sum of the reliable components of the interaction utility is:

[0046] In a preferred embodiment, for each order of interaction o, the strength of the reliable components of the interaction utility is:

[0047] In a preferred embodiment, the total strength of the reliable components of the interaction utility is:

[0048] In a preferred embodiment, for each order of interaction o, the sum of the unreliable components of the interaction utility is:

[0049] In a preferred embodiment, for each order of interaction o, the strength of the unreliable components of the interaction utility is:

[0050] In a preferred embodiment, the total strength of the unreliable components of the interaction utility is:

[0051] In a preferred embodiment, the following index represents the proportion of the strength of the reliable components of the interaction utility in the strength of the significant interaction utilities.

[0052] where the proportion of the strength of the reliable components of the interaction utility in the strength of the significant interactions is higher, the better the alignment of the model’s refined decision logic with human cognition. A model that is better aligned with human cognition should have a higher value of the index .

[0053] In a preferred embodiment, the following index represents the proportion of the strength of the reliable components of the interaction utility in the strength of all interactions.

[0054] where the proportion of the strength of the reliable components of the interaction utility in the strength of all interactions is higher, the better the alignment of the model’s refined decision logic with human cognition. A model that is better aligned with human cognition should have a higher value of the index .

[0055] In a preferred embodiment, the following indices and represent the proportion of the reliable components of the interaction utility in the sum of the positive and negative utilities, respectively, of the significant interactions of each order.

[0056] or

[0057] or

[0058] where ∈ denotes a very small positive real number to prevent division by zero operation. For each order of interaction o, compute the sum of positive interaction utilities of significant interactions and the sum of negative interaction utilities of significant interactions the positive interaction utility of the reliable component of interaction utility and the negative interaction utility of the reliable component of interaction utility indicator and The higher, i.e., the higher the proportion of the reliable component of interaction utility in the sum of positive and negative utilities of significant interactions for each order of interaction o, the better the alignment of the model’s fine-grained decision logic with human cognition. A model with a higher alignment with human cognition has a higher or indicator should be higher.

[0059] In a preferred embodiment, the following indicator denotes the proportion of positive and negative cancellation in the significant interaction utilities.

[0060] where the indicator is higher, the higher the proportion of effective interaction remaining after positive and negative cancellation of significant interaction utilities for each order of interaction o. This indicator is helpful to understand whether low-order or high-order interactions play a major role in the model’s output. A well-trained model tends to model more low-order interactions.

[0061] In a preferred embodiment, the following indicator O reliable denotes the weighted average order of “reliable interactions” in significant interactions.

[0062] or

[0063] where the weighted average order of “reliable interactions” in significant interactions O reliable is lower, the more susceptible the influence of the model’s fine-grained decision logic to low-order interactions; conversely, the higher the indicator O reliable , the more susceptible the influence of the model’s fine-grained decision logic to high-order interactions. A model with a higher alignment with human cognition has a lower weighted average order of “reliable interactions” in significant interactions Oreliable The lower the ∈ [0, |N|] is, the more the model tends to model lower-order "reliable interactions".

[0064] In a preferred embodiment, the differences between the model decision logic and human cognition can be represented by visualizing the statistical plots of the salient interaction utilities of each order, the reliable components of the interaction utilities, and the unreliable components of the interaction utilities, reflecting the potential representation defects of the model. Specifically, for each order of interaction o, the sum of the positive interaction utilities of the salient interactions Salient + (o), the sum of the negative interaction utilities of the salient interactions Salient - (o), the sum of the positive interaction utilities of the reliable interaction components Reliable + (o), the sum of the negative interaction utilities of the reliable interaction components Reliable - (o), the sum of the positive interaction utilities of the unreliable interaction components Unreliable the sum of the negative interaction utilities of the unreliable interaction components Unreliable - (o) := Salient - (o) - Reliable - (o).

[0065] For each order of interaction o, the higher the indicators and are, i.e., the higher the proportion of the reliable components of the interaction utilities in the salient interactions is, the higher the degree of alignment between the fine-grained decision logic of the model and human cognition is.

[0066] In summary, the statistical plots of the salient interaction utilities of each order, the reliable components of the interaction utilities, and the unreliable components of the interaction utilities, as well as the indicators and can be used to depict the differences between the model decision logic and human cognition, reflecting the potential representation defects of the model.

[0067] A second aspect of the present application discloses a system for evaluating the degree of alignment between the fine-grained decision logic of an artificial intelligence black box model and human cognition, comprising:

[0068] (1) an input module configured to accept a pre-trained black box artificial intelligence model and a set of data to be analyzed.

[0069] (2) a calculation module configured to calculate the "interaction utilities" between the input units of the data to be analyzed modeled by the model based on the model and the data in the input module. The "interaction utilities" are divided into "reliable components of the interaction utilities" and "unreliable components of the interaction utilities". The proportions of the "reliable components of the interaction utilities" and the "unreliable components of the interaction utilities" are calculated using the proposed interaction-based evaluation indicators.

[0070] (3) an output module configured to evaluate the alignment degree of the fine decision logic of the black-box model and human cognition based on the interaction utility evaluation index and the statistics of the significant interaction utility of each order, the reliable component of the interaction utility, the unreliable component of the interaction utility, and the alignment degree of the fine decision logic of the model and human cognition.

[0071] The present application has the advantages that:

[0072] 1) The present application classifies the input components reasonably, introduces the definition and calculation of the reliable component, thereby reliably evaluating the alignment degree of the fine decision logic of the black-box model and human cognition, and reflecting the potential representation defects of the model.

[0073] 2) The present application introduces multiple standards for evaluating the alignment degree of the fine decision logic of the black-box model and human cognition, thereby ensuring that the method used in the present application is applicable to most black-box models.

[0074] A large number of technical features are described in the specification of the present application, which are distributed in various technical solutions. If all possible combinations of technical features (i.e. technical solutions) of the present application are listed, the specification will be too long. In order to avoid this problem, each technical feature disclosed in the above invention content, each technical feature disclosed in the following embodiments and examples, and each technical feature disclosed in the drawings can be freely combined to form various new technical solutions (these technical solutions are considered to have been described in the specification), unless such combination of technical features is technically infeasible. For example, features A+B+C are disclosed in one example, features A+B+D+E are disclosed in another example, features C and D are equivalent technical means that play the same role, and can only be used at a time, and feature E can be combined with feature C technically. Therefore, the scheme of A+B+C+D should not be considered to have been described because it is technically infeasible, and the scheme of A+B+C+E should be considered to have been described. BRIEF DESCRIPTION OF DRAWINGS

[0075] FIG. 1 is a flowchart of the evaluation of the alignment degree of the fine decision logic of an artificial intelligence black-box model and human cognition according to the first embodiment of the present application;

[0076] FIG. 2a is a schematic diagram of a sample to be analyzed and input variables according to the present application;

[0077] FIG. 2b is a schematic diagram of the division of all input units of the sample to be analyzed into “relevant input units”, “irrelevant input units”, and “mutually exclusive units” according to the present application;

[0078] Fig. 3 is a schematic diagram showing the absolute values of the "and interaction utility" and "or interaction utility" in descending order according to the present application;

[0079] Fig. 4(a) is an evaluation index based on interaction utility according to the present application;

[0080] Fig. 4(b) is a statistical diagram of each order significant interaction utility, reliable component of interaction utility, and unreliable component of interaction utility. DETAILED DESCRIPTION

[0081] The present inventors have developed, for the first time, a method and system for evaluating the alignment degree of fine decision logic of an artificial intelligence black box model with human cognition.

[0082] It should be noted that the relationship terms such as first and second, etc. in the present patent document are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or device. Without more limitations, the element defined by the statement "including one" does not exclude the presence of another identical element in the process, method, article or device including the element. In the present patent document, if it is mentioned that a certain action is performed according to a certain element, it means that the action is performed at least according to the element, including two cases: the action is performed only according to the element, and the action is performed according to the element and other elements. The expressions of multiple, multiple times, multiple kinds, etc. include 2, 2 times, 2 kinds and more than 2, more than 2 times, more than 2 kinds.

[0083] The present application comprises the following steps:

[0084] (1) Selecting a black box artificial intelligence model

[0085] A black box artificial intelligence model to be analyzed is selected, which includes an artificial intelligence model pre-trained based on a certain data set.

[0086] (2) Selecting input samples and performing identification

[0087] selecting an input sample for the calculation of interaction utility, and identifying the input sample, thereby decomposing the input sample into n input units and combining the n input units to obtain 2n combinations of the input units; the input sample is selected from the group consisting of table data, pictures, text, voice, or a combination thereof;

[0088] (3) calculating the "interaction utility" and combining;

[0089] inputting the 2n combinations of the input units in step (2) into the black-box artificial intelligence model respectively, and obtaining the output of the black-box artificial intelligence model; based on the output of the black-box artificial intelligence model, modeling the interaction between the input units, thereby obtaining the "interaction utility" of the black-box artificial intelligence model for each combination of the input units; and explaining the output of the black-box artificial intelligence model on a certain combination of input units as the combination of the "interaction utility" of the black-box artificial intelligence model on the combination of input units;

[0090] (4) dividing the "interaction utility" into "reliable components of interaction utility" and "unreliable components of interaction utility";

[0091] According to whether each input unit is related to the output of the black-box artificial intelligence model in human cognition, all input units are divided into "relevant input units", "irrelevant input units" and "mutually exclusive units" by algorithm or user-defined.

[0092] The "interaction utility" of the combination of all input units in step (3) is divided into a "reliable component of interaction utility" and an "unreliable component of interaction utility". The "reliable component of interaction utility" is the interaction utility consistent with human cognition, which can be expressed as the interaction utility of the combination containing "relevant input units" and not containing "mutually exclusive units"; the "unreliable component of interaction utility" is the interaction utility inconsistent with human cognition, which can be expressed as the interaction utility of the combination containing "irrelevant input units" and "mutually exclusive units". Preferably, the interaction utility can be divided into "and interaction utility" and "or interaction utility", which respectively represent the "and relationship" between the input units modeled by the model and the "or relationship" between the input units modeled by the model. Preferably, the "reliable component of and interaction utility" can be calculated as the interaction utility of the combination containing "relevant input units" and not containing "mutually exclusive units"; the "unreliable component of and interaction utility" can be calculated as the interaction utility of the combination not containing "relevant input units" or containing "mutually exclusive units". Preferably, a "reliable component of or interaction utility" can be calculated as the interaction utility of the combination containing "relevant input units" and not containing "mutually exclusive units"; the "unreliable component of or interaction utility" can be calculated as the interaction utility of the combination not containing "relevant input units" or containing "mutually exclusive units". Another algorithm is to equally distribute the "or interaction utility" to each input unit contained in this interaction, so that the "reliable component of or interaction utility" can be calculated as the interaction utility component of the "relevant input units" in this "or interaction utility"; the "unreliable component of or interaction utility" can be calculated as the interaction utility component of the "irrelevant input units" and "mutually exclusive units" in this "or interaction utility".

[0093] (5) Using the interaction-based evaluation index, the alignment degree of the model decision logic and human cognition is evaluated.

[0094] According to the proposed interaction-based evaluation index, the proportion of the "reliable component of interaction utility" and the "unreliable component of interaction utility" is counted, and the alignment degree of the fine decision logic of the model and human cognition is evaluated.

[0095] Embodiments

[0096] In order to make the purpose, technical scheme and advantages of the present application more clear, the embodiments of the present application will be further described in detail below with reference to the drawings.

[0097] The first embodiment of the present application relates to a method for evaluating the alignment degree of the fine decision logic of an artificial intelligence black box model and human cognition, and the flowchart of the method is shown in FIG. 1. The method comprises the following steps:

[0098] In step 101: based on a certain data set, a black-box artificial intelligence model is trained as a model to be analyzed. Optionally, a model can be a deep neural network. In an embodiment, the model is selected as a white magnolia legal large model fine-tuned with Chinese legal documents, which is a deep neural network.

[0099] Then, step 102 is entered, which can be further divided into the following two sub-steps:

[0100] (a) Collecting data to be analyzed. Optionally, if the model in step (1) can be used for natural language generation, the data to be analyzed can be a text.

[0101] (b) Preprocessing the data to be analyzed. Optionally, the data to be analyzed can be processed by word segmentation, encoding into embedding vectors, etc. to adapt to the input of the above-mentioned model.

[0102] In an embodiment, as shown in FIG. 2(a), the data is selected as any Chinese legal case description sentence, which contains 10 input variables, representing the 1st to 10th word groups, respectively, denoted as word group 1, word group 2, …, word group 10. In an embodiment, inputting a Chinese legal case description sentence into the white magnolia legal large model can obtain the corresponding crime output of the case.

[0103] Then, step 103 is entered, which can be further divided into the following five sub-steps:

[0104] (a) Calculate the output value of the black-box model under any occlusion of the above-mentioned data to be analyzed;

[0105] (b) Calculate the “and interaction utility” in the data to be analyzed modeled by the black-box model;

[0106] (c) Calculate the “or interaction utility” in the data to be analyzed modeled by the black-box model;

[0107] (d) Based on the “and interaction utility” and the “or interaction utility”, explain the output of the black-box model as a combination of the “and interaction utility” and the “or interaction utility” between the input unit combinations;

[0108] (e) Based on the combination of the “and interaction utility” and the “or interaction utility”, further optimize the combination of the “and interaction utility” and the “or interaction utility” of the model, so that the interaction is more sparse and concise.

[0109] In the above sub-step (a), the black-box artificial intelligence model is denoted as v, and the data sample x to be analyzed contains n input units, denoted as a set N = {1, 2, …, n}, where the number of input units is generally 5 to 20. Optionally, each token in the input data can be regarded as an input unit, or the input data can be divided into multiple word groups or sentences, and the word embedding vector of all tokens covered by each word group or sentence can be regarded as an input unit. For extremely long input samples, all tokens covered by multiple sentences or multiple paragraph chapters can be counted as one input unit, and different input units correspond to different paragraph chapters in the input sample.

[0110] Any subset of the input units is called a combination of input units, v(x S ) represents the output of the black-box model when the input units in the combination S are given and the input units in N\S are blocked. In particular, if the input sample is input into the model without blocking, the obtained output is denoted as v(x N ), and if the input sample is completely blocked after all input units are completely blocked, the output obtained is denoted as

[0111] When calculating the output v(x S ) of the black-box model on a combination of input units S, the input values of the input units contained in S need to be kept in the input sample to be analyzed, and all input units in the complement of S N\S are replaced by a reference value. In particular, for a natural language processing task, it is assumed that the reference value is b, and it is assumed that each input unit contains only one token, then each input unit of the blocked input sample x S is defined by the following formula:

[0112] Here, (x S ) i = x i represents whether the word embedding vector corresponding to the i-th input unit (i.e., the i-th token) of the blocked sample is kept unchanged or the word embedding vector of the i-th input unit (i.e., the i-th token) in the original input sample; (x S ) i = b i represents that the word embedding vector corresponding to the i-th input unit (i.e., the i-th token) of the blocked sample is replaced by the word embedding vector reference value.

[0113] In this way, v(x S ) can be calculated as v(x SThe output value of the input model. Optionally, the baseline value can be set as the mean value of the sample set, a random value, a zero value, etc., or can be obtained through learning. For natural language processing tasks, the baseline value can be set as a certain special word embedding vector, and the masking of a single token can be implemented by replacing the word embedding vector of this token with the special word embedding vector.

[0114] If the i-th input unit includes multiple tokens (for example, the i-th input unit is a word, a phrase, a short sentence, etc.), the masking of the i-th input unit needs to replace the word embedding vectors corresponding to all tokens contained in the i-th input unit. Further, the output v(x S ) of the black box model calculated based on the baseline value b can be specifically denoted as v(x S |b), and it should be understood that v(x S |b) can be simply written as v(x S ) without causing ambiguity.

[0115] We can use "and interaction utility" and "or interaction utility" to explain the output of the model. Given the black box model output v and an input sample x, when calculating the output v(x S ) on a certain input unit combination S, the model output v(x S ) is decomposed into two terms, i.e. v(x S ) = v and (x S ) + v or (x S ). In this way, we can use and interaction to explain v and (x S ), and use or interaction to explain v or (x S ).

[0116] In the above sub-step (b), the contribution value of the and interaction between any input unit combination in the sample to be analyzed to the output v and (x) corresponding to the and interaction of each model is evaluated by calculating the "and interaction utility". The "and interaction utility" represents the additional utility generated by the and interaction of a black box model when the units in the combination are triggered at the same time (i.e. not masked). It should be understood that in a black box model, input units often do not contribute to the output of the model individually, but interact with each other to affect the output of the model.

[0117] Therefore, for , I and (S|v and ,x) is defined as the "and interaction utility" corresponding to the combination S of the input units of an artificial intelligence black box model, so that when all input variables in S are not masked, Iand (S|v and , x) is triggered and accumulated into the output v and (x) of the black-box model corresponding to the interaction. Specifically, I and (S|v and , x) can be calculated by the following formula.

[0118] The above "and interaction utility" satisfies, for any given input unit combination S of the sample to be analyzed the output v and (x T ) corresponding to the interaction of a person-made artificial intelligence model under this input unit combination and can be decomposed into the sum of all triggered and interactions I and (S|v T , x), as shown in the following formula.

[0119] Where e T represents whether the input variable is blocked. When i∈T, (e i ) T =1, indicating that the input variable i is not blocked (or called triggered), otherwise, (e i ) i∈S =0 indicates that the input variable i is blocked; ∧ T (e i ) i∈S represents whether all units in the input unit combination S are not blocked, and the result obtained by performing the "and" operation on the elements after the ∧ symbol, that is, ∧ T (e i ) i∈S =∏ T (e i )

[0120] Further, the above calculation formula of the and interaction can be equivalent to the following form:

[0121] Assuming and exhaust all input unit combinations of the interaction, then there exists such that I and =Av and .

[0122] In the above sub-step (c), the "or interaction utility" is calculated to evaluate the or relationship between any input unit combination of the sample to be analyzed and the or interaction output v orthe contribution value of (x). The "or-interaction utility" represents the additional utility generated by the or-interaction of the black-box model when at least one unit in the combination is triggered (i.e., not occluded). It should be understood that in the black-box model, the input units tend to not contribute to the output of the model individually, but rather interact with each other to affect the output of the model.

[0123] Therefore, for Definition I or (S|v or ,x) is the "or-interaction utility" corresponding to the combination S of input units, such that when at least one input variable in S is not occluded, I or (S|v or ,x) is triggered and added to the or-interaction output v or (x) of the black-box model. Specifically, I or (S|v or ,x) can be calculated by the following formula.

[0124] The "or-interaction utility" satisfies that for any given combination of input units S of the sample to be analyzed, the or-interaction output v or (x T ) of the black-box model under this combination of input units can be decomposed into the sum of all triggered or-interactions I or (S|v or ,x), as shown in the following formula.

[0125] where e T represents whether the input variable is not occluded. When i∈T, (e T ) i =1, which

[0126] indicates that the input variable i is not occluded, otherwise, (e T ) i =0, which indicates that the input variable i is occluded. ∨ i∈S (e T ) i represents whether at least one unit in the combination S of input units is not occluded, and the result obtained by performing the "or" operation on the elements after the ∨ symbol, i.e., ∨ i∈S (e T ) i =1-∏ i∈S [1-(e T ) i ].

[0127] Further, the calculation formula of the or-interaction can be equivalent to the following form, assuming After enumerating the interactions on all input unit combinations, there are such that or = Bv or .

[0128] In one embodiment of the present application, based on the Magnolia law large model in step 101, the "and interaction utility" and "or interaction utility" calculation formulas are used to calculate the "and interaction utility" and "or interaction utility" values of each model respectively. As shown in FIG. 3, the absolute values of the "and interaction utility" and "or interaction utility" of different input unit combinations obtained by the present application and the traditional method are respectively displayed in descending order to show the distribution.

[0129] In the above sub-step (d), based on the "and interaction utility" and "or interaction utility", the black box model output is explained as the combination of "and interaction utility" and "or interaction utility" between input unit combinations.

[0130] In sub-step (d), based on the "and interaction utility" and "or interaction utility" in step (3), the output v(x T ) of the black box model under any input unit combination T is explained as the combination of "and interaction utility" and "or interaction utility".

[0131] In sub-step (e), based on the combination of "and interaction utility" and "or interaction utility" of each model obtained in sub-step (d), in order to further obtain sparse interactions, the present application aims to minimize the sum of L1-norms of all interactions, that is,

[0132] wherein, ‖·‖1 represents the L1-norm of a vector.

[0133] Optionally, when optimizing the above expression, the following can be defined such that Then the above optimization problem can be converted into the following form. s.t. |q S | < τ S

[0134] wherein, τ S is a preset threshold. In this embodiment,

[0135] Then, step 104 is entered, which can be further divided into the following sub-steps:

[0136] (a) dividing all input units into "relevant input units", "irrelevant input units", and "mutually exclusive units";

[0137] (b) dividing the "interaction utility" of the combination of all input units in the step 103 into "reliable components of interaction utility" and "unreliable components of interaction utility".

[0138] In the above sub-step (a), all input units are divided into "relevant input units", "irrelevant input units", and "mutually exclusive units" according to whether each input unit is relevant to the output of the black-box artificial intelligence model in human cognition, using algorithms or user-defined.

[0139] [According to Rule 26, corrected on 30.04.2025] According to whether each input unit is relevant to the output of the black-box artificial intelligence model in human cognition, the user can customize / empirically divide the set N of all input units into the set R of "relevant input units", the set T of "irrelevant input units", and the set M of "mutually exclusive units" according to specific tasks. Among them, "relevant input units" refer to input units that are strongly related to the output of the model in human cognition, or input units that have a direct causal relationship with the output of the model; "irrelevant input units" refer to input units that are weakly related to the output of the model in human cognition, or input units that have no direct causal relationship with the output of the model; "mutually exclusive units" refer to input units that should not logically affect the output of the model in human cognition. In the Magnolia Legal Big Model, "relevant input units" can be set as input units that are strongly related to criminal behavior and sentencing, or input units that have a direct causal relationship, such as specific criminal behavior; "irrelevant input units" can be set as input units that are weakly related to criminal behavior and sentencing, or input units that have no direct causal relationship, such as crime time, crime location, and character name; "mutually exclusive units" can be set as the description of the case of other people in a multi-person case, such as if the behavior of Zhao 1 in the case description is unrelated to the behavior of Zhao 2, the behavior of Zhao 1 cannot be used as a basis for sentencing Zhao 2. As shown in Figure 2(b), the division of all input units of the sample to be analyzed into "relevant input units", "irrelevant input units", and "mutually exclusive units" is shown.

[0140] In the above sub-step (b), the "interaction utility" of the combination of all input units in the step 103 is divided into "reliable components of interaction utility" and "unreliable components of interaction utility". Among them, "reliable components of interaction utility" are interaction utilities consistent with human cognition, which can be represented as the interaction utility of the combination containing "relevant input units" and not containing "mutually exclusive units"; "unreliable components of interaction utility" are interaction utilities inconsistent with human cognition or in conflict, which can be represented as the interaction utility of the combination containing "irrelevant input units" and "mutually exclusive units".

[0141] In a preferred embodiment, the "and interaction utility" I and (S|v and , x) can be further divided into a "reliable component of and interaction utility" and an "unreliable component of and interaction utility" which can be calculated by the following formula:

[0142] where the set of all input units N is divided into the set of "relevant input units" R, the set of "irrelevant input units" T and the set of "mutually exclusive units" M.

[0143] In a preferred embodiment, the "or interaction utility" I or (S|v or , x) can be further divided into a "reliable component of or interaction utility" and an "unreliable component of or interaction utility" Preferably, the division of an "or interaction utility" is similar to the division of an "and interaction utility", which can be calculated by the following formula:

[0144] where the set of all input units N is divided into the set of "relevant input units" R, the set of "irrelevant input units" T and the set of "mutually exclusive units" M. Another division of an "or interaction utility" considers the proportion of "relevant input units" in the combination S, |S∩R|, to the number of input units in the combination S, |S|, which can be calculated by the following formula:

[0145] Then, step 105 is entered, which can be further divided into the following sub-steps:

[0146] (a) calculating interaction-based evaluation indicators, and counting the proportion of "reliable component of interaction utility" and "unreliable component of interaction utility";

[0147] (b) evaluating the degree of alignment between the model's fine-grained decision logic and human cognition through the statistics of each-order significant interaction utility, reliable component of interaction utility, and unreliable component of interaction utility.

[0148] In the above sub-step (a), the following interaction-based evaluation indicators are calculated, and the proportion of "reliable component of interaction utility" and "unreliable component of interaction utility" is counted. As shown in FIG. 4(a), the user can quantify the degree of alignment between the model's fine-grained decision logic and human cognition with the proportion of "reliable component of interaction utility" in the interaction utility.

[0149] A significant interaction is defined as follows. Given a pre-set threshold τ, a significant interaction is defined as the set of interactions Ω and and Ω or , which can be calculated by the following formula:

[0150] where the threshold τ is set as the 100th largest absolute value of interaction utility after sorting all interaction utilities in descending order of their absolute values.

[0151] The sum of the strength of all significant interaction utilities is:

[0152] The sum of the strength of all interaction utilities is:

[0153] The order of an interaction is defined as the number of input variables in the set S, i.e., order(S) = |S|. For each order o, the sum of the reliable components of interaction utility is:

[0154] For each order o, the strength of the reliable components of interaction utility is:

[0155] The total strength of the reliable components of interaction utility is:

[0156] For each order o, the sum of the unreliable components of interaction utility is:

[0157] For each order o, the strength of the unreliable components of interaction utility is:

[0158] The total strength of the unreliable components of interaction utility is:

[0159] The following index represents the proportion of the strength of the reliable components of interaction utility in the strength of significant interaction utility.

[0160] where the proportion of the strength of the reliable components of interaction utility in the strength of significant interaction utility is higher, indicating that the model's fine decision logic is more aligned with human cognition.

[0161] The following index represents the proportion of the strength of the reliable components of interaction utility in the strength of all interaction utility.

[0162] where the proportion of the strength of the reliable components of interaction utility in the strength of all interaction utility The higher, the more the model's fine-grained decision logic aligns with human cognition.

[0163] The following indicators and respectively represent the proportion of the reliable component of interaction utility in the sum of positive and negative utility of significant interactions at each order.

[0164] or

[0165] or

[0166] where ε represents a very small positive real number to prevent the division by 0 operation of the fraction. For each order of interaction o, the sum of positive interaction utility of significant interactions and the sum of negative interaction utility of significant interactions the positive interaction utility of the reliable component of interaction utility and the negative interaction utility of the reliable component of interaction utility The following indicators and The higher, i.e., the higher the proportion of the reliable component of interaction utility in the positive and negative utility of significant interactions for each order of interaction o, the more the model's fine-grained decision logic aligns with human cognition.

[0167] The following indicator represents the proportion of positive and negative cancellation in the utility of significant interactions at each order.

[0168] where the indicator The higher, the higher the proportion of effective interaction remaining after positive and negative cancellation of significant interaction utility for each order of interaction o. This indicator is helpful to understand whether low-order or high-order interactions play a major role in the model output. A well-trained model tends to model low-order interactions more.

[0169] The following indicator O reliable represents the weighted average order of "reliable interactions" in significant interactions.

[0170] or

[0171] where the weighted average order O reliable The lower, the more the influence of the model's fine-grained decision logic is susceptible to low-order interactions; conversely, the indicator O reliableThe higher, the more susceptible the impact of the model's fine decision logic to high-order interactions.

[0172] Therefore, as shown in FIG. 4(a), the proportion of the reliable component of interaction utility and the unreliable component of interaction utility can be calculated using the evaluation index based on interaction utility, to depict the difference between the model's decision logic and human cognition, and reflect the potential representation defects of the model.

[0173] In the above-mentioned sub-step (b), as shown in FIG. 4(b), the statistical charts of the significant interaction utility of each order, the reliable component of interaction utility, and the unreliable component of interaction utility are respectively displayed, to represent the difference between the model's decision logic and human cognition, and reflect the potential representation defects of the model.

[0174] For each order of interaction o, the sum of the positive interaction utility of the significant interaction Salient + (o), the sum of the negative interaction utility of the significant interaction Salient - (o), the sum of the positive utility of the reliable interaction component Reliable + (o), the sum of the negative utility of the reliable interaction component Reliable - (o), the sum of the positive utility of the unreliable interaction component Unreliable + (o): = Salient + (o)-Reliable + (o), the sum of the negative utility of the unreliable interaction component Unreliable - (o): = Salient - (o)-Reliable - (o).

[0175] For each order of interaction o, when the index and is higher, that is, the proportion of the reliable component of interaction utility in the significant interaction is higher, it indicates that the alignment degree of the model's fine decision logic and human cognition is higher.

[0176] Therefore, as shown in FIG. 4(b), the statistical charts of the significant interaction utility of each order, the reliable component of interaction utility, and the unreliable component of interaction utility can be used to depict the difference between the model's decision logic and human cognition, and reflect the potential representation defects of the model.

[0177] Note that when calculating the index mentioned in all the preferred examples above, a small positive real number ∈ can be added to the denominator to prevent the situation of "division by zero".

[0178] It should be noted that the implementation functions of each module shown in the above-mentioned embodiment of the system for measuring the alignment degree of the fine decision logic of the artificial intelligence black box model and human cognition can be understood with reference to the foregoing description of the method for measuring the alignment degree of the fine decision logic of the artificial intelligence black box model and human cognition. The functions of each module shown in the above-mentioned embodiment of the system for measuring the alignment degree of the fine decision logic of the artificial intelligence black box model and human cognition can be implemented by a program (executable instructions) running on a processor or by a specific logic circuit. The above-mentioned system for measuring the alignment degree of the fine decision logic of the artificial intelligence black box model and human cognition according to the embodiments of the present application, if implemented in the form of software function modules and sold or used as an independent product, can also be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the embodiments of the present application can be embodied in the form of a software product in essence or the part that contributes to the prior art, and the computer software product is stored in a storage medium, including a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the method of the embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read Only Memory), a magnetic disk or an optical disk, and various program code storage media. Thus, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0179] All the documents mentioned in the present application are considered to be included in the disclosure of the present application as a whole, so as to be used as a modification if necessary. In addition, it should be understood that the above is only a preferred embodiment of the present application, and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. within the spirit and principles of one or more embodiments of the present application should be included in the protection scope of one or more embodiments of the present application.

Claims

1. A method for evaluating the alignment degree of fine decision logic of an artificial intelligence black box model with human cognition, characterized in that, The method comprises the following steps: (1) providing a black-box artificial intelligence model; Providing a black-box artificial intelligence model to be analyzed, the black-box artificial intelligence model comprising a pre-trained artificial intelligence model based on a certain data set; (2) selecting input samples and performing identification; selecting an input sample for performing the interactive utility calculation, and identifying the input sample so as to decompose the input sample into n input units and combining the n input units to obtain 2 n combinations of the n input units; the input sample is selected from the group consisting of table data, pictures, text, voice, or a combination thereof; (3) calculating and combining "interaction utility"; 2 in step (2) n The black-box AI model is input with combinations of input units, and the output of the black-box AI model is obtained. Based on the output of the black-box AI model, the interaction utility between the input units is modeled to obtain the "interaction utility" of the black-box AI model for each combination of input units. The output of the black-box AI model on a certain combination of input units is interpreted as a combination of its "interaction utility" on that combination of input units. (4) coarsely dividing "interaction utility"; According to whether each input unit is related to the output of the black-box artificial intelligence model in human cognition, all input units are divided into "relevant input units" and "irrelevant input units" by using an algorithm or being defined by a user; (5) refining the relevant input units and performing further calculation The relevant input units are further divided into "irrelevant input units" and "mutually exclusive units", and the "interaction utility" of the combination of all input units in step (3) is divided into "reliable components of interaction utility" and "unreliable components of interaction utility" based on the "relevant input units", "irrelevant input units" and "mutually exclusive units"; (6) using an interaction-based evaluation index to evaluate the alignment degree of the model decision logic and human cognition; According to the preset interaction-based evaluation index, the proportion of "reliable components of interaction utility" and "unreliable components of interaction utility" is calculated to evaluate the alignment degree of the fine decision logic of the model and human cognition. The order of steps (1) and (2) can be replaced or performed at the same time.

2. The method of claim 1, wherein, The interaction utility can be divided into "and interaction utility" and "or interaction utility", representing the "and relationship" between the input units modeled by the model and the "or relationship" between the input units modeled by the model 3. The method of claim 2, wherein, The "reliable components of and interaction utility" can be calculated as the interaction utility of the combination containing "relevant input units" and not containing "mutually exclusive units"; the "unreliable components of and interaction utility" can be calculated as the interaction utility of the combination not containing "relevant input units" or containing "mutually exclusive units".

4. The method of claim 3, wherein, the reliable component of the interaction utility and the unreliable component of the interaction utility In particular, it can be calculated by the following formula: All input units are divided into a set R of "relevant input units", a set T of "irrelevant input units", and a set M of "mutually exclusive units".

5. The method of claim 2, wherein, The step (4) further comprising the step of: distributing the "or interaction utility" evenly to each input unit involved in the interaction, thereby obtaining a "reliable component of the or interaction utility", which can be calculated as the interaction utility component of the "or interaction utility" that is evenly distributed to "relevant input units"; and a "non-reliable component of the or interaction utility", which can be calculated as the interaction utility component of the "or interaction utility" that is evenly distributed to "irrelevant input units" and "mutually exclusive units", specifically, by the following formula: Where |S∩R| is the number of "relevant input units" in the combination S, and |S| is the number of input units in the combination S.

6. The method of claim 3, wherein, The reliable component of the or interaction utility can also be derived similarly to the division of the interaction utility, specifically, calculated by the following formula: All input units are divided into a set R of "relevant input units", a set T of "irrelevant input units", and a set M of "mutually exclusive units".

7. The method of claim 2, wherein, The evaluation index includes: calculating the strength or proportion of reliable and unreliable components of interaction utility and comparing them as an index; visualizing the statistical chart of each order of significant interaction utility, reliable components of interaction utility, and unreliable components of interaction utility and using it as an index; calculating the proportion of reliable components of interaction utility in significant interaction as an index.

8. The method of claim 7, wherein, The significant interactions are defined as the interaction set Ω larger than a threshold τ and and Ω or may be calculated by the following formula: Wherein, the threshold τ is set as the value of the Xth interaction utility absolute value in the descending order of all interaction utility absolute values, and the X is set by user experience; Ω and is a set of significant interactions in the interaction or is a set of significant interactions in the interaction or 9. The method of claim 8, wherein, The "mutually exclusive units" represent input units that should not affect the model output in human cognitive logic, or input units that are repulsive to the model output in cognitive logic.

10. The method of claim 1, wherein, a reliable component of the interaction utility in a significant interaction The calculation is made by the following formula:

Citation Information

Patent Citations

  • Method and system for explaining public interaction utility among multiple groups of black box artificial intelligence models

    CN117764193A

  • Interactive artificial intelligence value alignment method, system and related system

    CN118536539A

  • Algorithm for realizing artificial intelligence black box model fine decision logic and human cognition alignment degree evaluation

    CN119150077A

  • Explanations generation with different cognitive values using generative adversarial networks

    US20200125975A1

  • System and methods for safe alignment of superintelligence

    WO2024182819A2