Xai model evaluation method and device based on knowledge graph, equipment and medium

By using a knowledge graph-based approach to obtain the explanation results of the XAI model and calculate the score, the problem of time-consuming, labor-intensive, and ineffective traditional XAI evaluation methods is solved, achieving a more efficient and accurate evaluation.

CN115905558BActive Publication Date: 2026-05-29NANJING TRANSWARP INTELLIGENCE CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING TRANSWARP INTELLIGENCE CO LTD
Filing Date
2022-11-10
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing XAI model evaluation methods lack standardized quantitative measurement methods. Manual evaluation is resource-intensive and inaccurate and inefficient. Furthermore, user evaluations are easily influenced by subjective factors, making it impossible to guarantee the effectiveness of the evaluation.

Method used

A knowledge graph-based approach is adopted. By obtaining the explanation result pairs of the XAI model to be evaluated, matching the node sequence pairs in the knowledge graph, determining the subgraph set corresponding to each node sequence, and calculating the score pairs of explanation result pairs based on the subgraph set, including scores for explanation coherence, complexity, and credibility.

Benefits of technology

This improves the accuracy and efficiency of XAI model evaluation, ensures the validity of evaluation data, and reduces the resource consumption and subjective influence of manual evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115905558B_ABST
    Figure CN115905558B_ABST
Patent Text Reader

Abstract

The application discloses an XAI model evaluation method and device based on a knowledge graph, equipment and a storage medium. The method comprises the following steps: obtaining an explanation result pair corresponding to an XAI model to be evaluated; obtaining a node sequence pair matched with the explanation result pair in a knowledge graph; determining a subgraph set corresponding to each node sequence according to the node sequence pair and the knowledge graph; and determining a score pair corresponding to the explanation result pair according to the subgraph set corresponding to each node sequence. Through the technical scheme, the accuracy and efficiency of the evaluation selection method can be improved, and the effectiveness of the evaluation data is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of artificial intelligence technology, and in particular to a method, apparatus, device and medium for evaluating XAI models based on knowledge graphs. Background Technology

[0002] With the large-scale application of artificial intelligence, trustworthy AI technology has received increasing attention, and explainable AI (XAI) is an important branch of trustworthy AI technology. Similar to traditional machine learning models, XAI algorithms and models also require optimization and selection from a set of models. However, traditional machine learning model selection methods are not applicable to XAI models because XAI model evaluation metrics include not only the model's accuracy but also the expression and comprehensibility of the explanation results. Furthermore, there are evaluations of the ability of multiple explanation methods to explain the same result, and the evaluation content often involves the evaluator's cognitive abilities, making it difficult to form a unified standard or universal method. Most studies focusing on evaluating user effectiveness concentrate on subjective measurement, requiring people to manually rate the XAI model's explanation based on some given metrics. Some studies, however, measure both the subjective usability of the explanation and the participants' ability to make correct inferences based on the explanation, allowing people to distinguish between the behavioral effect and self-perceived effect of the explanation, thus emphasizing the value of objective measurement to some extent.

[0003] Therefore, existing XAI evaluation methods have significant drawbacks: on the one hand, there is a lack of standard quantitative measurement methods, and manual evaluation consumes a lot of resources and is inaccurate and inefficient; on the other hand, user evaluations of the model are easily affected by many (especially subjective) factors, and the validity of the evaluation data cannot be guaranteed. Summary of the Invention

[0004] This invention provides a knowledge graph-based method, apparatus, device, and storage medium for evaluating XAI models, which solves the problems of traditional XAI evaluation and selection methods being time-consuming, labor-intensive, lengthy, and having low evaluation effectiveness.

[0005] According to one aspect of the present invention, a knowledge graph-based XAI model evaluation method is provided, comprising:

[0006] Obtain the explanation result pairs corresponding to the XAI model to be evaluated;

[0007] Obtain the node sequence pairs in the knowledge graph that match the explanation result pair;

[0008] Determine the set of subgraphs corresponding to each node sequence based on the node sequence pairs and the knowledge graph;

[0009] The score pair corresponding to the interpretation result is determined based on the subgraph set corresponding to each node sequence.

[0010] According to another aspect of the present invention, a knowledge graph-based XAI model evaluation device is provided, the knowledge graph-based XAI model evaluation device comprising:

[0011] The explanation result acquisition module is used to acquire the explanation result pairs corresponding to the XAI model to be evaluated.

[0012] The node sequence pair acquisition module is used to acquire node sequence pairs in the knowledge graph that match the explanation result pair;

[0013] The subgraph set determination module is used to determine the subgraph set corresponding to each node sequence based on the node sequence pairs and the knowledge graph.

[0014] The rating pair determination module is used to determine the rating pair corresponding to the interpretation result pair based on the subgraph set corresponding to each node sequence.

[0015] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0016] At least one processor; and

[0017] A memory communicatively connected to the at least one processor; wherein,

[0018] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the knowledge graph-based XAI model evaluation method according to any embodiment of the present invention.

[0019] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the knowledge graph-based XAI model evaluation method according to any embodiment of the present invention.

[0020] This invention addresses the problems of traditional XAI evaluation and selection methods, such as time-consuming and labor-intensive processes, long processes, and low evaluation effectiveness, by obtaining explanation result pairs corresponding to the XAI model to be evaluated; obtaining node sequence pairs in the knowledge graph that match the explanation result pairs; determining the subgraph set corresponding to each node sequence based on the node sequence pairs and the knowledge graph; and determining the rating pair corresponding to each explanation result pair based on the subgraph set corresponding to each node sequence. This improves the accuracy and efficiency of the evaluation and selection method and ensures the validity of the evaluation data.

[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0022] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a flowchart of a knowledge graph-based XAI model evaluation method according to Embodiment 1 of the present invention;

[0024] Figure 2 This is a sub-graph set G obtained from a matching result in Embodiment 1 of the present invention. a A diagram of [].

[0025] Figure 3 This is a sub-graph set G obtained from a matching result in Embodiment 1 of the present invention. b A diagram of [].

[0026] Figure 4 This is a schematic diagram of the structure of a knowledge graph-based XAI model evaluation device according to Embodiment 2 of the present invention;

[0027] Figure 5 This is a schematic diagram of the structure of an electronic device according to Embodiment 3 of the present invention. Detailed Implementation

[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0030] Example 1

[0031] Figure 1 This is a flowchart of a knowledge graph-based XAI model evaluation method according to Embodiment 1 of the present invention. This embodiment is applicable to the comprehensive ranking evaluation of the interpretable performance of XAI models in multiple randomized comparative tests. This method can be executed by the knowledge graph-based XAI model evaluation device in this embodiment, which can be implemented in software and / or hardware, such as... Figure 1 As shown, the method specifically includes the following steps:

[0032] S110, Obtain the explanation result pair corresponding to the XAI model to be evaluated.

[0033] Among them, the explanation result pairs corresponding to the XAI model to be evaluated can be explanation result pairs from different training batches of the same XAI model, or explanation result pairs from different XAI models.

[0034] Specifically, one method to obtain the explanation result pair corresponding to the XAI model to be evaluated is to randomly select two explanation results from different training batches of the same XAI model to form an explanation result pair. Another method is to randomly select two explanation results from different XAI models to form an explanation result pair.

[0035] Specifically, the atomic representation of the interpretation result can be E = [X1...X2] n →Y], or simply atomic interpretation result, where X is the factor, Y is the result factor, and n is the number of factors. Therefore, the interpretation result can be represented as E pair ={E a [],E b []},E a []、E b[] represents the atomic interpretation result sequences with subscripts a and b, respectively. It should be noted that, based on the definition and combination of atomic interpretation results and atomic interpretation result sequences, various complex interpretations can be flexibly represented, such as one effect with one cause, one effect with multiple causes, multiple effects with one cause, and multiple effects with multiple causes.

[0036] S120, obtain the node sequence pairs in the knowledge graph that match the explanation result pair.

[0037] The knowledge graph (KG) can be used for matching, either by matching nodes within the KG or by matching attributes of the KG; this is not a specific limitation. Node sequence pairs are the corresponding results obtained after matching the interpretation result pairs with the KG.

[0038] Specifically, the method for obtaining the node sequence pairs that match the explanation result pairs in the knowledge graph can be as follows: obtain the explanation result pairs, match the explanation result pairs with the node sequences and attribute sequences of the KG, and then perform KG semantic retrieval matching processing to obtain the node sequence pairs that match the explanation result pairs.

[0039] Specifically, the nodes of KG can be denoted as V, the edges as E, and the node attributes as A. The node sequence of KG is V[], and the attribute sequence is A[]. First, if E a []、E b Factors in brackets [] are represented without semantic information. Therefore, based on the data's metadata, the specific factors are replaced with field names to make E... a []、E b The sequence [] is a sequence E containing semantic information. a ′[]、E b ′[], where the metadata of the data is the basic information in the data processing system and dataset management system. Next, E a ′[]、E b The semantic factor of ′[] calls KG's semantic retrieval interface for query matching, and obtains E pair The corresponding node sequence pair in KG is V. pair ={V a [],V b []}. It should be noted that, in order to improve matching efficiency, a reverse traversal search algorithm from the final effect to the original cause is adopted. The matching implementation algorithm is as follows: First, call the KG semantic retrieval interface to match i in KG. n The factor, in KG, is the node that matches V. n Second, with V n Starting from point D, match i again within the path depth D. n-1 Factors, if a match is successful, then use i n-1Repeat this sub-step starting from the previous step. If a match fails, then match within a depth of (number of unsuccessful consecutive matches + 1) * D, which is a depth of 2 * D. n-2 Factors, and use this rule to complete factor matching for all I[]; third, output the matching result V[] = {...V i ...}

[0040] It should be noted that, during the matching process with the interpretation results, considering that the structure and size of KG may differ in actual applications, the search depth can be adjusted according to actual needs when performing the reverse traversal search algorithm for node matching.

[0041] S130, determine the set of subgraphs corresponding to each node sequence based on the node sequence pairs and the knowledge graph.

[0042] The subgraph set can be one or more subgraphs, which are the paths corresponding to node sequence pairs in the KG. Multiple subgraphs indicate that the path is broken.

[0043] Specifically, the method for determining the subgraph set corresponding to each node sequence based on node sequence pairs and the knowledge graph can be as follows: obtain node sequence pairs, and obtain the subgraph set for each node sequence through the path lookup function of the knowledge graph.

[0044] Specifically, the node sequence pairs are V pair ={V a [],V b []}, using KG's pathfinding function, we obtain V pair The corresponding path is denoted as G. pair ={G a [],G b []},G a [] is V a The path corresponding to [] is one or more subgraphs, which determines V. a The set of subgraphs corresponding to [] is G. a [], Similarly, G b [] is V b The path corresponding to [] determines V. b The set of subgraphs corresponding to [] is G. b [].

[0045] S140, determine the corresponding score pair for the interpretation result based on the subgraph set corresponding to each node sequence.

[0046] The explanation result is paired with a rating pair, which is the total rating pair calculated based on the explanation coherence rating, explanation complexity rating, and explanation credibility rating of the subgraph set corresponding to each node sequence.

[0047] Specifically, the method for determining the score pair corresponding to the explanation result pair based on the subgraph set corresponding to each node sequence can be as follows: obtain the subgraph set corresponding to each node sequence, determine the score of explanation coherence, explanation complexity, and explanation credibility for each node sequence based on the subgraph set corresponding to each node sequence, calculate the total score of the subgraph set corresponding to each node sequence based on the score of explanation coherence, explanation complexity, and explanation credibility, and then determine the score pair corresponding to the explanation result pair.

[0048] Optionally, the scoring pairs corresponding to the explanation results are determined based on the subgraph set corresponding to each node sequence, including:

[0049] The explanation coherence score, explanation complexity score, and explanation credibility score are determined based on the subgraph set corresponding to each node sequence.

[0050] The target score for each node sequence is determined based on the explanation coherence score, explanation complexity score, and explanation credibility score corresponding to each node sequence.

[0051] The interpretation result is determined based on the target score corresponding to each node sequence, and the corresponding score pair is determined.

[0052] The target score is the total score calculated based on the explanation coherence score, explanation complexity score, and explanation credibility score corresponding to each node sequence.

[0053] Specifically, the method for determining the explanation coherence score, explanation complexity score, and explanation credibility score for each node sequence based on the subgraph set corresponding to each node sequence can be as follows: obtain the subgraph set corresponding to each node sequence; obtain the explanation coherence score by measuring the number of subgraphs in the subgraph set corresponding to each node sequence; obtain the explanation complexity score by measuring the number of nodes and edges in each subgraph in the subgraph set corresponding to each node sequence; and obtain the explanation credibility score by calculating the sum of the edge weights of all subgraphs in the subgraph set corresponding to each node sequence.

[0054] Specifically, the method for determining the target score for each node sequence based on the explanation coherence score, explanation complexity score, and explanation credibility score can be as follows: calculate the total score for each node sequence based on the explanation coherence score, explanation complexity score, and explanation credibility score corresponding to each node sequence.

[0055] Specifically, the total score for each node sequence can be calculated as follows: using G a For example, according to G a The determination of the number of neutron graphs explains the coherence score S. split According to Ga The complexity of the interpretation of [] determines G a [] corresponds to the explanation complexity score S complexity_a_total According to the subgraph G containing the outcome factors a_target The weights of each side determine the credibility score S. credit The total score is calculated as follows:

[0056] S Ga[] =standard(S split )+standard(S credit )+standard(S complexity_a_total );

[0057] Here, standard() is the data standardization function, which will not be elaborated on here.

[0058] Specifically, the method for determining the score pair corresponding to the explanation result pair based on the target score corresponding to each node sequence can be as follows: determine the total score corresponding to each node sequence based on the explanation coherence score, explanation complexity score, and explanation credibility score corresponding to each node sequence, and obtain the total score pair based on the total score corresponding to each node sequence, that is, determine the score pair corresponding to the explanation result pair.

[0059] It should be noted that after obtaining the scoring pairs, the scoring results are stored in Rsc[i]. One of the comparison results can be represented as an ordered positive integer pair (a, b), where a and b are the selected model labels to be evaluated, and i is the current comparison round number. All comparison results are stored in the result set Rsc[], the size of which is the hyperparameter N, i.e. the total number of comparison rounds. Steps S110 to S140 are repeated until the number of comparison rounds reaches the hyperparameter N.

[0060] Optionally, the explanatory coherence score for each node sequence is determined based on the set of subgraphs corresponding to each node sequence, including:

[0061] Get the number of subgraphs in the subgraph set corresponding to the node sequence;

[0062] The explanatory coherence score corresponding to the node sequence is determined based on the number of subgraphs in the subgraph set.

[0063] The coherence scoring strategy is as follows: a logically coherent explanation should have a coherent reasoning process, and the nodes corresponding to its causal factors in the KG should form a connected subgraph. The more disconnected subgraphs there are, the worse the coherence of the explanation is considered.

[0064] Specifically, the method to obtain the number of subgraphs in the subgraph set corresponding to the node sequence can be: obtain the subgraph set corresponding to the node sequence, and then obtain the number of subgraphs in the subgraph set.

[0065] Specifically, the method for determining the explanatory coherence score corresponding to a node sequence based on the number of subgraphs in the subgraph set can be as follows: obtain the number of subgraphs in the subgraph set, and the explanatory coherence score corresponding to the node sequence is the reciprocal of the number of subgraphs. For example, it could be based on G... a For example, G a The number of subgraphs in [] is |G a []|, the coherence score is explained as follows:

[0066]

[0067] That is, the coherence score is the reciprocal of the number of subgraphs.

[0068] Optionally, the explanation complexity score for each node sequence is determined based on the set of subgraphs corresponding to each node sequence, including:

[0069] The complexity of a subgraph is determined by the number of nodes and edges in each subgraph within the set of subgraphs corresponding to the node sequence.

[0070] The sum of the complexities of all subgraphs in the set of subgraphs corresponding to the node sequence is used as the interpretation complexity score for the node sequence.

[0071] The explanation complexity scoring strategy is as follows: the smaller the ratio of nodes to edges in a subgraph, the more effective, concise, and persuasive the explanation is. An edge can be understood as a connection between two nodes in each subgraph.

[0072] Specifically, the method to determine the complexity of a subgraph based on the number of nodes and edges in each subgraph of the subgraph set corresponding to the node sequence can be as follows: obtain each subgraph in the subgraph set corresponding to the node sequence, obtain the number of nodes and edges in each subgraph based on each subgraph in the subgraph set, and obtain the complexity of each subgraph based on the number of nodes / edges in each subgraph.

[0073] Specifically, the method to determine the interpretation complexity score of a node sequence by summing the complexities of all subgraphs in the subgraph set corresponding to the node sequence can be as follows: obtain the complexity of each subgraph in the subgraph set, and determine the interpretation complexity score of the node sequence by summing the complexities of all subgraphs in the subgraph set. For example, it could be based on G... a For example, G a Each subgraph in brackets [] is represented by a triplet. Let V be the set of nodes and E be the set of edges. Let G be the set of edge weights. a A subgraph G of [] a_i The corresponding triple is The specific metric for the complexity of this subgraph is:

[0074]

[0075] Among them, |V a_i |For subgraph G a_i The number of nodes, |E a_i | is the number of its edges (edges shared by multiple paths are counted multiple times), S complexity_a_i It is the complex graph score of the i-th subgraph. This is achieved by evaluating the set of subgraphs G. a The total score of [] is obtained by adding the scores of each subgraph, and the calculation method is as follows:

[0076]

[0077] Among them, S complexity_a_total For G a The overall score for the complexity of the explanation in brackets [].

[0078] Optionally, the explanatory credibility score for each node sequence is determined based on the subgraph set corresponding to each node sequence, including:

[0079] Obtain the target subgraph from the set of subgraphs corresponding to the node sequence, wherein the target subgraph is a subgraph that includes the result factor;

[0080] The sum of the weights of all edges in the target subgraph is used to determine the credibility score of the explanation corresponding to the node sequence.

[0081] The credibility scoring strategy is as follows: if edge weights reflect the strength of connections and logical relationships between node entities, then the sum of all edge weights in a subgraph can simply represent the credibility of that subgraph. The edge weights in the KG originate from the KG construction process, employing relation weights from a Bayesian structure construction method.

[0082] Specifically, the method of obtaining the target subgraph in the subgraph set corresponding to the node sequence, wherein the target subgraph is a subgraph including the result factor, can be: obtaining the subgraph set corresponding to the node sequence, and selecting the subgraph including the result factor according to the subgraph set.

[0083] Specifically, the method for determining the explanation credibility score corresponding to the node sequence by summing the weights of all edges in the target subgraph can be as follows: Select the subgraph containing the result factor from the set of subgraphs corresponding to the node sequence, and calculate the sum of the weights of all edges in the selected subgraph to determine the explanation credibility score corresponding to the node sequence. For example, it could be based on G... a For example, from G a [] Select the subgraph containing the outcome factors, denoted as G. a_target Calculate G a_target Sum of the weights of all edges in the array:

[0084] S credit=∑w e_i

[0085] Among them, w e_i This represents the weight of the i-th edge in the subgraph.

[0086] Optionally, the explanatory credibility score for each node sequence is determined based on the subgraph set corresponding to each node sequence, including:

[0087] Obtain the target subgraph from the set of subgraphs corresponding to the node sequence, wherein the target subgraph is a subgraph that includes the result factor;

[0088] The sum of the weights of the edges connected to the result factor in the target subgraph is used as the explanation credibility score for the node sequence.

[0089] Specifically, the method for determining the explanatory credibility score corresponding to the node sequence by the sum of the weights of the edges connected to the result factors in the target subgraph can be as follows: select the edges connected to the result factors in the target subgraph, calculate the edge proportion, calculate the edge weight based on the edge proportion, and determine the explanatory credibility score corresponding to the node sequence by the sum of the weights of the edges connected to the result factors.

[0090] It should be noted that when determining the corresponding rating pairs for the interpretation results, technical personnel can adjust the specific calculation and combination methods in the rating strategy according to different evaluation scenarios and the actual utility of the ratings.

[0091] Optionally, after determining the score pairs corresponding to the explanation results based on the subgraph set corresponding to each node sequence, the method further includes:

[0092] By inputting the ratings into the target model, we can obtain the ranking of the explanatory power of each XAI model, or the ranking of the explanatory power of each XAI explanation result.

[0093] The target model can be a model that takes rating pairs as input and outputs a ranking result of the explanatory power of the XAI model to be evaluated. It should be noted that, in order to improve accuracy and reduce error, the explanatory results are compared in a random manner. Therefore, the number of comparisons for different XAI models is not the same, and a target model is established to obtain the ranking order based on the imbalanced comparison results.

[0094] Specifically, the rating pairs are input into the target model to obtain the ranking of the explanatory power of each XAI model. Alternatively, the ranking of the explanatory power of each XAI explanation result can be as follows: if the initially selected XAI models to be evaluated are two different XAI models, then the rating pairs of the explanation results corresponding to the two different XAI models are input into the target model to obtain the ranking of the explanatory power of each XAI model; if the initially selected XAI models to be evaluated are different training batches of the same XAI model, then the rating pairs of the explanation results from different training batches of the same XAI model are input into the target model to obtain the ranking of the explanatory power of each XAI explanation result.

[0095] For example, if the input to the target model is the result set Rsc[] of the stored rating pairs, and the output is the ranking result of the explanatory power of different XAI models, denoted as {...α i ...α j ...}, where α i Let be the ranking order of the explanatory power of the i-th XAI model. Its order in the ranking results is its ranking order of explanatory power. The specific implementation method is as follows: First, the comparison results stored in the result set array Rsc[] are converted into a matrix A that stores the comparison results. n×n In the middle, a ij The position in the matrix is ​​(i,j), which represents the number of times the i-th XAI model outperforms the j-th XAI model; secondly, the explanatory power α of each model is calculated. i The specific value of α, and its specific solution method: Given any α... i After setting the explanatory power baseline value to 1, ArgMax(L) is solved, and the calculation of L is expressed as follows:

[0096]

[0097] Among them, a j Let represent the ranking of the explanatory power of the j-th XAI model.

[0098] It's important to note that when weighted equally, for any j, summing the terms of this loss function over i is simply a constant minus the "rank-sum test," a non-parametric test that achieves the highest statistical power and is proportional to the common AUC metric. Therefore, optimizing this loss function can be understood as optimizing the weighted sum of AUCs. This gives it strong extensibility and practicality. After solving ArgMax(L), the specific explanatory power of each model can be obtained, sorted, and stored as a sequence of model explanatory power {α}. i ...α j ...}, here α i This corresponds to the i-th explanatory model ranked first in terms of explanatory power. iBased on the above approach, we can obtain the ranking of the explanatory power of each XAI model, or the ranking of the explanatory power of each XAI explanation result.

[0099] It should be noted that technicians can adjust the storage method of the model comparison results according to actual needs, and are not limited to storing them in matrix form.

[0100] In a specific example, if the recommendation business department of an internet company H stores data of 10,000 users in its sample library, Data[], for a specific user, the data is represented as P = {Pur[], Foo[], rec}, where Pur[] is the set of the user's past purchase records, represented as strings (e.g., "umbrella", "banana"); Foo[] is the set of browsing (but not purchasing) records, represented as strings; and rec is the product recommended to the user by the model based on the above two sets of arrays, represented as strings. The department wants to add explanations to the original recommendation system to improve user acceptance of the recommended products. Several XAI explanation models, MODEL1, MODEL2, and MODEL3, have been selected to provide reasonable explanations for their recommendation results. The evaluation focuses on the path explanation process—inferring the product the user is most likely to want to buy based on information about the products the user has already purchased and browsed—and the model with the strongest explanation ability is selected for implementation. The system administrator can set the comparison rounds parameter N=20. First, a random algorithm is used to select two XAI models to be compared, MODEL1 and MODEL2 (hereinafter referred to as a and b respectively). Then, a user's data P is randomly selected from the sample library Data[] and fed into a and b to generate the interpretation results. The atomic representation of the output interpretation results is E=[X1……X n →Y], where X i The explanatory model selects the dependent factors to participate in the interpretation of the results; therefore, the explanatory results can be represented as E. pair ={E a [],E b []},E a []、E b [] represents the sequence of atomic interpretation results with subscripts a and b, respectively.

[0101] If user P's data is Pur a [] = {"diapers", "Microeconomics", "baby crib"...}, Foo[] = {"SK-II", "water bottle", "maternity clothes"...}, rec = "milk powder". E a The data in brackets is as follows:

[0102] X1 chewing gum <![CDATA[X2]]> cold medicine <![CDATA[X3]]> crib <![CDATA[X4]]> Parenting Bible <![CDATA[X5]]> diapers <![CDATA[X6]]> baby bottle Y milk powder

[0103] E bThe data in brackets is as follows:

[0104] <![CDATA[X1]]> chewing gum <![CDATA[X2]]> cold medicine <![CDATA[X3]]> 100,000 Whys <![CDATA[X4]]> C++ Primer <![CDATA[X5]]> bird's nest Y milk powder

[0105] After obtaining the interpreted sequence pairs, they are matched with KG. To improve matching efficiency, E... a For example, matching [].

[0106] 1. Call the Sophon KG semantic retrieval interface to match the factor corresponding to Y in the KG. The node matched in the KG is V. y .

[0107] 2. With V y Starting from point D, match node V6 corresponding to X6 within a path depth of D. If the match is successful, repeat this step starting from V6; if the match fails, match nodes corresponding to factor X5 within a depth of (number of unsuccessful consecutive matches + 1) * D, assuming that the subgraph containing the previous node is not connected to the subgraph containing the next node. Continue this rule until all nodes are matched.

[0108] Final output matching result: E pair The corresponding node sequence pair in KG is V. pair ={V a [],V b []}.

[0109] The pathfinding function of Sophon KG can be used to obtain the node sequence pair V. pair ={V a [],V b The set of subgraphs corresponding to []} is G. pair ={G a [],G b The specific matching result is as follows: []}

[0110] G a []={{X1,X2},{X3,X4,X5,X6,Y}}

[0111] G b [] = {{X1,X2},{X3,X4},{X5,Y}}

[0112] Figure 2 This is a sub-graph set G obtained from a matching result in Embodiment 1 of the present invention. a A diagram of [], as shown Figure 2 As shown, G a There are 2 subgraphs in [].

[0113] Figure 3 This is a sub-graph set G obtained from a matching result in Embodiment 1 of the present invention. bA diagram of [], as shown Figure 3 As shown, G b There are 3 subgraphs in [].

[0114] According to G pair ={G a [],G b []} is used for scoring, in order to Figure 2 G in a For example, []

[0115] according to Measuring G a [] indicates the coherence score of the explanation.

[0116] according to The calculation result is accurate to five decimal places and is used to measure G. a The complexity score for the explanation of [].

[0117] according to Figure 2 subgraph G in a_2 [] represents the subgraph containing the result factor Y. The weights of each edge in the subgraph are already indicated. S is calculated. credit =7.623, used to measure G a [] represents the credibility score of the explanation.

[0118] According to G a The total score is calculated using the explanation coherence score, explanation complexity score, and explanation credibility score within the brackets []. If the standard() function uses the logistic function... The calculation result is S Ga[] =0.99996, similarly, S Gb[] =0.97629, accurate to five decimal places. After comparison, MODEL1 wins over MODEL2 in this round.

[0119] After repeating N=20 times, a result array Rsc[] containing the comparison results of 20 times is obtained. The data contained in Rsc[] is transformed into a 3×3 matrix A, where the element a in A is... ij This represents the number of times the i-th model outperforms the j-th model across all comparison rounds. The data for A is as follows:

[0120]

[0121] Let α1, α2, and α3 be the specific values ​​of the explanatory power of MODEL1, MODEL2, and MODEL3, respectively. Specify the baseline value of explanatory power α1 = 1.00. α can be obtained by solving ArgMax(L). i The calculation of L is expressed as follows:

[0122]

[0123] Using optimization algorithms such as BFGS (but not limited to), we can find that when L reaches its maximum value, α1 = 1.00, α2 = 0.89, and α3 = 0.66. The explanatory power sequence of these three models is {α1 = 1.00, α2 = 0.89, α3 = 0.66}. Therefore, the XAI model MODEL1 with the strongest explanatory power is selected for use. It should be noted that only three XAI models were compared in this specific example; in actual applications, the number of XAI models to be evaluated is not limited.

[0124] The technical solution of this embodiment solves the problems of time-consuming, labor-intensive, lengthy, and low evaluation effectiveness of traditional XAI evaluation and selection methods by obtaining the explanation result pairs corresponding to the XAI model to be evaluated; obtaining the node sequence pairs that match the explanation result pairs in the knowledge graph; determining the subgraph set corresponding to each node sequence based on the node sequence pairs and the knowledge graph; and determining the rating pair corresponding to the explanation result pairs based on the subgraph set corresponding to each node sequence. This improves the accuracy and efficiency of the evaluation and selection method and ensures the validity of the evaluation data.

[0125] Example 2

[0126] Figure 4 This is a schematic diagram of a knowledge graph-based XAI model evaluation device according to Embodiment 2 of the present invention. This embodiment is applicable to the comprehensive ranking evaluation of the interpretable performance of XAI models in multiple randomized comparative tests. The device can be implemented in software and / or hardware, and can be integrated into any device that provides knowledge graph-based XAI model evaluation functionality, such as… Figure 4 As shown, the device for evaluating the XAI model based on knowledge graph specifically includes: an explanation result pair acquisition module 210, a node sequence pair acquisition module 220, a subgraph set determination module 230, and a score pair determination module 240.

[0127] Among them, the explanation result acquisition module 210 is used to acquire the explanation result pair corresponding to the XAI model to be evaluated;

[0128] The node sequence pair acquisition module 220 is used to acquire node sequence pairs in the knowledge graph that match the explanation result pair;

[0129] Subgraph set determination module 230 is used to determine the subgraph set corresponding to each node sequence based on the node sequence pairs and the knowledge graph;

[0130] The rating pair determination module 240 is used to determine the rating pair corresponding to the interpretation result pair based on the subgraph set corresponding to each node sequence.

[0131] Optionally, the scoring module is specifically used for:

[0132] The explanation coherence score, explanation complexity score, and explanation credibility score are determined based on the subgraph set corresponding to each node sequence.

[0133] The target score for each node sequence is determined based on the explanation coherence score, explanation complexity score, and explanation credibility score corresponding to each node sequence.

[0134] The interpretation result is determined based on the target score corresponding to each node sequence, and the corresponding score pair is determined.

[0135] Optionally, the scoring module is specifically used for:

[0136] Get the number of subgraphs in the subgraph set corresponding to the node sequence;

[0137] The explanatory coherence score corresponding to the node sequence is determined based on the number of subgraphs in the subgraph set.

[0138] Optionally, the scoring module is specifically used for:

[0139] The complexity of a subgraph is determined by the number of nodes and edges in each subgraph within the set of subgraphs corresponding to the node sequence.

[0140] The sum of the complexities of all subgraphs in the set of subgraphs corresponding to the node sequence is used as the interpretation complexity score for the node sequence.

[0141] Optionally, the scoring module is specifically used for:

[0142] Obtain the target subgraph from the set of subgraphs corresponding to the node sequence, wherein the target subgraph is a subgraph that includes the result factor;

[0143] The sum of the weights of all edges in the target subgraph is used to determine the credibility score of the explanation corresponding to the node sequence.

[0144] Optionally, the scoring module is specifically used for:

[0145] Obtain the target subgraph from the set of subgraphs corresponding to the node sequence, wherein the target subgraph is a subgraph that includes the result factor;

[0146] The sum of the weights of the edges connected to the result factor in the target subgraph is used as the explanation credibility score for the node sequence.

[0147] Optional, also includes:

[0148] The ranking module is used to input the scores into the target model to obtain the ranking order of the explanatory power of each XAI model, or the ranking order of the explanatory power of each XAI explanation result.

[0149] The above-described products can perform the methods provided in any embodiment of the present invention, and have the corresponding functional modules and beneficial effects for performing the methods.

[0150] The technical solution of this embodiment solves the problems of time-consuming, labor-intensive, lengthy, and low evaluation effectiveness of traditional XAI evaluation and selection methods by obtaining the explanation result pairs corresponding to the XAI model to be evaluated; obtaining the node sequence pairs that match the explanation result pairs in the knowledge graph; determining the subgraph set corresponding to each node sequence based on the node sequence pairs and the knowledge graph; and determining the rating pair corresponding to the explanation result pairs based on the subgraph set corresponding to each node sequence. This improves the accuracy and efficiency of the evaluation and selection method and ensures the validity of the evaluation data.

[0151] Example 3

[0152] Figure 5 This is a schematic diagram of an electronic device according to Embodiment 3 of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0153] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0154] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0155] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as the XAI model evaluation method based on knowledge graphs.

[0156] In some embodiments, the knowledge graph-based XAI model evaluation method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the knowledge graph-based XAI model evaluation method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the knowledge graph-based XAI model evaluation method by any other suitable means (e.g., by means of firmware).

[0157] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0158] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0159] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0160] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0161] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0162] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0163] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0164] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A knowledge graph-based evaluation method for XAI models, characterized in that, include: Obtain the explanation result pairs corresponding to the XAI model to be evaluated, including: A user data point is randomly selected from the recommendation sample library. The user data includes the user's past purchase records and browsing but not purchasing records. The set of past purchase records and the set of browsing but not purchasing records are used as input data and simultaneously fed into two XAI models to be evaluated. Each XAI model infers and determines the products to recommend to the user based on the input data. Each XAI model selects a factor from the input data to explain the rationality of the recommendation, and generates a corresponding atomic explanation result based on the factor and the recommended product. The atomic explanation results generated by each model are combined into a corresponding explanation result sequence, resulting in an explanation result pair consisting of two explanation result sequences. The XAI model to be evaluated is used to assess the path explanation in the recommendation system, specifically inferring the product a user currently wants to purchase based on information about products the user has already purchased and browsed. The evaluation also involves obtaining node sequence pairs in the knowledge graph that match the explanation result pair, including: Obtain the explanation result pairs, match the explanation result pairs with the node sequence and attribute sequence of KG, and after KG semantic retrieval matching processing, obtain the node sequence pairs that match the explanation result pairs; Determine the set of subgraphs corresponding to each node sequence based on the node sequence pairs and the knowledge graph; The interpretation result is determined based on the subgraph set corresponding to each node sequence, including: The explanation coherence score, explanation complexity score, and explanation credibility score are determined based on the subgraph set corresponding to each node sequence. The target score for each node sequence is determined based on the explanation coherence score, explanation complexity score, and explanation credibility score corresponding to each node sequence. Determine the corresponding score pair for the interpretation result based on the target score corresponding to each node sequence; Based on the rating pairs, a sequence of explanatory power corresponding to each XAI model to be evaluated is determined. Based on multiple explanatory power sequences, the XAI model with the strongest explanatory power is selected for use in the recommendation system.

2. The method according to claim 1, characterized in that, The explanatory coherence score for each node sequence is determined based on the subgraph set corresponding to each node sequence, including: Get the number of subgraphs in the subgraph set corresponding to the node sequence; The explanatory coherence score corresponding to the node sequence is determined based on the number of subgraphs in the subgraph set.

3. The method according to claim 1, characterized in that, The explanation complexity score for each node sequence is determined based on the set of subgraphs corresponding to each node sequence, including: The complexity of a subgraph is determined by the number of nodes and edges in each subgraph within the set of subgraphs corresponding to the node sequence. The sum of the complexities of all subgraphs in the set of subgraphs corresponding to the node sequence is used as the interpretation complexity score for the node sequence.

4. The method according to claim 1, characterized in that, The explanation credibility score for each node sequence is determined based on the subgraph set corresponding to each node sequence, including: Obtain the target subgraph from the set of subgraphs corresponding to the node sequence, wherein the target subgraph is a subgraph that includes the result factor; The sum of the weights of all edges in the target subgraph is used to determine the credibility score of the explanation corresponding to the node sequence.

5. The method according to claim 1, characterized in that, The explanation credibility score for each node sequence is determined based on the subgraph set corresponding to each node sequence, including: Obtain the target subgraph from the set of subgraphs corresponding to the node sequence, wherein the target subgraph is a subgraph that includes the result factor; The sum of the weights of the edges connected to the result factor in the target subgraph is used as the explanation credibility score for the node sequence.

6. The method according to claim 1, characterized in that, After determining the score pairs corresponding to the interpretation results based on the subgraph set corresponding to each node sequence, the process also includes: By inputting the ratings into the target model, we can obtain the ranking of the explanatory power of each XAI model, or the ranking of the explanatory power of each XAI explanation result.

7. A knowledge graph-based XAI model evaluation device, characterized in that, The knowledge graph-based XAI model evaluation device includes: The explanation result acquisition module is used to acquire the explanation result pairs corresponding to the XAI model to be evaluated. The interpretation result is specifically used by the acquisition module for: A user data point is randomly selected from the recommendation sample library. The user data includes the user's past purchase records and browsing but not purchasing records. The set of past purchase records and the set of browsing but not purchasing records are used as input data and simultaneously fed into two XAI models to be evaluated. Each XAI model infers and determines the products to recommend to the user based on the input data. Each XAI model selects a factor from the input data to explain the rationality of the recommendation, and generates a corresponding atomic explanation result based on the factor and the recommended product. The atomic explanation results generated by each model are combined into a corresponding explanation result sequence, thereby obtaining an explanation result pair consisting of two explanation result sequences. The XAI model to be evaluated is used to evaluate the path explanation in the recommendation system, which infers the product that the user wants to buy based on the product information that the user has purchased and browsed. The node sequence pair acquisition module is used to acquire node sequence pairs in the knowledge graph that match the explanation result pair; The node sequence pair acquisition module is specifically used to acquire interpretation result pairs, match the interpretation result pairs with the node sequence and attribute sequence of KG, and obtain node sequence pairs that match the interpretation result pairs after KG semantic retrieval matching processing. The subgraph set determination module is used to determine the subgraph set corresponding to each node sequence based on the node sequence pairs and the knowledge graph. The rating pair determination module is used to determine the rating pair corresponding to the interpretation result pair based on the subgraph set corresponding to each node sequence. The scoring and determination module is specifically used for: The explanation coherence score, explanation complexity score, and explanation credibility score are determined based on the subgraph set corresponding to each node sequence. The target score for each node sequence is determined based on the explanation coherence score, explanation complexity score, and explanation credibility score corresponding to each node sequence. Determine the corresponding score pair for the interpretation result based on the target score corresponding to each node sequence; Based on the scoring pairs, a sequence of explanatory power corresponding to each XAI model to be evaluated is determined, and the XAI model with the strongest explanatory power is selected for use based on multiple explanatory power sequences.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the knowledge graph-based XAI model evaluation method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the knowledge graph-based XAI model evaluation method according to any one of claims 1-6.