Data processing method and related apparatus
By constructing an interpretable logical model and using SHAP and integral gradient algorithms, the contribution of changes in financial indicators is quantified, solving the problem of quantitative analysis of complex financial indicators across periods and improving the reliability of the analysis.
Patent Information
- Application Number
- PCT/CN2025/087427
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-30
- Filing Date
- 2025-04-07
- Publication Date
- 2025-12-04
AI Technical Summary
Existing technologies struggle to quantitatively analyze the drivers of changes in complex financial indicators across periods, particularly failing to separate the impact of year-on-year changes in contribution profit on contribution profit.
By employing pre-defined interpretable algorithms such as SHAP and integral gradient algorithm, an interpretable logical model is constructed based on the logical relationship between multiple data indicators. This model quantifies the contribution of each data indicator to the changes in financial indicator values. Through consistency verification of contribution values and processing of interaction effects, the reliability of attribution results is improved.
It enables quantitative analysis of changes in financial indicators over time, improving the quantification of the causes of changes in data indicator values and the reliability of the attribution results.
Smart Images

Figure CN2025087427_04122025_PF_FP_ABST
Abstract
Description
A data processing method and related apparatus
[0001] The present application claims priority from the Chinese patent application No. 202410694074.3 filed on May 30, 2024, and entitled "A data processing method and related apparatus", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] The present application relates to the field of data analysis, and in particular, to a data processing method and related apparatus. BACKGROUND
[0003] Currently, the internal financial management of large enterprises presents the characteristics of rich index definition, multiple dimensions, and complex reconciliation relationships, which brings great difficulty to the cause analysis of the difference (GAP) of index cross-period changes.
[0004] The existing solutions mainly decompose the GAP value of cross-period changes to some key business elements based on the experience of analysts in different fields, for example, an analyst can attribute the GAP value of company revenue to relevant factors such as shipment volume and unit price according to business experience.
[0005] However, the existing analysis method can only analyze the cause of GAP based on artificial experience, and when the elements affecting the financial index are complex, it is unable to perform quantitative analysis on complex elements. For example, since the contribution profit is affected by multiple factors such as income, cost, and expense in the same period, it is unable to separate the influence of the same-period change of the R&D expense allocation rate on the contribution profit. SUMMARY
[0006] Embodiments of the present application provide a data processing method and related apparatus, which can use the contribution value of a second data index in an interpretable logic model obtained by using a preset interpretable algorithm to quantify the reason for the change in the value of a first data index.
[0007] Therefore, in a first aspect, the present application provides a data processing method, in which a plurality of data indicators have a logical relationship, and a plurality of second data indicators will affect the change of a first data indicator value. First, a plurality of data indicators can be obtained, including one or more first data indicators and a plurality of second data indicators, the second data indicators being influencing factors affecting the change of the first value, and the first value being the value of the first data indicator. Then, an interpretable logical model can be obtained according to the logical relationship between the first data indicator and the plurality of second data indicators, the interpretable logical model being used to obtain the first value according to a plurality of second values, the second values being the values of the second data indicators. According to the first value and the plurality of second values, a preset interpretable algorithm is used to calculate the contribution value of the second data indicator in the interpretable logical model, which can be used to quantify the influence degree of the second data indicator on the change of the value of the first data indicator.
[0008] In the embodiments of the present application, the interpretable algorithm can generally be used to explain the contribution of each feature in the machine learning model to the prediction result, so that the contribution values of the plurality of second data indicators in the interpretable logical model can be calculated based on the interpretable algorithm, and the influence degree of the second data indicators on the change of the value of the first data indicator can be quantified by the contribution values, so as to realize the quantitative analysis of the reasons for the change of the data indicator value.
[0009] In a possible implementation, the foregoing operation of calculating the contribution value of the second data indicator in the interpretable logical model according to the first value and the plurality of second values by using the preset interpretable algorithm can include: calculating a first contribution value of the second data indicator according to the first value and the plurality of second values by using a SHAP (SHapley Additive exPlanations) interpretable algorithm, the first contribution value being obtained according to the shapley value of the second data indicator; and calculating a second contribution value of the second data indicator according to the first value and the plurality of second values by using an integral gradient algorithm, the second contribution value being obtained according to the integral gradient value of the second data indicator.
[0010] In the embodiments of the present application, the plurality of contribution values of the plurality of second data indicators can be calculated based on the SHAP algorithm and the integral gradient algorithm respectively, so that the reasons for the change of the data indicator value can be quantitatively analyzed according to the contribution values calculated by different algorithms.
[0011] In a possible implementation, after the foregoing operation of calculating the contribution value of the second data indicator in the interpretable logical model according to the first value and the plurality of second values by using the interpretable algorithm, the method can further include: determining an attribution result of the change of the value of the first data indicator according to the plurality of first contribution values and the plurality of second contribution values, the attribution result including the plurality of contribution values of the plurality of second data indicators.
[0012] In the embodiments of the present application, the consistency of the contribution values obtained by different explainable algorithms can be compared, and the attribution result of the numerical change of the data indicators is determined, so as to improve the consistency and reliability of the obtained attribution result.
[0013] In a possible implementation, the aforementioned determining the attribution result of the numerical change of the first data indicators according to the plurality of first contribution values and the plurality of second contribution values can include: when the first contribution value and the second contribution value of each second data indicator are consistent, taking the plurality of first contribution values of the plurality of second data indicators or the plurality of second contribution values of the plurality of second data indicators as the attribution result; when there is a second data indicator whose first contribution value and second contribution value are inconsistent, determining the attribution result according to a plurality of data indicator combinations, each data indicator combination in the plurality of data indicator combinations including a plurality of third data indicators, the plurality of third data indicators being obtained by combining the plurality of second data indicators.
[0014] In the embodiments of the present application, when the contribution values of the data indicators obtained by different explainable algorithms are consistent, the contribution values can be used to quantitatively analyze the reasons for the numerical change of the data indicators, and when the contribution values are inconsistent, the data indicators can be processed for interaction effect, and the data indicators are recombined, so that the attribution result can be determined according to the new data indicator combination.
[0015] In a possible implementation, the aforementioned determining the attribution result according to the plurality of data indicator combinations can include: using the SHAP algorithm to obtain a plurality of third contribution values of each data indicator combination in the plurality of data indicator combinations, the third contribution value being obtained according to the shapley value of the third data indicator; using the integral gradient algorithm to obtain a plurality of fourth contribution values of each data indicator combination in the plurality of data indicator combinations, the fourth contribution value being obtained according to the integral gradient value of the third data indicator; determining a target data indicator combination from the plurality of data indicator combinations according to the plurality of third contribution values and the plurality of fourth contribution values, the third contribution value and the fourth contribution value of each data indicator in the target data indicator combination being equal and each data indicator being independent of each other; determining the attribution result according to the target data indicator combination.
[0016] In the embodiments of the present application, when the contribution values obtained by different explainable algorithms are inconsistent, the data indicators can be recombined, and a target data indicator combination with consistent third contribution values and fourth contribution values is selected, so that the attribution result can be determined according to the contribution value of each data indicator in the target data indicator combination subsequently, thereby improving the reliability of the obtained attribution result.
[0017] In a possible implementation, the aforementioned determining the attribution result according to the target data indicator combination can include: determining the attribution result according to the third contribution value or the fourth contribution value of the plurality of target data indicators in the target data indicator combination.
[0018] In a second aspect, the present application provides a data processing apparatus, comprising:
[0019] an index logic expansion module, configured to obtain a plurality of data indexes, the plurality of data indexes comprising a first data index and a plurality of second data indexes, the plurality of second data indexes being influence factors affecting a change of a first value of the first data index;
[0020] an index logic encapsulation module, configured to obtain an interpretable logic model according to a logical relationship between the first data index and the plurality of second data indexes, the interpretable logic model being used to obtain the first value according to a plurality of second values, the second values being values of the second data indexes;
[0021] an index interpretable analysis module, configured to obtain a contribution value of the second data index in the interpretable logic model according to the first value and the plurality of second values by using a preset interpretable algorithm, the contribution value being used to quantify an influence degree of the second data index on the value change of the first data index.
[0022] In a possible implementation, the index interpretable analysis module is specifically configured to: obtain a first contribution value of the second data index according to the first value and the plurality of second values by using a SHAP (SHapley Additive exPlanations) interpretable algorithm, the first contribution value being obtained according to a SHAP value of the second data index; and obtain a second contribution value of the second data index according to the first value and the plurality of second values by using an integral gradient algorithm, the second contribution value being obtained according to an integral gradient value of the second data index.
[0023] In a possible implementation, after the index interpretable analysis module obtains the plurality of contribution values of the plurality of second data indexes in the interpretable logic model by using the interpretable algorithm, the data processing apparatus further comprises an index analysis result processing and interaction effect decomposition module, configured to determine an attribution result of the value change of the first data index according to the plurality of first contribution values and the plurality of second contribution values, the attribution result comprising the plurality of contribution values of the plurality of second data indexes.
[0024] In a possible implementation, the index analysis result processing and interaction effect decomposition module is specifically configured to: when the first contribution value of each second data index is consistent with the second contribution value, take the plurality of first contribution values of the plurality of second data indexes or the plurality of second contribution values of the plurality of second data indexes as the attribution result; and when the first contribution value of the second data index is inconsistent with the second contribution value, determine the attribution result according to a plurality of data index combinations, each data index combination in the plurality of data index combinations comprising a plurality of third data indexes, the plurality of third data indexes being obtained by combining the plurality of second data indexes.
[0025] In a possible implementation, the index analysis result processing and interaction effect decomposition module is specifically configured to: obtain, by using the SHAP algorithm, a plurality of third contribution values of each data index combination in the plurality of data index combinations, the third contribution value being obtained according to a shapley value of the third data index; obtain, by using the integral gradient algorithm, a plurality of fourth contribution values of each data index combination in the plurality of data index combinations, the fourth contribution value being obtained according to an integral gradient value of the third data index; determine, according to the plurality of third contribution values and the plurality of fourth contribution values, a target data index combination from the plurality of data index combinations, the third contribution value and the fourth contribution value of each data index in the target data index combination being equal and each data index being independent of each other; and determine the attribution result according to the target data index combination.
[0026] In a possible implementation, the index analysis result processing and interaction effect decomposition module is specifically configured to: determine the attribution result according to the third contribution value or the fourth contribution value of each target data index in the target data index combination.
[0027] In a third aspect, an embodiment of the present application provides a computing device, including a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device executes the method in the first aspect or any possible implementation manner of the first aspect.
[0028] In a fourth aspect, an embodiment of the present application provides a computing device cluster, including at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method in the first aspect or any possible implementation manner of the first aspect.
[0029] In a fifth aspect, an embodiment of the present application provides a computer readable storage medium, including computer program instructions, when the computer program instructions are executed by a computing device cluster, the computing device cluster executes the method in the first aspect or any possible implementation manner of the first aspect.
[0030] In a sixth aspect, an embodiment of the present application provides a computer program product including instructions, when the instructions are run by a computing device cluster, the computing device cluster executes the method in the first aspect or any possible implementation manner of the first aspect.
[0031] The technical effects brought by the second aspect to the sixth aspect or any possible implementation manner thereof can refer to the technical effects brought by the first aspect or the related possible implementation manner of the first aspect, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS
[0032] Fig. 1 is a schematic diagram of a system framework according to an embodiment of the present application;
[0033] Fig. 2 is a schematic diagram of an apparatus architecture for root cause analysis of a multi-dimensional, multi-level data indicator GAP according to an embodiment of the present application;
[0034] Fig. 3 is a schematic diagram of a data processing method according to an embodiment of the present application;
[0035] Fig. 4 is a schematic diagram of a data processing apparatus according to an embodiment of the present application;
[0036] Fig. 5 is a schematic diagram of a computing device according to an embodiment of the present application;
[0037] Fig. 6 is a schematic diagram of a computing device cluster according to an embodiment of the present application;
[0038] Fig. 7 is a schematic diagram of another computing device cluster according to an embodiment of the present application. DETAILED DESCRIPTION
[0039] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of the present application.
[0040] In the present application, "at least one" means one or more, and "multiple" means two or more. "And / or" describes the association between the associated objects, which means that there can be three relationships, for example, A and / or B, which means that A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. The terms "first", "second", etc. in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged under appropriate circumstances, which is only a distinguishing way used in the description of the embodiments of the present application to describe the objects with the same properties. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that the process, method, system, product or device containing a series of units does not have to be limited to those units, but can include other units not clearly listed or inherent to these processes, methods, products or devices.
[0041] First, some concepts related to the embodiments of the present application are introduced.
[0042] 1. SHAP algorithm
[0043] The SHAP (shapley additive explanations) algorithm is a method based on the shapley value in game theory to explain the prediction results of machine learning models. Specifically, the SHAP algorithm is a post-hoc explanation framework that can provide a shapley value for each feature, indicating the contribution of that feature to the model prediction.
[0044] 2. Integrated Gradients algorithm
[0045] The Integrated Gradients algorithm is a method used in machine learning to explain the prediction results of deep neural networks. It is based on the concept of gradient, which is used to quantify the contribution of input features to the model output. Specifically, the Integrated Gradients algorithm calculates the integral of the gradient path from the reference input (or baseline input) to the target input to evaluate the impact of each feature on the model output.
[0046] In many application scenarios, it is necessary to analyze the changes in the numerical values of data indicators to find the reasons for the changes in the numerical values of data indicators. In the financial field, with more and more enterprises internal financial management showing characteristics such as multi-dimensional financial indicators and complex reconciliation relationships, the analysis of the causes of the GAP between financial indicators across time periods has become more and more complex.
[0047] Currently, the analysis of the GAP between financial indicators across time periods is mainly based on the experience of analysts in different fields, and quantitative analysis cannot be performed for complex financial indicators with multiple dimensions and multiple levels.
[0048] Based on this, the present application proposes a data analysis method, which can obtain an interpretable logic model according to the logical relationship between a first data indicator and a plurality of second data indicators in a plurality of data indicators, and obtain a plurality of contribution values of the plurality of second data indicators in the interpretable logic model by using a preset interpretable algorithm, to quantify and analyze the numerical value change of the first data indicator by using the contribution value.
[0049] The data processing method provided by the present application can be applied to root cause analysis of changes in the numerical values of data indicators, for example, root cause analysis of financial indicators when a GAP occurs across time periods.
[0050] The system architecture provided by the embodiments of the present application is introduced below.
[0051] Referring to FIG. 1, a system architecture 100 provided by the present application is shown. As shown in FIG. 1, the system architecture can include a client 110 and a computing device cluster 120 including at least one computing device. Any of the computing devices can be a terminal device such as a desktop computer, a notebook computer or other terminal device, or a server, a cloud server or the like. When the computing device cluster includes multiple computing devices, the types of the different computing devices can be the same or different.
[0052] Specifically, any of the computing devices in the computing device cluster 120 can obtain data sent by the client 110, for example, data of multiple financial indicators in different periods, and implement the data processing method provided by the present application through an installed application program or plug-in or the like.
[0053] It is worth noting that the system architecture shown in FIG. 1 is only an example and is not intended to limit the specific implementation to this example. For example, in other possible system architectures, the system architecture 100 can not include the client 110 and directly obtain data from a data storage system by any of the computing devices in the computing device cluster 120.
[0054] Specifically, the device for root cause analysis of GAP of multi-dimensional and multi-level data indicators in the embodiments of the present application is shown in FIG. 2. The root cause analysis device 200 can include an indicator logic expansion module 201, an indicator logic optimization module 202, an indicator logic packaging module 203, an indicator data correlation module 204, an indicator explainable analysis module 205 and an indicator analysis result processing and interaction effect decomposition module 206.
[0055] The functions of each module include:
[0056] The indicator logic expansion module 201 is configured to obtain multiple data indicators having complex logical relationships therebetween, which can be in a tree structure, and obtain an indicator calculation formula between the multiple data indicators according to a parent-child node logical relationship therebetween, for example, obtaining data indicators A, B and C, and obtaining an expression of the indicator calculation formula A = B + C according to a logical relationship therebetween.
[0057] The indicator logic optimization module 202 is configured to simplify the indicator calculation formula obtained by the foregoing indicator logic expansion module 201, for example, B = D / E, and the foregoing indicator calculation formula can be simplified as A = D / E + C.
[0058] The indicator logic packaging module 203 is configured to package the foregoing obtained indicator calculation formula into an explainable logic model.
[0059] The index data association module 204 is configured to associate the data indexes obtained by the index logic expansion module 201 to obtain a logical relationship between the data indexes.
[0060] The index explainable analysis module 205 is configured to calculate the contribution values of the data indexes in the explainable logic model by using different preset explainable algorithms.
[0061] The index analysis result processing and interaction effect decomposition module 206 is configured to analyze the calculated contribution values to obtain an attribution result, and perform interaction effect processing, i.e., decomposing and checking consistency for different attribution result differences caused by non-independent features in the indexes.
[0062] The method provided by the present application will be described below in combination with the foregoing system architecture.
[0063] Referring to FIG. 3, the present application provides a flowchart of a data processing method, which is described as follows.
[0064] 301, obtaining a plurality of data indexes;
[0065] The data indexes are generally used to measure, evaluate and analyze a certain situation or phenomenon, such as financial indexes, sales indexes, operation indexes, etc. The financial indexes can include profit margin, asset-liability ratio, cash flow, etc. The sales indexes can include sales, sales growth rate, etc. The operation indexes can include production efficiency, work efficiency, product yield, etc. The specific embodiments are not limited herein. First, a plurality of data indexes can be obtained, which include one or more first data indexes and a plurality of second data indexes. The plurality of second data indexes are influence factors that affect the value change of the first data indexes.
[0066] Exemplarily, the plurality of data indexes can be a plurality of financial indexes, which include A product revenue, unit price and shipment volume of a and b sub-modules, turnover rate, adjustment rate and module revenue proportion, etc. The A product revenue is the first data index, and the unit price and shipment volume of the a and b sub-modules, the turnover rate, the adjustment rate and the module revenue proportion are the plurality of second data indexes. The unit price and shipment volume of the a and b sub-modules, the turnover rate, the adjustment rate and the module revenue proportion jointly affect the value change of the A product revenue.
[0067] In addition, under multi-dimensional cross management, the indexes are usually multi-dimensional composite indexes. For example, the A product revenue described above can be the A product revenue of A region, B industry and C company. In this case, the multi-dimensional composite indexes need to be analyzed as a whole.
[0068] 302, obtaining an explainable logic model according to the logical relationship between the first data index and the plurality of second data indexes;
[0069] In the embodiments of the present application, the obtained multiple data indicators are generally multi-dimensional and multi-level data indicators, and the multiple data indicators have complex logical relationships and a tree structure. Therefore, the index calculation formula of the first data indicator and the multiple second data indicators can be obtained according to the logical relationships between the multiple data indicators, and then an interpretable logical model is obtained, which includes the logical relationships between the first data indicator and the multiple second data indicators. Moreover, the interpretable logical model can obtain the value corresponding to the first data indicator according to the values of the multiple second data indicators.
[0070] For example, according to the logical relationship between the A product revenue and the unit price, the delivery quantity, the transfer rate, the adjustment rate and the module revenue proportion of the a and b sub-modules, the expression of the obtained index calculation formula can be: A product revenue = (sub-module a unit price * sub-module a delivery quantity + sub-module b unit price * sub-module b delivery quantity) * transfer rate * (1-adjustment rate) / (1-module revenue proportion). In addition, for a multi-dimensional composite indicator, the multi-dimensional composite indicator can be regarded as a whole to obtain a composite index calculation formula, such as [A product revenue of A region, B industry and C company] = sub-module a unit price * sub-module a delivery quantity + sub-module b unit price * sub-module b delivery quantity) * transfer rate * (1-adjustment rate) / (1-module revenue proportion). For another example, multiple data indicators A, B and C are obtained, wherein B and C are influence factors that cause the value of indicator A to change, and according to the logical relationship between A, B and C, the expression can be A = B 2 +C. In addition, if indicator B is influenced by lower-level indicators D and E, and satisfies the formula B = D / E, the previously obtained index calculation formula should be refined, and the expression of the new index calculation formula can be A = (D / E) 2 +C.
[0071] After obtaining the index calculation formula of the multiple data indicators, the index calculation formula can be used as a function of the interpretable logical model, so as to obtain the interpretable logical model.
[0072] 303、adopting a preset interpretable algorithm to obtain the contribution value of the second data indicator in the interpretable logical model.
[0073] The interpretable algorithm is usually applied to the field of machine learning to explain the prediction result of the machine learning model. The interpretable algorithm includes a SHAP algorithm and an integral gradient algorithm, etc. The SHAP algorithm and the integral gradient algorithm can be adopted to respectively calculate the shapley value and the integral gradient value of each second data index in the interpretable logical model, and the contribution value of each second data index can be determined according to the shapley value and the integral gradient value of the second data index. The contribution value is used to quantify the influence degree of the second data index on the numerical change of the first data index.
[0074] Specifically, based on the numerical value (first numerical value) of the first data index and the numerical value (second numerical value) of the plurality of second data indexes, the SHAP algorithm and the integral gradient algorithm can be used to calculate the shapley value and the integral gradient value of each second data index in the model, respectively. The contribution value of each second data index can be calculated according to the shapley value or the integral gradient value of each second data index. The contribution value of each second data index can represent the contribution of each second data index to the difference of the numerical value (output value of the model) of the first data index. The sum of the contribution values of each second data index is equal to the difference between the actual value and the predicted value of the model output. For example, the average predicted price of an apartment is 310,000 yuan, and the actual predicted price of a certain apartment (with an area of 50 square meters, located on the 2nd floor, with a park nearby, and with a cat prohibition) is 300,000 yuan. The difference between the average predicted price is 10,000 yuan. It is known that the characteristics affecting the price of the apartment include surrounding facilities, whether to allow cats, area, and floor. Based on the SHAP algorithm or the integral gradient algorithm, the contribution value of each characteristic can be calculated, for example, the contribution value of “near a park” is 30,000 yuan, the contribution value of “50 square meters” is 10,000 yuan, the contribution value of “located on the 2nd floor” is 0 yuan, and the contribution value of “prohibition of keeping cats” is -50,000 yuan. The total contribution value is -10,000 yuan, which is equal to the difference between the actual predicted price and the average predicted price of the apartment.
[0075] Optionally, the SHAP algorithm can be used to calculate a plurality of first contribution values of a plurality of second index data. The first contribution value is obtained according to the shapley value of the second data index. The integral gradient algorithm can also be used to calculate a plurality of second contribution values of the plurality of second index data. The second contribution value is obtained according to the integral gradient value of the second data index.
[0076] The shapley value of each second data index in the interpretable logical model can be calculated according to the following expression:
[0077] wherein φ j (val) is the shapley value of the jth second index data, S is the plurality of second index data excluding xj where x is a vector of the second data indicators, M is the number of the second data indicators, |S| is the number of the second data indicators in the subset S, val(S∪x j ) is the value (or benefit) created by the cooperation of the subset S and x j where val(S) is the value (or benefit) created by the cooperation of the subset S.
[0078] The expression of the output value of the explainable logic model based on a certain sample data (which includes the value of the first data indicator and the values of the plurality of second data indicators) can be:
[0079] wherein, can be regarded as a fixed term, is the shapley value of the indicator x (any second data indicator), is the shapley value of the indicator y, is the shapley value of the indicator z.
[0080] Therefore, the output value corresponding to a certain sample data can be the sum of the shapley values of the plurality of data indicators, and the difference between the values of the first data indicator in different sample data can be the sum of the first contribution values of the plurality of second data indicators, wherein the first contribution value of the second data indicator can be the difference between the shapley value based on the first sample data and the shapley value based on the second sample data, and the expression thereof can be:
[0081] wherein, is the first contribution value of the indicator x, is the first contribution value of the indicator y, is the first contribution value of the indicator z.
[0082] In addition, the integral gradient value of each second data indicator in the explainable logic model can be calculated by the expression:
[0083] At this time, the expression of the output value of the explainable logic model based on a certain sample data can be:
[0084] At this time, the difference between the values of the first data indicator in different sample data can be the sum of the second contribution values of the plurality of second data indicators, wherein the second contribution value of the second data indicator can be the difference between the integral gradient value based on the first sample data and the integral gradient value based on the second sample data, and the expression thereof can be:
[0085] wherein, is a second contribution value of the indicator x, is a second contribution value of the indicator y, is a second contribution value of the indicator z.
[0086] Optionally, the attribution result of the numerical change of the first data indicator can be determined according to the plurality of first contribution values and the plurality of second contribution values, and the attribution result includes a plurality of contribution values of the plurality of second data indicators.
[0087] In the embodiments of the present application, the SHAP algorithm and the integral gradient algorithm can be used to calculate the plurality of contribution values of the plurality of second data indicators, so that the consistency of the contribution values calculated by different algorithms can be compared subsequently, the attribution result of the numerical change of the data indicator is determined, thereby improving the consistency and reliability of the obtained attribution result.
[0088] Specifically, after obtaining the plurality of first contribution values and the plurality of second contribution values, it can be judged whether the first contribution value and the second contribution value corresponding to each second data indicator are consistent, if the first contribution value and the second contribution value of each second data indicator are consistent, the plurality of first contribution values or the plurality of second contribution values of the plurality of second data indicators can be taken as the attribution result. If the first contribution value and the second contribution value of the second data indicator are inconsistent, the plurality of second data indicators can be recombined to obtain a plurality of data indicator combinations, each data indicator combination includes a plurality of third data indicators, and then the attribution result can be determined according to the plurality of data indicator combinations.
[0089] Alternatively, when the first contribution value and the second contribution value of the second data indicator are inconsistent, the plurality of first contribution values obtained based on the SHAP algorithm can also be directly selected as the contribution values of the plurality of second data indicators affecting the numerical change of the first data indicator after expert analysis and determination.
[0090] In the embodiments of the present application, since the integral gradient algorithm is used to find the integral path to obtain the corresponding integral gradient value, when the data indicator is not suitable for the straight integral path for solving, the integral gradient algorithm will cause the contribution of the data indicator to be unfairly distributed, therefore, when the first contribution value and the second contribution value of the second data indicator are inconsistent, the first contribution value obtained by the SHAP algorithm can be selected as the contribution value of the data indicator.
[0091] When the first contribution value and the second contribution value of the second data indicator are inconsistent, the SHAP algorithm and the integral gradient algorithm can be used to obtain a plurality of shapley values and a plurality of integral gradient values of the plurality of data indicator combinations respectively, and a plurality of third contribution values are obtained according to the plurality of shapley values, and a plurality of fourth contribution values are obtained according to the plurality of integral gradient values. Further, by judging whether the plurality of third contribution values and the plurality of fourth contribution values of the plurality of data indicators are consistent, a target data indicator combination is determined from the plurality of data indicator combinations, and the attribution result is determined according to the target data indicator combination. If there is no target data indicator combination in the plurality of data indicator combinations in which the third contribution value and the fourth contribution value are consistent, the plurality of first contribution values of the plurality of second data indicators obtained by using the SHAP algorithm can be selected as the contribution values of the plurality of second data indicators affecting the numerical change of the first data indicator.
[0092] Specifically, the SHAP algorithm can be used to obtain a plurality of third contribution values of each data indicator combination in the plurality of data indicator combinations, and the third contribution value is obtained according to the shapley value of the third data indicator, that is, the third contribution value can be the difference between the shapley value of the third data indicator based on the first data sample and the shapley value based on the second data sample. The integral gradient algorithm is used to obtain a plurality of fourth contribution values of each data indicator combination in the plurality of data indicator combinations, and the fourth contribution value is obtained according to the integral gradient value of the third data indicator, that is, the fourth contribution value can be the difference between the integral gradient value of the third data indicator based on the first data sample and the integral gradient value based on the second data sample. The third contribution value and the fourth contribution value of the third data indicator in each data indicator combination are compared, and a target data indicator combination is determined from the plurality of data indicator combinations, and the shapley value and the integral gradient value of each data indicator in the target data indicator combination are equal and each data indicator is independent.
[0093] After determining the target data indicator combination, the attribution result affecting the numerical change of the first data indicator can be determined according to the third contribution value and the fourth contribution value of the plurality of target data indicators in the target data combination. Specifically, the contribution value of the plurality of second data indicators to the first data indicator can be determined according to the third contribution value or the fourth contribution value of the plurality of target data indicators, and then the attribution result affecting the numerical change of the first data indicator is obtained.
[0094] For example, a plurality of data indicators X, Y and Z are obtained, wherein Z is the first data indicator, and X and Y are the second data indicators. According to the logical relationship between X, Y and Z data indicators, the function expression of the interpretable logical model is Z = f(X, Y) = X + 2XY 2+3Y+1, at this time, X in sample 0 is 3, Y is 4, X in sample 1 is 6, and Y is 10. First, the first contribution value and the second contribution value of X and the first contribution value and the second contribution value of Y are respectively calculated based on the SHAP algorithm and the integral gradient algorithm, when the first contribution value and the second contribution value of X are inconsistent, and the first contribution value and the second contribution value of Y are inconsistent, the data index combination 1 includes the data indexes X, Y and XY, the data index combination 2 includes the data indexes X, Y and Y 2 , the data index combination 3 includes the data indexes X, Y and XY 2 . Subsequently, the shapley value and the integral gradient value of each data index in the three data index combinations can be calculated based on sample 0 and sample 1 respectively, and the third contribution value and the fourth contribution value corresponding to each data index are calculated according to the shapley value and the integral gradient value of each data index based on different sample data, wherein the third contribution value represents the contribution of X to the difference of f(X, Y) based on different sample data, and the fourth contribution value represents the contribution of Y to the difference of f(X, Y) based on different sample data. The specific calculation results are as follows.
[0095] The third contribution value and the fourth contribution value of each third data index in the data index combination 1 are as follows: The third contribution value and the fourth contribution value of each third data index in the data index combination 2 are as follows: The third contribution value and the fourth contribution value of each third data index in the data index combination 3 are as follows:
[0096] According to the calculation results, the third contribution value and the fourth contribution value of each data index in the three data index combinations are consistent, and since the SHAP algorithm has the feature independence requirement, the data indexes in the data index combination 2 in the multiple data index combinations are independent of each other, and thus the data index combination 2 can be used as the target data index combination. According to the contribution values of the target data indexes X, Y and Y 2 in the data index combination 2, the contribution values of the indexes X and Y are obtained, as shown in Table 1.
[0097] Table 1
[0098] As shown in Table 1, the difference between the output values (the first data index value, i.e., the first value) corresponding to sample 0 and sample 1 is Δ = f(X1, Y1) - f(X0, Y0) = 1237 - 112 = 1125, and according to the contribution value of the third data index in the data index combination 2, the contribution value of the index X to the difference between the output values of sample 0 and sample 1 is 351, and the contribution value of the index Y to the difference between the output values of sample 0 and sample 1 is 774.
[0099] In the embodiments of the present application, the contribution values of the plurality of second data indexes in the interpretable logic model can be calculated by using an interpretable algorithm, and the contribution values are used to quantify the influence degree of the plurality of second data indexes on the value change of the first data index, so as to realize the quantitative analysis of the value change reason of the data index. In addition, the contribution values of the data indexes can be calculated by using two different interpretable algorithms, and the attribution result is determined according to the consistency of the contribution values, so as to improve the consistency and reliability of the attribution result.
[0100] The foregoing describes the method flow provided by the present application, and the following describes the device provided by the present application based on the foregoing method flow.
[0101] Referring to FIG. 4, the structure of a data processing device provided by the present application is shown as follows.
[0102] The index logic expansion module 201 is configured to obtain a plurality of data indexes, wherein the plurality of data indexes include a first data index and a plurality of second data indexes, the plurality of second data indexes are influence factors that cause the first value to change, and the first value is the value of the first data index.
[0103] The index logic encapsulation module 203 is configured to obtain an interpretable logic model according to the logical relationship between the first data index and the plurality of second data indexes, and the interpretable logic model is used to obtain the first value according to a plurality of second values, wherein the second values are the values of the second data indexes.
[0104] The index interpretable analysis module 205 is configured to obtain the contribution values of the second data indexes in the interpretable logic model by using a preset interpretable algorithm according to the first value and the plurality of second values, and the contribution values are used to quantify the influence degree of the second data indexes on the value change of the first data index.
[0105] In a possible implementation, the index explainable analysis module 205 is specifically configured to: according to the first value and the plurality of second values, obtain a first contribution value of the second data index by using a SHAP algorithm of a Shapley Additive exPlanation (SHAP) model, the first contribution value being obtained according to a shapley value of the second data index; and according to the first value and the plurality of second values, obtain a second contribution value of the second data index by using an integrated gradient algorithm, the second contribution value being obtained according to an integrated gradient value of the second data index.
[0106] In a possible implementation, after the index explainable analysis module 205 obtains the plurality of contribution values of the plurality of second data indexes in the explainable logic model by using the explainable algorithm, the data processing apparatus further includes an index analysis result processing and interaction effect decomposition module 206, configured to determine an attribution result of a value change of the first data index according to the plurality of first contribution values and the plurality of second contribution values, the attribution result including the plurality of contribution values of the plurality of second data indexes.
[0107] In a possible implementation, the index analysis result processing and interaction effect decomposition module 206 is specifically configured to: when the first contribution value and the second contribution value of each second data index are consistent, take the plurality of first contribution values of the plurality of second data indexes or the plurality of second contribution values of the plurality of second data indexes as the attribution result; and when the first contribution value and the second contribution value of the second data index are inconsistent, determine the attribution result according to a plurality of data index combinations, each data index combination in the plurality of data index combinations including a plurality of third data indexes, the plurality of third data indexes being obtained by combining the plurality of second data indexes.
[0108] In a possible implementation, the index analysis result processing and interaction effect decomposition module 206 is specifically configured to: obtain, by using the SHAP algorithm, a plurality of third contribution values of each data index combination in the plurality of data index combinations, the third contribution value being obtained according to a shapley value of the third data index; obtain, by using the integrated gradient algorithm, a plurality of fourth contribution values of each data index combination in the plurality of data index combinations, the fourth contribution value being obtained according to an integrated gradient value of the third data index; determine, according to the plurality of third contribution values and the plurality of fourth contribution values, a target data index combination from the plurality of data index combinations, each data index in the target data index combination having equal third contribution values and fourth contribution values and each data index being independent of each other; and determine the attribution result according to the target data index combination.
[0109] In a possible implementation, the index analysis result processing and interaction effect decomposition module 206 is specifically configured to: determine the attribution result according to the third contribution value or the fourth contribution value of each target data index in the target data index combination.
[0110] The index logic unfolding module, the index logic packaging module, the index explainable analysis module, and the index analysis result processing and interaction effect decomposition module can be implemented by software or by hardware. For example, the implementation of the index logic unfolding module is described below. Similarly, the implementation of the index logic packaging module, the index explainable analysis module, and the index analysis result processing and interaction effect decomposition module can be implemented by referring to the implementation of the index logic unfolding module.
[0111] As an example of a software functional unit, the index logic unfolding module can include code running on a computing instance. The computing instance can include at least one of a physical host (computing device), a virtual machine, and a container. Further, the computing instance can be one or more. For example, the A module can include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers running the code can be distributed in the same region, or in different regions. Further, the multiple hosts / virtual machines / containers running the code can be distributed in the same availability zone (AZ), or in different AZs, each of which includes a data center or multiple data centers in close geographical proximity. Generally, one region can include multiple AZs.
[0112] Similarly, the multiple hosts / virtual machines / containers running the code can be distributed in the same virtual private cloud (VPC), or in multiple VPCs. Generally, one VPC is set in one region, and communication between two VPCs in the same region or between VPCs in different regions needs to be set in each VPC to realize the interconnection between VPCs through a communication gateway.
[0113] As an example of a hardware functional unit, the metric logic unfolding module can include at least one computing device, such as a server or the like. Alternatively, the metric logic unfolding module can also be a device implemented using a central processing unit (CPU), or implemented using an application-specific integrated circuit (ASIC), or implemented using a programmable logic device (PLD), and the like. The PLD can be implemented as a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), a data processing unit (DPU), a neural network processing unit (NPU), a system on chip (SoC), an offload card, an acceleration card, or any combination thereof.
[0114] The plurality of computing devices included in the metric logic unfolding module can be distributed in the same region or in different regions. The plurality of computing devices included in the metric logic unfolding module can be distributed in the same AZ or in different AZs. Similarly, the plurality of computing devices included in the metric logic unfolding module can be distributed in the same VPC or in multiple VPCs. The plurality of computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, DPUs, NPUs, SoCs, offload cards, acceleration cards, and the like.
[0115] It should be noted that in other embodiments, the metric logic unfolding module can be used to perform any step in the data processing method, the metric logic packaging module can be used to perform any step in the data processing method, the metric interpretable analysis module can be used to perform any step in the data processing method, the metric analysis result processing and interaction effect decomposition module can be used to perform any step in the data processing method, and the metric logic unfolding module, the metric logic packaging module, the metric interpretable analysis module, and the metric analysis result processing and interaction effect decomposition module are responsible for implementing the steps specified as needed by the data processing method. The metric logic unfolding module, the metric logic packaging module, the metric interpretable analysis module, and the metric analysis result processing and interaction effect decomposition module implement different steps in the data processing method to realize the overall function of the data processing device.
[0116] This application also provides a computing device 500. As shown in FIG5, the computing device 500 includes: a bus 502, a processor 504, a memory 506, and a communication interface 508. The processor 504, the memory 506, and the communication interface 508 communicate with each other via the bus 502. The computing device 500 may be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 500.
[0117] Bus 502 can be a Peripheral Component Interconnect Express (PCIe) bus, an Extended Industry Standard Architecture (EISA) bus, a Unified Bus (Ubus or UB), a Compute Express Link (CXL), a Cache Coherent Interconnect for Accelerators (CCIX), etc. The Unified Bus is also known as the Lingqu Bus. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, only one line is used in Figure 5, but this does not indicate that there is only one bus or one type of bus. Bus 504 can include pathways for transmitting information between various components of the computing device 500 (e.g., memory 506, processor 504, communication interface 508). The Unified Bus can also be referred to as the Lingqu Bus.
[0118] Processor 504 may include any one or more computing devices such as a central processing unit (CPU), graphics processing unit (GPU), microprocessor (MP) or digital signal processor (DSP), ASIC, FPGA, CPLD, NPU, SoC, offload card, accelerator card, etc.
[0119] Memory 506 may include volatile memory, such as random access memory (RAM). Processor 504 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD). Furthermore, memory 506 may also be implemented using storage class memory (SCM), phase change memory (PCM), or other types of storage media.
[0120] It is worth noting that the same type of storage medium can be configured in the same computing device to realize the function of memory 506, or two or more types of storage media can be configured to realize the function of memory 506. This application does not limit this.
[0121] The memory 506 stores executable program code, which the processor 504 executes to implement the functions of the aforementioned indicator logic expansion module, indicator logic encapsulation module, indicator interpretability analysis module, indicator analysis result processing module, and interaction effect decomposition module, thereby realizing the data processing method. In other words, the memory 506 stores instructions for executing the data processing method.
[0122] The communication interface 508 uses transceiver modules, such as, but not limited to, network interface cards and transceivers, to enable communication between the computing device 500 and other devices or communication networks.
[0123] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0124] As shown in Figure 6, the computing device cluster includes at least one computing device 500. The memory 506 of one or more computing devices 500 in the computing device cluster may store the same instructions for executing data processing methods.
[0125] In some possible implementations, the memory 506 of one or more computing devices 500 in the computing device cluster may also store partial instructions for executing data processing methods. In other words, a combination of one or more computing devices 500 can jointly execute instructions for executing data processing methods.
[0126] It should be noted that the memory 506 in different computing devices 500 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the data processing device. That is, the instructions stored in the memory 506 of different computing devices 500 can implement the functions of one or more modules among the indicator logic expansion module, indicator logic encapsulation module, indicator interpretability analysis module, indicator analysis result processing module, and interaction effect decomposition module.
[0127] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 7 illustrates one possible implementation. As shown in Figure 7, two computing devices 500A and 500B are connected via a network. Specifically, they are connected to the network through communication interfaces in each computing device. In this type of possible implementation, the memory 506 in computing device 500A stores instructions for executing the functions of the indicator logic expansion module. Simultaneously, the memory 506 in computing device 500B stores instructions for executing the functions of the indicator logic encapsulation module, the indicator interpretability analysis module, the indicator analysis result processing module, and the interaction effect decomposition module.
[0128] It should be understood that the functions of computing device 500A shown in Figure 7 can also be performed by multiple computing devices 500. Similarly, the functions of computing device 500B can also be performed by multiple computing devices 500.
[0129] This application embodiment also provides another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similarly referred to the connection method of the computing device cluster described in Figures 6 and 7. The difference is that the memory 506 of one or more computing devices 500 in this computing device cluster can store the same instructions for executing data processing methods.
[0130] In some possible implementations, the memory 506 of one or more computing devices 500 in the computing device cluster may also store partial instructions for executing data processing methods. In other words, a combination of one or more computing devices 500 can jointly execute instructions for executing data processing methods.
[0131] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform the method provided in this application.
[0132] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform the method provided in this application.
[0133] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.
Claims
1. A data processing method, characterized in that, include: Multiple data indicators are acquired, including a first data indicator and multiple second data indicators. The multiple second data indicators are influencing factors that affect the change of the first value, and the first value is the value of the first data indicator. An interpretable logical model is obtained based on the logical relationship between the first data indicator and the plurality of second data indicators. The interpretable logical model is used to obtain the first value based on the plurality of second values, where the second values are the values of the second data indicators. Based on the first value and the plurality of second values, a preset interpretable algorithm is used to calculate the contribution value of the second data indicator in the interpretable logic model. The contribution value is used to quantify the degree of influence of the second data indicator on the numerical change of the first data indicator.
2. The method according to claim 1, characterized in that, The step of calculating the contribution value of the second data indicator in the interpretable logic model based on the first value and the plurality of second values using a preset interpretable algorithm includes: Based on the first value and the plurality of second values, the first contribution value of the second data indicator is obtained by using the Shapley Additivity Model Interpretation Algorithm (SHAP). The first contribution value is obtained based on the Shapley value of the second data indicator. Based on the first value and the plurality of second values, the second contribution value of the second data index is obtained by using an integral gradient algorithm. The second contribution value is obtained based on the integral gradient value of the second data index.
3. The method according to claim 2, characterized in that, After calculating the contribution value of the second data indicator in the interpretable logic model using a preset interpretable algorithm based on the first value and the plurality of second values, the method further includes: Based on multiple first contribution values and multiple second contribution values, an attribution result for the numerical change of the first data indicator is determined, wherein the attribution result includes multiple contribution values of the multiple second data indicators.
4. The method according to claim 3, characterized in that, The step of determining the attribution result of the numerical change of the first data indicator based on multiple first contribution values and multiple second contribution values includes: When the first contribution value of each of the second data indicators is consistent with the second contribution value, the multiple first contribution values of the multiple second data indicators or the multiple second contribution values of the multiple second data indicators are used as the attribution result; When the first contribution value of the second data indicator is inconsistent with the second contribution value, the attribution result is determined based on a combination of multiple data indicators. Each combination of multiple data indicators includes multiple third data indicators, which are obtained by combining the multiple second data indicators.
5. The method according to claim 4, characterized in that, Determining the attribution result based on a combination of multiple data indicators includes: The SHAP algorithm is used to calculate multiple third contribution values for each of the multiple data indicator combinations, and the third contribution values are obtained based on the shapley value of the third data indicator. The integral gradient algorithm is used to calculate multiple fourth contribution values for each of the multiple data indicator combinations, and the fourth contribution values are obtained based on the integral gradient value of the third data indicator. Based on the plurality of third contribution values and the plurality of fourth contribution values, a target data indicator combination is determined from the plurality of data indicator combinations, wherein the third contribution value and the fourth contribution value of each data indicator in the target data indicator combination are equal and each data indicator is independent of the others; The attribution result is determined based on the combination of the target data indicators.
6. The method according to claim 5, characterized in that, Determining the attribution result based on the target data indicator combination includes: The attribution result is determined based on the third or fourth contribution value of multiple target data indicators in the target data indicator combination.
7. A data processing apparatus, characterized in that, include: The indicator logic expansion module is used to obtain multiple data indicators, including a first data indicator and multiple second data indicators. The multiple second data indicators are influencing factors that affect the change of the first value, and the first value is the value of the first data indicator. The indicator logic encapsulation module is used to obtain an interpretable logic model based on the logical relationship between the first data indicator and the plurality of second data indicators. The interpretable logic model is used to obtain the first value based on the plurality of second values, where the second values are the values of the second data indicators. The indicator interpretability analysis module is used to calculate the contribution value of the second data indicator in the interpretable logic model based on the first value and the plurality of second values using a preset interpretable algorithm. The contribution value is used to quantify the degree of influence of the second data indicator on the numerical change of the first data indicator.
8. The apparatus according to claim 7, characterized in that, The indicator interpretation and analysis module is specifically used for: Based on the first value and the plurality of second values, the first contribution value of the second data indicator is obtained by using the Shapley Additivity Model Interpretation Algorithm (SHAP). The first contribution value is obtained based on the Shapley value of the second data indicator. Based on the first value and the plurality of second values, the second contribution value of the second data index in the interpretable logic model is obtained by using the integral gradient algorithm. The second contribution value is obtained based on the integral gradient value of the second data index.
9. The apparatus according to claim 8, characterized in that, After obtaining the multiple contribution values of the multiple second data indicators in the interpretable logic model using an interpretable algorithm, the device further includes: The indicator analysis result processing and interaction effect decomposition module is used to determine the attribution result of the numerical change of the first data indicator based on multiple first contribution values and multiple second contribution values, wherein the attribution result includes multiple contribution values of the multiple second data indicators.
10. The apparatus according to claim 9, characterized in that, The module for processing the index analysis results and decomposing the interaction effect is specifically used for: When the first contribution value of each of the second data indicators is consistent with the second contribution value, the multiple first contribution values of the multiple second data indicators or the multiple second contribution values of the multiple second data indicators are used as the attribution result; When the first contribution value of the second data indicator is inconsistent with the second contribution value, the attribution result is determined based on a combination of multiple data indicators. Each combination of multiple data indicators includes multiple third data indicators, which are obtained by combining the multiple second data indicators.
11. The apparatus according to claim 10, characterized in that, The module for processing the index analysis results and decomposing the interaction effect is specifically used for: The SHAP algorithm is used to calculate multiple third contribution values for each of the multiple data indicator combinations, and the third contribution values are obtained based on the shapley value of the third data indicator. The integral gradient algorithm is used to calculate multiple fourth contribution values for each of the multiple data indicator combinations, and the fourth contribution values are obtained based on the integral gradient value of the third data indicator. Based on the plurality of third contribution values and the plurality of fourth contribution values, a target data indicator combination is determined from the plurality of data indicator combinations, wherein the third contribution value and the fourth contribution value of each data indicator in the target data indicator combination are equal and each data indicator is independent of the others; The attribution result is determined based on the combination of the target data indicators.
12. The apparatus according to claim 11, characterized in that, The module for processing the index analysis results and decomposing the interaction effect is specifically used for: The attribution result is determined based on the third or fourth contribution value of multiple target data indicators in the target data indicator combination.
13. A computing device, characterized in that, The computing device includes a processor and memory; The processor is configured to execute instructions stored in the memory to cause the computing device to perform the method as described in any one of claims 1 to 6.
14. A computing device cluster, characterized in that, It includes at least one computing device, said at least one computing device including a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the operational steps of the method as described in any one of claims 1 to 6.
15. A computer-readable storage medium, characterized in that, It includes computer program instructions, which, when executed by a cluster of computing devices, perform the operational steps of the method as described in any one of claims 1 to 6.
16. A computer program product containing instructions, characterized in that, When the instruction is executed by the computing device cluster, the computing device cluster causes the computing device cluster to perform the operation steps of the method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Risk control method, system and device and storage medium
CN116416073A
Numerical feature discretization attribution analysis method and device based on clustering binning
CN116738261A
Attribution method and device, electronic equipment and storage medium
CN117252450A
Model interpretation method and device, equipment and storage medium
CN117372179A
Systems and methods for risk factor predictive modeling with model explanations
US11983777B1