Model attenuation attribution method and device, electronic equipment and storage medium

By determining key features in the sample data and analyzing the change in probability density distribution between them and the application data, accurately locate the root cause of machine learning model performance attenuation, and solving the problem that model performance attenuation cannot be effectively solved in the prior art.

CN120104384APending Publication Date: 2025-06-06JINGDONG TECH HLDG CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311667694.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-06
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing technology is difficult to accurately determine the root cause of performance decay of machine learning models after they are launched, which makes it impossible to effectively solve the problem of model performance decay.

Method used

The root information of model attenuation is determined by determining the N key features with the greatest importance in the sample data and based on the change in the probability density distribution between these key features and similar features in the application data.

Benefits of technology

It can accurately and quickly determine the root cause of model attenuation, thereby effectively solving the problem of model performance attenuation after launch, and improving production efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104384A_ABST
    Figure CN120104384A_ABST
Patent Text Reader

Abstract

The invention provides a model attenuation attribution method and device, electronic equipment and a storage medium, and relates to the technical field of computers and the Internet. The method comprises the following steps: determining N key features with the maximum feature importance in a plurality of features included in sample data; determining root information of model attenuation based on probability density distribution change between each key feature and application features of the same kind in the application data; wherein the application data is data obtained after the model is online. In the embodiment of the invention, the source information of model attenuation is determined by analyzing the probability density distribution change between each key feature with the maximum feature importance and the application features in the application data, and the variables drifting can be accurately and quickly determined when the model is attenuated; therefore, the root cause (root information) of model attenuation can be accurately determined, and the problem of model attenuation after online can be effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer and Internet technology, and in particular to a model attenuation attribution method, device, electronic device and storage medium. Background Art

[0002] With the development of artificial intelligence (AI) technology, various machine learning and deep learning models are increasingly widely used in various production environments. However, most models will experience performance degradation, performance decay, or even complete failure after running online for a period of time. The more common solution at present is to re-collect and process the latest data, retrain and test the model, and then put it online. The above method does not fully analyze the root cause of model decay, but only retrains the model with updated data. The final effect may not completely improve the performance of the online model. It cannot solve the problem of performance decay of the model after going online. Summary of the invention

[0003] The embodiments of the present application provide a model attenuation attribution method, device, electronic device, and storage medium, which can accurately determine the root cause of model attenuation and effectively solve the problem of model performance attenuation after going online.

[0004] The technical solution of this application is implemented as follows:

[0005] The present application embodiment provides a model attenuation attribution method, including:

[0006] Determine N key features with the greatest feature importance among multiple features included in the sample data; wherein N is an integer greater than 0;

[0007] Based on the change in probability density distribution between each of the key features and similar application features in the application data, the root cause information of the model attenuation is determined; wherein the application data is the data obtained after the model is launched.

[0008] In the above scheme, the sample data includes: label sample data; the label sample data corresponds to N key label features; the N key features belong to the N features; the root cause information of the model attenuation is determined based on the probability density distribution change between each key feature and the same application features in the application data, including:

[0009] Determine a first probability density distribution change between each of the key features and the corresponding application feature in the application data;

[0010] Determine a second probability density distribution change between each of the key tag features and the corresponding application feature in the application data;

[0011] The root cause information is determined based on the correlation between the performance attenuation of the model and the first probability density distribution change corresponding to each of the key features and the second probability density distribution change corresponding to each of the key tag features.

[0012] In the above solution, the determining of a first probability density distribution change between each of the key features and the corresponding application feature in the application data includes:

[0013] Determine a first probability density distribution of each of the key features, and a second probability density distribution of the application features corresponding to each of the key features in the application data;

[0014] Determine a first relative entropy between a first probability density distribution corresponding to each of the key features and the corresponding second probability density distribution; wherein the first relative entropy is used to characterize a change in the first probability density distribution between each of the key features and the corresponding application feature.

[0015] In the above solution, the determining of the second probability density distribution change between each of the key tag features and the corresponding application feature in the application data includes:

[0016] Determine a third probability density distribution of each of the key tag features, and a fourth probability density distribution of the application features corresponding to each of the key tag features in the application data;

[0017] Determine a second relative entropy between the third probability density distribution corresponding to each of the key tag features and the corresponding fourth probability density distribution; wherein the second relative entropy is used to characterize the change in the second probability density distribution between each of the key tag features and the corresponding application features.

[0018] In the above scheme, the first probability density distribution change includes: a first relative entropy; the second probability density distribution change includes: a second relative entropy; the correlation between the performance attenuation based on the model and the first probability density distribution change corresponding to each of the key features and the second probability density distribution change corresponding to each of the key tag features, determining the root information includes:

[0019] determining a plurality of performance information of the model within a predetermined period of time;

[0020] determining performance degradation information of the model based on a plurality of the performance information;

[0021] The root cause information is determined based on the correlation between the performance degradation information and the first relative entropy corresponding to each of the key features, and the correlation between the performance degradation information and the second relative entropy corresponding to each of the key tag features.

[0022] In the above solution, determining the root information based on the correlation between the performance degradation information and the first relative entropy corresponding to each of the key features, and the correlation between the performance degradation information and the second relative entropy corresponding to each of the key tag features, includes:

[0023] Determine a first correlation coefficient between the performance degradation information and the first relative entropy corresponding to each of the key features;

[0024] Determine a second correlation coefficient between the performance decay information and the second relative entropy corresponding to each of the key tag features;

[0025] Based on the first correlation coefficient corresponding to each of the key features and the second correlation coefficient corresponding to each of the key label features, sorting each of the key features and each of the key label features to obtain a root feature set;

[0026] The root source information is determined based on the root source feature set.

[0027] In the above solution, the N key features with the greatest feature importance among the multiple features included in the sample data are determined, including:

[0028] Using a preset model, respectively calculate the importance of multiple features included in the sample data to obtain the importance of each feature;

[0029] It is determined that the features corresponding to the top N greatest importances are the N key features.

[0030] The embodiment of the present application also provides a model attenuation attribution device, including:

[0031] A determination unit, used to determine N key features with the greatest feature importance among the multiple features included in the sample data; wherein N is an integer greater than 0;

[0032] A determination unit is used to determine the root cause information of the model attenuation based on the change in probability density distribution between each of the key features in the N features and similar application features in the application data; wherein the application data is the data obtained after the model is launched.

[0033] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and the processor implements the steps in the above method when executing the computer program.

[0034] An embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method are implemented.

[0035] In an embodiment of the present application, N key features with the greatest feature importance are determined among the multiple features included in the sample data; wherein N is an integer greater than 0; based on the change in probability density distribution between each key feature and the same type of application features in the application data, the root cause information of model attenuation is determined; wherein the application data is the data acquired after the model is launched. In an embodiment of the present application, the root cause information of model attenuation is determined by analyzing the change in probability density distribution between each key feature with the greatest feature importance and the application features in the application data. When the model is attenuated, it is possible to accurately and quickly determine which variables have drifted, and then accurately determine the root cause (root cause information) of the model attenuation, and then effectively solve the problem of model attenuation after going online. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 An optional flowchart of the model attenuation attribution method provided in an embodiment of the present application;

[0037] Figure 2 An optional flowchart of the model attenuation attribution method provided in an embodiment of the present application;

[0038] Figure 3 An optional flowchart of the model attenuation attribution method provided in an embodiment of the present application;

[0039] Figure 4 An optional effect schematic diagram of the model attenuation attribution method provided in an embodiment of the present application;

[0040] Figure 5 An optional effect schematic diagram of the model attenuation attribution method provided in an embodiment of the present application;

[0041] Figure 6 An optional effect schematic diagram of the model attenuation attribution method provided in an embodiment of the present application;

[0042] Figure 7 An optional flowchart of the model attenuation attribution method provided in an embodiment of the present application;

[0043] Figure 8 An optional effect schematic diagram of the model attenuation attribution method provided in an embodiment of the present application;

[0044] Fig. 9 An optional flowchart of the model attenuation attribution method provided in an embodiment of the present application;

[0045] Fig.10 An optional effect schematic diagram of the model attenuation attribution method provided in an embodiment of the present application;

[0046] Fig.11 An optional flowchart of the model attenuation attribution method provided in an embodiment of the present application;

[0047] Fig.12 A schematic diagram of the structure of a model attenuation attribution device provided in an embodiment of the present application;

[0048] Fig.13 A hardware entity schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0049] In order to make the purpose, technical solutions and advantages of the present application clearer, the technical solutions of the present application are further elaborated in detail below in conjunction with the drawings and embodiments. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.

[0050] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0051] If similar descriptions of "first / second" appear in the application documents, the following instructions are added. In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0052] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0053] In an embodiment of the present application, the entity that executes the model attenuation attribution may be an intelligent terminal or server with data processing capabilities.

[0054] This application embodiment provides a model attenuation attribution method, see Figure 1 , which is an optional flow chart of the model attenuation attribution method provided in the embodiment of the present application, will be combined with Figure 1 The steps shown are explained.

[0055] S101. Determine N key features with the greatest feature importance among multiple features included in sample data.

[0056] In the embodiment of the present application, the model is a model that has been put into use after being launched. The model can be a model of various machine learning and deep learning. The model needs to be trained with sample data before going online. Each sample data may include multiple features. The importance of each feature of the sample data can be detected to determine the N key features with the greatest feature importance.

[0057] For example, the sample data may be parameter information of an item, and multiple features of the sample data may include: item price, item size, item location, etc. The sample data may also be user-related information, and multiple features of the sample data may include: user gender, user age, user height, and number of user transactions, etc. In the embodiments of the present application, when obtaining user-related information, it is obtained within the scope permitted by law and with the user's consent.

[0058] For example, the importance of the user's gender, user's age, user's height and user's transaction times corresponding to each user's information may be tested to determine the top N most important key features.

[0059] S102: Determine the root cause information of the model attenuation based on the change in probability density distribution between each of the key features and similar application features in the application data.

[0060] In the embodiment of the present application, the root cause information of model attenuation can be determined based on the change in probability density distribution between each key feature and similar application features in the application data, where the application data is the data obtained after the model is launched.

[0061] In some other embodiments, label sample data may be determined from the sample data. The label data may characterize risky data in the sample data. Since the label sample data belongs to the sample data, the label sample data also includes N key label features. The root cause information of model attenuation may be determined based on the change in probability density distribution between each key label feature and similar application features in the application data.

[0062] For example, the application data may be user-related information processed after the model goes online, or parameter information of an item.

[0063] In the embodiment of the present application, each key feature can be ranked based on the change in probability density distribution between each key feature and similar application features. The key feature ranked in the front has a greater causal relationship with the model attenuation.

[0064] In an embodiment of the present application, after determining several key features that have the greatest causal relationship with model attenuation, they can be displayed so that the target object can accurately locate the root cause of the model performance attenuation problem and make targeted improvements based on the displayed key features.

[0065] The purpose of this application is to use a variety of univariate analysis, data drift, concept drift, and drift and model performance correlation coefficients and other technical means to accurately and quickly detect which variables have drifted when the model performance decays, so as to accurately attribute the root cause of the model performance decay, and provide model developers with a basis for improving the model and corresponding improvement suggestions. The entire process of this solution is relatively complete and applicable to the attribution of almost all machine learning model performance decay. It can adapt to most algorithm scenarios without the need for algorithm personnel to perform a lot of complex configurations. Therefore, it can be promoted and widely used, allowing algorithm personnel to accurately locate the cause of model performance decay, saving a lot of time for trial and error exploration and repeated training, and improving production efficiency.

[0066] In an embodiment of the present application, N key features with the greatest feature importance are determined among the multiple features included in the sample data; wherein N is an integer greater than 0; based on the change in probability density distribution between each key feature and the same type of application features in the application data, the root cause information of model attenuation is determined; wherein the application data is the data acquired after the model is launched. In an embodiment of the present application, the root cause information of model attenuation is determined by analyzing the change in probability density distribution between each key feature that has the greatest impact on the model and the application features in the application data. When the model is attenuated, it is possible to accurately and quickly determine which variables have drifted, and then accurately determine the root cause (root cause information) of the model attenuation, and then effectively solve the problem of model attenuation after going online.

[0067] In some embodiments, see Figure 2 , Figure 2 An optional flow chart of the model attenuation attribution method provided in an embodiment of the present application, Figure 1 The illustrated S102 can also be implemented through S103 to S105, which will be described in conjunction with each step.

[0068] S103: Determine a first probability density distribution change between each of the key features and the corresponding application feature in the application data.

[0069] In the embodiment of the present application, a corresponding probability density distribution can be determined for each key feature in the sample data, and a corresponding probability density distribution can be determined for the application feature corresponding to each key feature in the application data. The first probability density distribution change is determined based on the change between the probability density distribution of each key feature and the probability density distribution of the corresponding application feature.

[0070] In an embodiment of the present application, a corresponding probability density distribution graph may be determined for each key feature. A corresponding probability density distribution graph may be determined for an application feature corresponding to each key feature in the application data. A first probability density distribution change between each key feature and the corresponding application feature in the application data is determined based on the probability density distribution graph of each key feature and the probability density distribution graph of the corresponding application feature.

[0071] The first probability density distribution change can also be called univariate data drift: Data drift refers to the quantitative change of observed data relative to the training data, which can have a huge impact on the quality of model predictions over time, and often for the worse. Tracking drift indicators related to training features and predictions should be an important part of model monitoring and identifying when the model should be retrained.

[0072] S104: Determine a second probability density distribution change between each of the key tag features and the corresponding application feature in the application data.

[0073] In the embodiment of the present application, the sample data includes: label sample data; the label sample data corresponds to N key label features; the N key features belong to the N features. A corresponding probability density distribution can be determined for each key label feature in the label sample data, and a corresponding probability density distribution can be determined for each application feature corresponding to each key label feature in the application data. Based on the change in the probability density distribution of each key label feature and the probability density distribution of the corresponding application feature, the second probability density distribution change is determined.

[0074] In an embodiment of the present application, a corresponding probability density distribution graph may be determined for each key tag feature. A corresponding probability density distribution graph may be determined for an application feature corresponding to each key tag feature in the application data. A second probability density distribution change between each key tag feature and the corresponding application feature in the application data is determined based on the probability density distribution graph of each key tag feature and the probability density distribution graph of the corresponding application feature.

[0075] The second probability density distribution change can also be called univariate concept drift detection: concept drift is a phenomenon in predictive analysis and machine learning that the statistical characteristics of the target variable change in an unpredictable way over time. Over time, the prediction accuracy of the model will decrease.

[0076] S105. Determine the root information based on the correlation between the performance attenuation of the model and the first probability density distribution change corresponding to each of the key features and the second probability density distribution change corresponding to each of the key tag features.

[0077] In the embodiment of the present application, based on the correlation between the performance decay of the model and the first probability density distribution change corresponding to each key feature and the second probability density distribution change corresponding to each key label feature, each key feature and each key label feature in the same sequence are sorted according to the size of the correlation. The greater the causal relationship between the key feature or key label feature ranked first and the performance decay of the model, the greater the causal relationship between the key feature or key label feature ranked first and the performance decay of the model. The key feature or key label feature ranked first can be determined as the root information.

[0078] In an embodiment of the present application, the root cause information of model attenuation is determined by analyzing the changes in the probability density distribution between each key feature and each key label feature that has the greatest impact on the model and the application features in the application data. When the model attenuates, it is possible to accurately and quickly determine which variables have drifted, and thus quickly and accurately determine the cause of the model attenuation.

[0079] In some embodiments, see Figure 3 , Figure 3 An optional flow chart of the model attenuation attribution method provided in an embodiment of the present application, Figure 3 It is shown that S103 to S105 can also be implemented through S106 to S112, which will be explained in combination with each step.

[0080] S106: Determine a first probability density distribution of each of the key features, and a second probability density distribution of the application features corresponding to each of the key features in the application data.

[0081] In the embodiment of the present application, the corresponding first probability density distribution may be determined based on each key feature of the sample data, and the second probability density distribution may be determined based on the application feature corresponding to each key feature in the application data.

[0082] In the embodiment of the present application, the first probability density distribution and the second probability density distribution can also be determined in the same probability density distribution diagram, so as to more intuitively show the probability density distribution changes between each key feature and the corresponding application feature.

[0083] Exemplary, combined Figure 4 , you can Figure 4 The probability density distribution diagrams corresponding to the key features of the sample data in different time periods and the corresponding application features are determined. Figure 4 The horizontal axis is the value of the feature, and the vertical axis is the probability density of the corresponding feature value. Figure 4 It shows that the probability density distribution between the X application features from January to March 2023 and the key features for the whole year of 2022, January to May 2022, and September to December 2022 has different degrees of drift, and the specific degree of drift is measured by the corresponding KL (Kullback-Leibler divergence, KLD) divergence. Among them, the distribution drift of the first probability density distribution and the second probability density is the largest when the feature value is -1.0 (where the __ line corresponds to the first probability density distribution, the _ line corresponds to the second probability density distribution, and the _____ line corresponds to the third probability density distribution).

[0084] S107: Determine a first relative entropy between a first probability density distribution corresponding to each of the key features and the corresponding second probability density distribution.

[0085] In the embodiment of the present application, the corresponding first relative entropy can be determined based on the first probability density distribution corresponding to each key feature and the second probability density distribution corresponding to the application feature.

[0086] Among them, the first relative entropy is also called KL divergence (Kullback-Leibler divergence, abbreviated as KLD): information divergence, information gain. KL divergence is a measure of the asymmetry of the difference between two probability distributions P (first probability density distribution) and Q (second probability density distribution). If the two distributions P and Q are very far apart and have no overlap at all, then the KL divergence value is meaningless. KL divergence is used to measure the number of additional bits required to encode the average of "samples from P" using "Q-based coding". Typically, P represents the true distribution of the data, and Q represents the theoretical distribution of the data, the model distribution, or the approximate distribution of P. The definition is as follows:

[0087]

[0088] Among them, P(x), Q(x) are probability distributions.

[0089] S108: Determine a third probability density distribution of each of the key tag features and a fourth probability density distribution of the application features corresponding to each of the key tag features in the application data.

[0090] In the embodiment of the present application, the corresponding third probability density distribution can be determined based on each key label feature of the sample label data, and the fourth probability density distribution can be determined based on the application feature corresponding to each key label feature in the application data.

[0091] In the embodiment of the present application, the third probability density distribution and the fourth probability density distribution can also be determined in the same probability density distribution diagram, so as to more intuitively show the probability density distribution changes between each key tag feature and the corresponding application feature.

[0092] S109: Determine a second relative entropy between the third probability density distribution corresponding to each of the key tag features and the corresponding fourth probability density distribution.

[0093] In the embodiment of the present application, the corresponding second relative entropy can be determined based on the third probability density distribution corresponding to each key tag feature and the second probability density distribution of the corresponding application feature.

[0094] For example, univariate concept drift detection: Compared with the full univariate drift, we take the data with label 1 (risky) to perform corresponding univariate drift detection, which is to detect the drift of x when the corresponding y value remains unchanged, that is, the concept drift situation. Figure 5 is the total univariate drift, Figure 6 is the univariate concept drift. Figure 5 It shows that the probability density distribution between the X application features from January to March 2023 and the key features for the whole year of 2022, from January to May 2022, and from September to December 2022 all drift to varying degrees. Figure 6 It shows that the probability density distribution between the X application features in the labeled sample data from January to March 2023 and the key features during the whole year of 22, January to May 22, and September to December 22 all have different degrees of drift. Comparing the drift of the two groups of x without the constraint y = 1 and with the constraint y = 1, we found that the drift of P(x) has a kl range of about 0.0005, while the drift of P(x|y = 1) has a kl range of about 0.02, which shows that the latter is about 40 times the former. Therefore, it can be judged that there is an obvious concept drift between the variable from January to March 2023 and other time windows.

[0095] S110: Determine a plurality of performance information of the model within a predetermined period of time.

[0096] In an embodiment of the present application, performance testing can be performed on the model during the training period and the initial online stage to determine multiple performance information corresponding to the model during the training period and the initial online stage.

[0097] In the embodiment of the present application, the model effect evaluation index detection can be performed on the model performance during the training period and the initial launch period to determine the model's (Receiver Operating Characteristic, ROC) and (Area under the curve, AUC) scores, and then multiple performance scores can be obtained.

[0098] S111. Determine performance degradation information of the model based on the multiple performance information.

[0099] In an embodiment of the present application, multiple performance information may be analyzed to determine the attenuation information of the model performance from the training period to the initial launch period.

[0100] In the embodiment of the present application, the performance degradation information can be determined by subtracting the lowest value from the highest value of the multiple performance information.

[0101] S112: Determine the root cause information based on the correlation between the performance degradation information and the first relative entropy corresponding to each of the key features, and the correlation between the performance degradation information and the second relative entropy corresponding to each of the key tag features.

[0102] In the embodiment of the present application, the correlation coefficient between the performance attenuation information and the first relative entropy corresponding to each of the key features can be determined, and the correlation coefficient between the performance attenuation information and the second relative entropy corresponding to each of the key label features can be determined. Based on the size of the correlation coefficient, each key feature and each key label feature are sorted, and the key feature or key label feature with the largest correlation coefficient is determined as the root information.

[0103] In an embodiment of the present application, the root cause information of model decay is determined by analyzing the relative entropy between each key feature and each key tag feature and the application feature in the application data. When the model decays, it is possible to accurately and quickly determine which variables have drifted, and thus the cause of the model decay can be quickly and accurately determined.

[0104] In some embodiments, see Figure 7 , Figure 7 An optional flow chart of the model attenuation attribution method provided in an embodiment of the present application, Figure 3 It is shown that S112 can also be implemented through S113 to S116, which will be explained in combination with each step.

[0105] S113: Determine a first correlation coefficient between the performance degradation information and the first relative entropy corresponding to each of the key features.

[0106] In the embodiment of the present application, a correlation calculation may be performed on the first relative entropy corresponding to each key feature and the performance degradation information to determine the first correlation coefficient corresponding to each key feature.

[0107] In the embodiment of the present application, the corresponding Pearson correlation coefficient (first correlation coefficient) can be calculated for the first relative entropy corresponding to each key feature and the performance decay information.

[0108] Among them, Pearson correlation coefficient: In statistics, the Pearson product-moment correlation coefficient (abbreviation: PPMCC, or PCCs, sometimes referred to as correlation coefficient) is used to measure the degree of linear correlation between variables X and Y of two sets of data. It is the ratio of the covariance of the two variables to the product of their standard deviation; therefore, it is essentially a normalized measure of covariance, so the result always has a value between -1 and 1. Like the covariance itself, this measure can only reflect the linear correlation of the variables, ignoring many other types of relationships or correlations. As a simple example, it can be expected that the Pearson product-moment correlation coefficient of age and height of a sample of high school teenagers is significantly greater than 0, but less than 1 (because 1 represents an unrealistic perfect correlation).

[0109] S114: Determine a second correlation coefficient between the performance degradation information and the second relative entropy corresponding to each of the key tag features.

[0110] In the embodiment of the present application, a correlation calculation may be performed on the second relative entropy corresponding to each key label feature and the performance decay information to determine the second correlation coefficient corresponding to each key label feature.

[0111] In the embodiment of the present application, the corresponding Pearson correlation coefficient (second correlation coefficient) can be calculated for the first relative entropy and performance decay information corresponding to each key tag feature.

[0112] S115 . Based on the first correlation coefficient corresponding to each key feature and the second correlation coefficient corresponding to each key label feature, sort each key feature and each key label feature to obtain a root feature set.

[0113] In the embodiment of the present application, each key feature and each key label feature may be sorted based on the first correlation coefficient corresponding to each key feature and the second correlation coefficient corresponding to each key label feature to obtain a root feature set.

[0114] Exemplary, combined Figure 8, showing the Pearson correlation coefficient of each key feature of the model, each key label feature and the model performance decay information, and sorting the corresponding features from large to small according to the Pearson correlation coefficient. The strength of the influence of single variable drift on model decay can be obtained, so as to pay attention to the variables with the most serious influence and make further detection and attribution. Among them, the Pearson correlation coefficient corresponding to feature X is 0.943975, which is the largest, indicating that feature X is the variable with the most serious influence on the model.

[0115] S116: Determine the root source information based on the root source feature set.

[0116] In the embodiment of the present application, the first M features in the root feature set may be determined as the root information, where M is an integer greater than 0.

[0117] In an embodiment of the present application, each key feature and each key label feature are sorted by the first correlation coefficient between the first relative entropy of each key feature and the corresponding application feature, and the second correlation coefficient between the second relative entropy of each key label feature and the corresponding application feature to obtain a root feature set, which is convenient for more intuitively determining which features have the greatest causal relationship with the attenuation of model performance, and can quickly and accurately determine the root information among all features in the root feature set.

[0118] In some embodiments, see Fig. 9 , Fig. 9 An optional flow chart of the model attenuation attribution method provided in an embodiment of the present application, Figure 1 It is shown that S101 can also be implemented through S117 to S118, which will be explained in combination with each step.

[0119] S117, performing importance calculation processing on multiple features included in the sample data respectively through a preset model to obtain the importance corresponding to each of the features.

[0120] In the embodiment of the present application, the importance corresponding to each feature in the sample data can be calculated by the SHAP analysis algorithm.

[0121] For example, the height feature in the user's relevant information may be used to calculate the corresponding importance through the SHAP analysis algorithm.

[0122] S118: Determine the features corresponding to the top N greatest importances as the N key features.

[0123] In the embodiment of the present application, the features corresponding to the top N greatest importances may be the N key features.

[0124] Exemplary, combined Fig.10,Univariate analysis: We sort multiple features according to their importance. Univariate analysis is performed on the top 10 or 20 features, and the WL divergence score is given. The higher the WL divergence, the more severe the drift. Among them, feature X has the highest importance.

[0125] In the embodiment of the present application, by calculating the importance of multiple features respectively, the features corresponding to the top N greatest importances are determined as the N key features, thereby reducing the noise interference of the features that have less impact on the model performance among the multiple features.

[0126] In some embodiments, see Fig.11 , Fig.11 An optional flow chart of the model attenuation attribution method provided for an embodiment of the present application will be described in conjunction with each step.

[0127] S201. Evaluate the feature importance in the model.

[0128] Model feature importance detection: Before detecting the model performance degradation caused by data drift and concept drift, we must first detect which features are most important to the model. Changes in these most important features will have a more significant impact on the model performance, causing the model performance to decline significantly. There are many algorithmic techniques to measure the model feature importance, such as the totalgain of the Extreme Gradient Boosting (XGBoost) model. Here we use the SHAP value to measure feature importance. First, it can more scientifically measure the impact of each feature on the final result of the model. Second, it can be model-independent, that is, it is applicable to traditional statistical models and various deep models.

[0129] S202: Detect data drift of the most important N key features.

[0130] Univariate data drift detection: After the feature importance is determined, we mainly detect the drift of the top 10 and 20 feature data in terms of importance. We mainly use univariate analysis techniques to detect the drift of the top 10 and top 20 feature data before and after the model is launched. We can draw a probability density distribution graph of this part of the data before and after the model is launched to compare the distribution changes before and after the launch. This part of the drift is the drift of P(x), and the WL divergence can be used to measure the degree of distribution change.

[0131] S203, detecting the concept drift of the most important key tag features.

[0132] Single variable concept drift detection: Based on the second step, we only select the key label features corresponding to the label sample data with label 1 (in the risk control scenario, label 1 means risky data). Detect the change in the probability density distribution between the key label features and the corresponding application features, that is, the drift of P(x|y=1). This part of the drift means the drift of the relationship between x and y when y is fixed, which is the so-called concept drift. If the WL divergence of this part of the drift is much larger than the drift of P(x) detected in the second step, it can be considered that the single variable has undergone a relatively obvious concept drift.

[0133] S204. Detect the correlation between the drift of the most important features and the degradation of model performance.

[0134] Sorting the correlation between univariate concept drift and model performance: ROC and AUC scores are performed on model performance to detect the decline coefficient of ROC, AUC and other scores of the model during the relative stability period of the attenuation period (training and initial launch). At the same time, indicators such as the univariate WL divergence detected in the above steps measure the degree of univariate drift. The Pearson correlation coefficient between these two indicators can be calculated to obtain the relationship between univariate drift and model performance attenuation. The larger the Pearson correlation coefficient, the more correlated the drift of the univariate is with the model attenuation, thereby guiding the algorithm personnel to check the variable and ultimately accurately locate the cause of model attenuation.

[0135] S205. Based on the correlation between drift and model performance, locate the cause of model performance degradation and make improvements.

[0136] Through the above steps, we can analyze: 1. Whether the features that have a key impact on model performance have drifted, 2. Whether the relationship between these features and label data has drifted, that is, concept drift, 3. The direct relationship between the degree of data drift and the degradation of model performance. From the conclusions of these analyses, we can clearly understand which feature data drifts are the real cause of model performance degradation, so as to further check the underlying reasons for data changes and make corresponding processing to improve the performance of the model and ensure that it can function stably after going online.

[0137] This application provides algorithm and engineering personnel with a set of attribution technology means and processes for data drift, concept drift and model performance degradation, so as to accurately interpret the root causes of model performance degradation problems, so that algorithm personnel can accurately locate the root causes of model performance degradation problems and make targeted improvements. This solution provides a complete set of tools packaged into a python algorithm library, and provides functions such as data visualization, and the entire pipeline solves problems related to data drift and concept drift detection. This solution provides common best practice examples of related algorithms in the form of jupyter notebook code to guide algorithm engineering personnel to use it, making the entire algorithm pipeline easier to operate and promote.

[0138] See also Fig.12 , Fig.12 A schematic diagram of the structure of a model attenuation attribution device provided in an embodiment of the present application.

[0139] The embodiment of the present application also provides a model attenuation attribution device 800 , including: a determination unit 801 .

[0140] A determination unit 801 is used to determine N key features with the greatest feature importance among the multiple features included in the sample data; wherein N is an integer greater than 0;

[0141] The determination unit 801 is used to determine the root cause information of the model attenuation based on the change in probability density distribution between each of the key features in the N features and similar application features in the application data; wherein the application data is the data obtained after the model is launched.

[0142] In an embodiment of the present application, the sample data includes: label sample data; the label sample data corresponds to N key label features; the N key features belong to the N features; the determination unit 801 in the model attenuation attribution device 800 is used to determine the first probability density distribution change between each of the key features and the corresponding application features in the application data; determine the second probability density distribution change between each of the key label features and the corresponding application features in the application data; based on the correlation between the performance attenuation of the model and the first probability density distribution change corresponding to each of the key features and the second probability density distribution change corresponding to each of the key label features, determine the root information.

[0143] In an embodiment of the present application, the determination unit 801 in the model attenuation attribution device 800 is used to determine a first probability density distribution of each of the key features, and a second probability density distribution of the application feature corresponding to each of the key features in the application data; determine a first relative entropy between the first probability density distribution corresponding to each of the key features and the corresponding second probability density distribution; wherein the first relative entropy is used to characterize the change in the first probability density distribution between each of the key features and the corresponding application feature.

[0144] In the embodiment of the present application, the determination unit 801 in the model attenuation attribution device 800 is used to determine the third probability density distribution of each of the key label features, and the fourth probability density distribution of the application feature corresponding to each of the key label features in the application data;

[0145] Determine a second relative entropy between the third probability density distribution corresponding to each of the key tag features and the corresponding fourth probability density distribution; wherein the second relative entropy is used to characterize the change in the second probability density distribution between each of the key tag features and the corresponding application features.

[0146] In an embodiment of the present application, the first probability density distribution change includes: a first relative entropy; the second probability density distribution change includes: a second relative entropy; the determination unit 801 in the model attenuation attribution device 800 is used to determine multiple performance information of the model within a predetermined time period; determine the performance attenuation information of the model based on the multiple performance information; determine the root information based on the correlation between the performance attenuation information and the first relative entropy corresponding to each of the key features, and the correlation between the performance attenuation information and the second relative entropy corresponding to each of the key label features.

[0147] The determination unit 801 in the model attenuation attribution device 800 is used to determine the first correlation coefficient between the performance attenuation information and the first relative entropy corresponding to each of the key features; determine the second correlation coefficient between the performance attenuation information and the second relative entropy corresponding to each of the key label features; based on the first correlation coefficient corresponding to each of the key features and the second correlation coefficient corresponding to each of the key label features, sort each of the key features and each of the key label features to obtain a root feature set; determine the root information based on the root feature set.

[0148] The determination unit 801 in the model attenuation attribution device 800 is used to calculate the importance of multiple features included in the sample data through a preset model to obtain the importance of each feature; and determine that the features corresponding to the top N largest importances are the N key features.

[0149] It should be noted that in the embodiments of the present application, if the above-mentioned model attenuation attribution method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application can be essentially or partly reflected in the form of a software product, which is stored in a storage medium and includes several instructions for a model attenuation attribution device (which can be a personal computer, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a disk or an optical disk. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0150] Correspondingly, an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the steps in the above method are implemented.

[0151] Correspondingly, an embodiment of the present application provides an electronic device 900, including a memory 902 and a processor 901, wherein the memory 902 stores a computer program that can be executed on the processor 901, and the processor 901 implements the steps in the above method when executing the program.

[0152] It should be noted here that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.

[0153] It should be noted that Fig.13 A hardware entity diagram of an electronic device provided in an embodiment of the present application, such as Fig.13 As shown, the hardware entity of the electronic device 900 includes: a processor 901 and a memory 902, wherein;

[0154] The processor 901 generally controls the overall operation of the electronic device 900 .

[0155] The memory 902 is configured to store instructions and applications executable by the processor 901, and can also cache data to be processed or processed by the processor 901 and various modules in the electronic device 900 (for example, image data, audio data, voice communication data, and video communication data), which can be implemented through flash memory (FLASH) or random access memory (Random Access Memory, RAM).

[0156] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned sequence numbers of the embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.

[0157] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0158] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be electrical, mechanical or other forms.

[0159] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0160] In addition, all functional units in the embodiments of the present application may be integrated into one processing unit, or each unit may be a separate unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0161] A person skilled in the art can understand that all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: a mobile storage device, a read-only memory (ROM), a magnetic disk or an optical disk, and other media that can store program codes.

[0162] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application can be essentially or partly embodied in the form of a software product that contributes to the relevant technology. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a mobile storage device, a ROM, a magnetic disk, or an optical disk.

[0163] The above is only an implementation method of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A model decay attribution method, It is characterized in that include: Determine N key features with the greatest feature importance among multiple features included in the sample data; wherein N is an integer greater than 0; Based on the change in probability density distribution between each of the key features and similar application features in the application data, the root cause information of the model attenuation is determined; wherein the application data is the data obtained after the model is launched.

2. The model attenuation attribution method according to claim 1, It is characterized in that The sample data includes: label sample data; the label sample data corresponds to N key label features; the N key features belong to the N features; the root cause information of the model attenuation is determined based on the probability density distribution change between each of the key features and the same application features in the application data, including: Determine a first probability density distribution change between each of the key features and the corresponding application feature in the application data; Determine a second probability density distribution change between each of the key tag features and the corresponding application feature in the application data; The root cause information is determined based on the correlation between the performance attenuation of the model and the first probability density distribution change corresponding to each of the key features and the second probability density distribution change corresponding to each of the key tag features.

3. The model attenuation attribution method according to claim 2, It is characterized in that The determining of a first probability density distribution change between each of the key features and the corresponding application feature in the application data includes: Determine a first probability density distribution of each of the key features, and a second probability density distribution of the application features corresponding to each of the key features in the application data; Determine a first relative entropy between a first probability density distribution corresponding to each of the key features and the corresponding second probability density distribution; wherein the first relative entropy is used to characterize a change in the first probability density distribution between each of the key features and the corresponding application feature.

4. The model attenuation attribution method according to claim 2, It is characterized in that The determining of a second probability density distribution change between each of the key tag features and the corresponding application feature in the application data includes: Determine a third probability density distribution of each of the key tag features, and a fourth probability density distribution of the application features corresponding to each of the key tag features in the application data; Determine a second relative entropy between the third probability density distribution corresponding to each of the key tag features and the corresponding fourth probability density distribution; wherein the second relative entropy is used to characterize the change in the second probability density distribution between each of the key tag features and the corresponding application features.

5. The model attenuation attribution method according to any one of claims 2 to 4, It is characterized in that The first probability density distribution change includes: a first relative entropy; the second probability density distribution change includes: a second relative entropy; the correlation between the performance attenuation based on the model and the first probability density distribution change corresponding to each of the key features and the second probability density distribution change corresponding to each of the key tag features, determining the root information, includes: determining a plurality of performance information of the model within a predetermined period of time; determining performance degradation information of the model based on a plurality of the performance information; The root cause information is determined based on the correlation between the performance degradation information and the first relative entropy corresponding to each of the key features, and the correlation between the performance degradation information and the second relative entropy corresponding to each of the key tag features.

6. The model attenuation attribution method according to claim 5, It is characterized in that The determining the root source information based on the correlation between the performance decay information and the first relative entropy corresponding to each of the key features, and the correlation between the performance decay information and the second relative entropy corresponding to each of the key tag features, includes: Determine a first correlation coefficient between the performance degradation information and the first relative entropy corresponding to each of the key features; Determine a second correlation coefficient between the performance decay information and the second relative entropy corresponding to each of the key tag features; Based on the first correlation coefficient corresponding to each of the key features and the second correlation coefficient corresponding to each of the key label features, sorting each of the key features and each of the key label features to obtain a root feature set; The root source information is determined based on the root source feature set.

7. The model attenuation attribution method according to claim 1, It is characterized in that The determining of N key features with the greatest feature importance among the multiple features included in the sample data includes: Using a preset model, respectively calculate the importance of multiple features included in the sample data to obtain the importance of each feature; It is determined that the features corresponding to the top N greatest importances are the N key features.

8. A model attenuation attribution device, It is characterized in that include: A determination unit, used to determine N key features with the greatest feature importance among the multiple features included in the sample data; wherein N is an integer greater than 0; A determination unit is used to determine the root cause information of the model attenuation based on the change in probability density distribution between each of the key features in the N features and similar application features in the application data; wherein the application data is the data obtained after the model is launched.

9. An electronic device, It is characterized in that The method comprises a memory and a processor, wherein the memory stores a computer program executable on the processor, and the processor implements the steps in the method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the steps in the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Data mining method applied to hardware product control system

    CN121579853A