Model performance improvement method and device, electronic equipment and storage medium

By detecting the change in the probability density distribution of key features and training the sub-model, it integrates into the original model, solving the problem of model performance attenuation after running online, and achieving long-term stability and adaptability of model performance.

CN120106240APending Publication Date: 2025-06-06JINGDONG TECH HLDG CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311667388.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-06
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively solve the problem of model performance degradation after running online, resulting in model performance degradation or complete failure.

Method used

The feature detection results are determined by detecting the change in the probability density distribution between key features and similar features in the application data. When key features drift, the sub-model is trained based on the new sample data acquired by timing and integrate the sub-model into the original model to form a new model.

Benefits of technology

This method can detect feature drift in time, train the model as needed, and adapt to data changes, thereby effectively preventing model performance attenuation and improving the long-term operation performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106240A_ABST
    Figure CN120106240A_ABST
Patent Text Reader

Abstract

The invention provides a model performance improvement method and device, electronic equipment and a storage medium, and relates to the technical field of computers and the Internet. The method comprises the steps of determining a feature detection result based on probability density distribution change between each key feature and similar application features in application data; the application data is data obtained after the model is online; the key features are M features with the maximum feature importance in the sample data; under the condition that the feature detection result represents key feature drift, training based on new sample data obtained according to a time sequence to obtain a corresponding sub-model; a new model is determined based on the sub-model and the model. According to the method, the integrated training process of the model is triggered only when drifting of the key features is detected, so that the purposes that the model is not retrained blindly, the model is trained as required to adapt to the feature drifting condition, the model is adjusted in time to prevent performance degradation of the model, and then the problem of model attenuation after online can be effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer and Internet technology, and in particular to a model performance improvement method, device, electronic device and storage medium. Background Art

[0002] With the development of artificial intelligence (AI) technology, various machine learning and deep learning models are increasingly widely used in various production environments. However, most models will experience performance degradation, performance decay, or even complete failure after running online for a period of time. The more common solution at present is to take measures such as retraining the model only after the model performance has obviously decayed. However, the final effect of this method may not completely improve the performance of the online model. It cannot solve the problem of model performance decay after going online. Summary of the invention

[0003] The embodiments of the present application provide a model performance improvement method, device, electronic device and storage medium, which can effectively solve the problem of model performance degradation after going online.

[0004] The technical solution of this application is implemented as follows:

[0005] The present application embodiment provides a method for improving model performance, including:

[0006] Determine the feature detection result based on the change in probability density distribution between each key feature and similar application features in the application data; wherein the application data is data acquired after the model is launched; the key features are the M features with the greatest feature importance in the sample data; M is an integer greater than 0;

[0007] In the case where the feature detection result represents the drift of the key feature, a corresponding sub-model is obtained by training based on new sample data acquired in time series;

[0008] A new model is determined based on the sub-models and the model.

[0009] In the above solution, determining the feature detection result based on the probability density distribution change between each key feature and similar application features in the application data includes:

[0010] Determine the relative entropy based on the probability density changes corresponding to each of the key features and the corresponding application features; wherein the relative entropy is used to characterize the probability density distribution changes between each key feature and the corresponding application feature;

[0011] The feature detection result is determined based on the relative entropy corresponding to each of the key features.

[0012] In the above solution, determining the feature detection result based on the relative entropy corresponding to each of the key features includes one of the following:

[0013] If the relative entropy corresponding to any of the key features is outside a preset range, determining the feature detection result characterizing the drift of the key feature;

[0014] If the relative entropy corresponding to each of the key features is within the preset range, the feature detection result indicating that the key feature is normal is determined.

[0015] In the above solution, the corresponding sub-model is obtained based on the training of new sample data acquired in time series, including:

[0016] Acquire the new sample data within each predetermined time period in a time sequence;

[0017] The initial sub-model is trained based on the new sample data until the training condition is met, thereby obtaining the sub-model of the new sample data.

[0018] In the above solution, determining a new model based on the sub-model and the model includes:

[0019] Integrating the sub-models obtained in time sequence into the model to form an intermediate model;

[0020] The new model is determined based on at least one of the number and performance information of the sub-models in the intermediate model.

[0021] In the above solution, the intermediate model includes: N sub-models; N is an integer greater than 0;

[0022] The determining of the new model based on at least one of the number and performance information of the sub-models in the intermediate model comprises:

[0023] If the N sub-models reach a predetermined number, then deleting the sub-model with the earliest time sequence in the intermediate model;

[0024] Training the initial sub-model based on new sample data within a next predetermined period of time to obtain a first new sub-model;

[0025] The first new sub-model is integrated into the intermediate model to obtain the new model.

[0026] In the above solution, the intermediate model includes: N sub-models; N is an integer greater than 0;

[0027] The determining of the new model based on at least one of the number and performance information of the sub-models in the intermediate model comprises:

[0028] If the N sub-models do not reach the predetermined number, determining the performance information respectively corresponding to the N sub-models;

[0029] The new model is obtained based on the performance information respectively corresponding to the N sub-models and the integration of the second new sub-model; wherein the second new sub-model is trained on the initial sub-model based on the new sample data within the next predetermined time period.

[0030] In the above solution, before determining the feature detection result based on the change in probability density distribution between each key feature and similar application features in the application data, the method further includes:

[0031] The importance of each feature in the sample data is calculated by using a preset model to obtain the importance of each feature;

[0032] The features corresponding to the top M greatest importances are determined as the M key features.

[0033] The present application also provides a model performance improvement device, including:

[0034] A feature detection unit, used to determine a feature detection result based on a change in probability density distribution between each key feature and similar application features in the application data; wherein the application data is data acquired after the model is launched; the key features are M features with the greatest feature importance in the sample data; M is an integer greater than 0;

[0035] A model training unit, configured to obtain a corresponding sub-model by training based on new sample data acquired in time series when the feature detection result represents the drift of the key feature;

[0036] A model determination unit is used to determine a new model based on the sub-model and the model.

[0037] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and the processor implements the steps in the above method when executing the computer program.

[0038] An embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method are implemented.

[0039] In the embodiment of the present application, the feature detection result is determined based on the change in probability density distribution between each key feature and similar application features in the application data; wherein the application data is the data obtained after the model is launched; the key features are the M features with the greatest feature importance in the sample data; M is an integer greater than 0; in the case where the feature detection result represents the drift of the key feature, the corresponding sub-model is obtained based on the training of the new sample data obtained in time series; and the new model is determined based on the sub-model and the model. In the embodiment of the present application, the integrated training process of the model is triggered when the drift of the key feature is detected, so as to achieve the goal of not blindly retraining the model, but training the model on demand to adapt to the situation of feature drift, and timely adjusting the model to prevent the model performance from decaying, thereby effectively solving the problem of model decay after going online. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 An optional flowchart of a method for improving model performance provided in an embodiment of the present application;

[0041] Figure 2 An optional flowchart of a method for improving model performance provided in an embodiment of the present application;

[0042] Figure 3 An optional effect diagram of the model performance improvement method provided in the embodiment of the present application;

[0043] Figure 4 An optional flowchart of a method for improving model performance provided in an embodiment of the present application;

[0044] Figure 5 An optional flowchart of a method for improving model performance provided in an embodiment of the present application;

[0045] Figure 6 An optional flowchart of a method for improving model performance provided in an embodiment of the present application;

[0046] Figure 7 An optional flowchart of a method for improving model performance provided in an embodiment of the present application;

[0047] Figure 8 An optional flowchart of a method for improving model performance provided in an embodiment of the present application;

[0048] Fig. 9 An optional effect diagram of the method for improving model performance provided in an embodiment of the present application;

[0049] Fig.10 An optional flowchart of a method for improving model performance provided in an embodiment of the present application;

[0050] Fig.11 A schematic diagram of the structure of a model performance improvement device provided in an embodiment of the present application;

[0051] Fig.12 A hardware entity schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical solutions and advantages of the present application clearer, the technical solutions of the present application are further elaborated in detail below in conjunction with the drawings and embodiments. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.

[0053] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0054] If similar descriptions of "first / second" appear in the application documents, the following instructions are added. In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0056] In the embodiment of the present application, the subject that performs the model performance improvement may be an intelligent terminal or server with data processing capabilities.

[0057] This application embodiment provides a method for improving model performance. Figure 1 , is an optional flow chart of the model performance improvement method provided in the embodiment of the present application, which will be combined with Figure 1 The steps shown are explained.

[0058] S101. Determine a feature detection result based on a change in probability density distribution between each key feature and similar application features in the application data.

[0059] In the embodiment of the present application, the feature detection result of whether each key feature has drifted can be determined based on the change in the probability density distribution between each key feature and the same application features in the application data. The application data is the data obtained after the model is launched. The key features are the M features with the greatest feature importance in the sample data. M is an integer greater than 0.

[0060] Among them, drift: Concept drift is an important data phenomenon in the field of AI learning, which is manifested as the inconsistency between online reasoning data (real-time distribution) and the training stage (historical distribution). Concept drift detection can detect changes in data distribution in a timely manner and predict signs of model failure in advance, which is of great significance for the timely adjustment of AI models. Concept drift detection is essentially to detect changes in data distribution. This example proposes a method to detect data changes by comparing whether the features of the new window data deviate sufficiently from the historical window features. If the degree of deviation is greater than a certain threshold, the data has undergone concept drift.

[0061] In the embodiment of the present application, the model is a model that has been put into use after being launched. The model can be a model of various types of machine learning and deep learning. The model needs to be trained with sample data before going online. Each sample data may include multiple features. The importance of each feature of the sample data can be detected to determine the N key features that are most important to the model performance.

[0062] For example, the sample data may be parameter information of an item, and multiple features of the sample data may include: item price, item size, item location, etc. The sample data may also be user-related information, and multiple features of the sample data may include: user gender, user age, user height, and number of user transactions, etc. In the embodiments of the present application, when obtaining user-related information, it is obtained within the scope permitted by law and with the user's consent.

[0063] S102: When the feature detection result represents the drift of the key feature, obtain a corresponding sub-model through training based on new sample data acquired in time series.

[0064] In the embodiment of the present application, when the key features drift, the application data used for detection can be obtained at intervals from the current moment as new sample data at intervals, thereby obtaining each new sample data. The initial sub-model is trained based on each new sample data to obtain the corresponding sub-model.

[0065] In an embodiment of the present application, if a key feature drifts, it can be said that the model performance is decaying, and then a process of training N sub-models using the latest acquired new sample data is performed. Among them, the data type of the new sample data and the sample data can be the same. Model performance decay: "Model decay" refers to a phenomenon in machine learning, that is, the predictions of the model become less accurate over time. The root cause is that the data changes over time due to changes in the underlying environment. Model drift, also known as model decay, refers to the gradual degradation of the performance of a machine learning model over time. This occurs when the model suddenly or gradually begins to provide predictions with lower accuracy than during training.

[0066] S103: Determine a new model based on the sub-model and the model.

[0067] In an embodiment of the present application, the earliest sub-model among the sub-models obtained by training based on the new sample data can be deleted, or the sub-model with the worst performance among the sub-models obtained by training with the new sample data can be deleted, and the remaining sub-models can be integrated and learned into the original model to obtain a new model with improved performance.

[0068] Among them, ensemble learning: In statistics and machine learning, ensemble learning (English: Ensemble learning) methods use multiple learning algorithms to obtain better predictive performance than any single learning algorithm alone. That is, the technology of using multiple compatible learning algorithms / models to perform a single task is to obtain better predictive performance. The main methods of ensemble learning can be classified into three categories: stacking, boosting, and bagging. The most popular methods include random forest, gradient boosting, AdaBoost, gradient boosted decision tree (GBDT) and XGBoost.

[0069] In the embodiment of the present application, the feature detection result is determined based on the change in probability density distribution between each key feature and similar application features in the application data; wherein the application data is the data obtained after the model is launched; the key features are the M features with the greatest feature importance in the sample data; in the case where the feature detection result represents the drift of the key features, the corresponding sub-model is obtained based on the training of the sample data acquired in time series; the new model is determined based on the sub-model and the model. In the embodiment of the present application, the integrated training process of the model is triggered when the drift of the key features is detected, so as to achieve the goal of not blindly retraining the model, but training the model on demand to adapt to the situation of feature drift, and timely adjusting the model to prevent the model performance from decaying, thereby effectively solving the problem of model decay after going online.

[0070] In some embodiments, see Figure 2 , Figure 2 An optional flow chart of the model performance improvement method provided in the embodiment of the present application, Figure 1 The illustrated S101 can also be implemented through S104 to S105, which will be described in conjunction with each step.

[0071] S104. Determine a relative entropy based on the change in probability density between each of the key features and the corresponding application features; wherein the relative entropy is used to characterize the change in probability density distribution between each of the key features and the corresponding application features.

[0072] In an embodiment of the present application, a corresponding first probability density distribution can be determined based on each key feature of the sample data. A second probability density distribution can be determined based on the application feature corresponding to each key feature in the application data. A corresponding relative entropy can be determined based on the first probability density distribution corresponding to each key feature and the second probability density distribution of the corresponding application feature. The relative entropy is used to characterize the change in the probability density distribution between each key feature and the corresponding application feature.

[0073] In the embodiment of the present application, the first probability density distribution and the second probability density distribution can also be determined in the same probability density distribution diagram, so as to more intuitively show the probability density distribution changes between each key feature and the corresponding application feature.

[0074] Exemplary, combined Figure 3 , you can Figure 3 The probability density distribution diagrams corresponding to the key features of the sample data in different time periods and the corresponding application features are determined. Figure 3 The horizontal axis is the value of the feature, and the vertical axis is the probability density of the corresponding feature value. Figure 3 It shows that the probability density distributions between the X application features from January to March 2023 and the key features for the whole year of 2022, January to May 2022, and September to December 2022 all have different degrees of drift, and the specific degree of drift is measured by the corresponding KL (Kullback-Leibler divergence, KLD) divergence. Among them, the first probability density distribution and the second probability density distribution have the largest drift when the feature value is -1.0 (where, The line corresponds to the first probability density distribution, The line corresponds to the second probability density distribution, and the - line corresponds to the third probability density distribution).

[0075] Among them, relative entropy is also called KL divergence: information divergence, information gain. The larger the KL divergence between two sets of data, the greater the difference in the distribution of the two sets of data. KL divergence is a measure of the asymmetry of the difference between two probability distributions P (first probability density distribution) and Q (second probability density distribution). If the two distributions P and Q are very far apart and there is no overlap at all, then the KL divergence value is meaningless. KL divergence is used to measure the number of additional bits required to encode the average of "samples from P" using "Q-based coding". Typically, P represents the true distribution of the data, and Q represents the theoretical distribution, model distribution, or approximate distribution of P.

[0076] S105. Determine the feature detection result based on the relative entropy corresponding to each of the key features.

[0077] In the embodiment of the present application, the relative entropy corresponding to each key feature can be compared with a preset range to determine the feature detection result. The preset range is a preset normal value range of the relative entropy. The feature detection result characterizes whether each key feature has drifted.

[0078] In the embodiment of the present application, if the relative entropy corresponding to any of the key features is outside a preset range, the feature detection result characterizing the drift of the key feature is determined.

[0079] In the embodiment of the present application, if the relative entropy corresponding to each of the key features is within the preset range, the feature detection result indicating that the key feature is normal is determined.

[0080] In the embodiment of the present application, the relative entropy is determined based on the change in the probability density corresponding to each of the key features and the corresponding application features, and the sub-model is trained based on whether the relative entropy is within a preset range, and the process of determining the new model is performed. The integrated training process of the model is triggered only when the drift of the key features is detected, so as to achieve the goal of not blindly retraining the model, but training the model on demand to adapt to the situation of feature drift, and adjusting the model in time to prevent the model performance from decaying, thereby effectively solving the problem of model decay after going online.

[0081] In some embodiments, see Figure 4 , Figure 4 An optional flow chart of the model performance improvement method provided in the embodiment of the present application, Figure 1 The illustrated S102 may also be implemented through S106 to S107, which will be described in conjunction with each step.

[0082] S106. Acquire new sample data within every predetermined time period in a time series.

[0083] In the embodiment of the present application, when the feature detection result indicates the drift of key features, the application data within every predetermined period can be determined as new sample data starting from the current moment, thereby obtaining N new sample data with time series features.

[0084] The new sample data acquired at every predetermined time period may also be parameter information of an object or relevant information of a user. The predetermined time period may be one day or one month. In the embodiment of the present application, there is no specific limitation on the length of the predetermined time period.

[0085] S107 . Train the initial sub-model based on the new sample data until the training condition is met, thereby obtaining the sub-model of the new sample data.

[0086] In an embodiment of the present application, the corresponding initial sub-model can be trained based on the new sample data obtained at predetermined intervals until the training condition is met and the sub-model corresponding to each new sample data is stopped. Since the new sample data obtained is arranged in time series, the N sub-models obtained based on the sample data training are also arranged in time series. N is an integer greater than 0. Among them, the initial sub-model can be an untrained XGBoost model.

[0087] In the embodiment of the present application, the corresponding sub-model is trained using the new sample data obtained at predetermined time intervals, so that the newly trained sub-model can match the latest sample data. Furthermore, the new model determined using the sub-model of the latest trip can also match the latest data, thereby improving the data matching of the new model.

[0088] In some embodiments, see Figure 5 , Figure 5 An optional flow chart of the model performance improvement method provided in the embodiment of the present application, Figure 1 The illustrated S103 can also be implemented through S108 to S109, which will be described in conjunction with each step.

[0089] S108, integrating the sub-models obtained in time sequence into the model to form an intermediate model.

[0090] In an embodiment of the present application, the sub-model trained based on each new sample data can be integrated into the original model to obtain an intermediate model.

[0091] S109: Determine the new model based on at least one of the number and performance information of the sub-models in the intermediate model.

[0092] In an embodiment of the present application, different types of screening processing can be performed on the N sub-models based on the number of N sub-models and at least one of the performance information, and the remaining sub-models after screening and the new sub-model obtained by new training can be integrated into the original model to obtain a new model with improved performance.

[0093] Exemplarily, the preset number may be determined to be 3, and the sub-model may be an XGBoost model. When the number of N sub-models is greater than 3, the XGBoost model with the earliest time sequence is eliminated, and the latest trained XGBoost sub-model is added to the sub-models in the time sequence.

[0094] In the embodiment of the present application, the sub-model of the new trip is integrated into the original model to obtain an intermediate model, and then based on at least one of the number of sub-models and performance information included in the intermediate model, a new model with improved performance is determined. This process can effectively combine the performance information of each sub-model and the number of sub-models to perform integrated learning of the model, thereby ensuring the integrated performance of the new model.

[0095] In some embodiments, see Figure 6 , Figure 6 An optional flow chart of the model performance improvement method provided in the embodiment of the present application, Figure 5 The illustrated S109 can also be implemented through S110 to S112, which will be described in conjunction with each step.

[0096] S110 , if the N sub-models reach a predetermined number, deleting the sub-model with the earliest timing in the intermediate model.

[0097] In the embodiment of the present application, the intermediate model includes N sub-models. If the number of the N sub-models reaches 3, the sub-model with the earliest time sequence among the N sub-models arranged in time sequence in the intermediate model is deleted.

[0098] In some other embodiments, if the quantity detection result indicates that the number of N sub-models reaches 3, the K sub-models with the earliest timing among the N sub-models arranged in time sequence in the intermediate model are deleted.

[0099] The number of the predetermined number may be 3, 4 or 10. In the embodiment of the present application, no specific limitation is imposed on the predetermined number.

[0100] S111. Train the initial sub-model based on sample data in a next predetermined time period to obtain a first new sub-model.

[0101] In an embodiment of the present application, after deleting the model with the earliest time sequence among the N sub-models, the initial sub-model can also be trained based on new sample data within a recent predetermined time period until the training conditions are met to obtain the first new sub-model.

[0102] In an embodiment of the present application, after deleting the K models with the earliest timing among the N sub-models, the initial sub-model can also be trained based on new sample data within the most recent K predetermined time periods until the training conditions are met, thereby obtaining K first new sub-models.

[0103] S112. Integrate the first new sub-model into the intermediate model to obtain the new model.

[0104] In an embodiment of the present application, the first new sub-model may be integrated into the intermediate model to obtain the new model with improved performance.

[0105] In the embodiment of the present application, the earliest sub-model formed in the intermediate model is deleted, and then the most recently formed new sub-model is integrated into the intermediate model, thereby obtaining a new model that best matches the latest data, thereby improving the performance of the new model.

[0106] In some embodiments, see Figure 7 , Figure 7 An optional flow chart of the model performance improvement method provided in the embodiment of the present application, Figure 5 The illustrated S109 can also be implemented through S113 to S114, which will be described in conjunction with each step.

[0107] S113: If the N sub-models do not reach the predetermined number, determine the performance information corresponding to each of the N sub-models.

[0108] In the embodiment of the present application, the intermediate model includes N sub-models. If the N sub-models do not reach the predetermined number, the performance scores corresponding to the N sub-models are determined.

[0109] In the embodiment of the present application, the model effect evaluation index detection can be performed on the N sub-models to determine the model's (Receiver Operating Characteristic, ROC) and (Area under the curve, AUC) scores. Then, the performance scores corresponding to the N sub-models are obtained.

[0110] S114, obtaining the new model based on the performance information respectively corresponding to the N sub-models and the second new sub-model integration; wherein the second new sub-model is trained on the initial sub-model based on new sample data within the next predetermined time period.

[0111] In the embodiment of the present application, the sub-model with the worst performance can be determined based on the performance scores corresponding to the N sub-models, and then the sub-model with the worst performance is deleted from the N sub-models, and the second sub-model is integrated into the intermediate sub-model to obtain a new model.

[0112] In some other embodiments, if the performance scores corresponding to the N sub-models respectively exceed the score threshold, the second sub-model is integrated into the intermediate sub-model to obtain a new model.

[0113] In the embodiment of the present application, the sub-models are integrated by soft voting. Exemplarily, the performance comparison between the new model finally obtained and an XGBoost sub-model trained with new sample data is as follows: sub-XGBoost model test results: AUC: 0.9, KS (Kolmogorov-Smirnov) value: 0.63; the new model integrated in this solution has the following test results: AUC: 0.92, KS: 0.66. It can be seen that the performance of the new model after ensemble learning is much better than that of a single sub-model.

[0114] In an embodiment of the present application, based on the performance information of each sub-model, the model with the worst performance can be deleted at an appropriate time, and then a newly formed new sub-model can be integrated into the intermediate model to obtain a new model that best matches the latest data, thereby improving the performance of the new model.

[0115] In some embodiments, see Figure 8 , Figure 8 An optional flow chart of the model performance improvement method provided in the embodiment of the present application, Figure 1 The illustrated S101 may also include S115 to S116, which will be described in conjunction with each step.

[0116] S115, performing importance calculation processing on multiple features included in the sample data respectively through a preset model to obtain the importance corresponding to each of the features.

[0117] In the embodiment of the present application, the importance corresponding to each feature in the sample data can be calculated by the SHAP analysis algorithm.

[0118] For example, the height feature in the user's relevant information may be used to calculate the corresponding importance through the SHAP analysis algorithm.

[0119] S116. Determine the features corresponding to the top M greatest importances as M key features.

[0120] In the embodiment of the present application, the features corresponding to the top N greatest importances may be M key features, where M is an integer greater than 0.

[0121] Exemplary, combined Fig. 9 , analyze the importance of model features. In this scenario, since the sub-model is XGBoost, we use XGBoost's total gain to calculate the model's feature importance and sort it. Then, we perform data drift detection and monitoring on the top 20 key features. The feature importance sorting results Fig. 9 As shown: Among them, feature X has the highest importance.

[0122] In some embodiments, see Fig.10 , Fig.10 An optional flow chart of the model performance improvement method provided for the embodiment of the present application will be described in conjunction with each step.

[0123] S201, detect and sort the importance of model features to select topN features.

[0124] Model feature importance detection: Before detecting the model performance degradation caused by data drift and concept drift, we must first detect which key features are most important to the model. Changes in these most important features will have a more significant impact on model performance, causing the model performance to decline significantly. There are many algorithmic techniques to measure model feature importance, such as the total gain of the XGBoost model and the model-independent SHAP value.

[0125] S202: Perform real-time drift monitoring on topN feature data.

[0126] Data concept drift detection: This step mainly involves continuously collecting application data used by the model while performing real-time reasoning after the model is launched, and comparing and calculating the collected application data with the important features obtained in the first step of the sample data before the model is launched. There are many comparison algorithms, such as KL divergence. If the KL divergence is within a reasonable range (based on the KL divergence of the sample data), it is considered that the important features have not drifted and the model continues to be used. If it is found that the key features have drifted, real-time data is collected and stored, and the model is prepared to be reintegrated and trained and updated.

[0127] S203: Whether feature data drift occurs.

[0128] Whether drift occurs is determined based on the KL divergence of each key feature.

[0129] S204: Collect the latest data and use the latest data to train a new sub-model.

[0130] Model retraining: To address the problem of key feature drift, we use the latest collected new sample data to train a sub-model, and then integrate the sub-model into the existing large model to participate in the reasoning of the entire model to improve the performance of the entire model.

[0131] S205: Whether the number of sub-models exceeds a preset number.

[0132] Check whether the number of sub-models currently formed exceeds the preset number.

[0133] S206. Eliminate the earliest sub-model.

[0134] If the number of sub-models currently formed exceeds the preset number, the earliest sub-model formed will be eliminated.

[0135] S207. Integrate the latest sub-model into the large model to participate in reasoning.

[0136] Model integration management: Assume that we set the number of sub-models to N. When the number of sub-models is less than or equal to N, the entire model determines the final result based on a comprehensive score based on the inference scores of these sub-models. When the number of sub-models is greater than N, we eliminate the oldest sub-model. This first-in-first-out algorithm for managing sub-models ensures that the N sub-models are the most suitable models for the latest data (the shortest time distance from the latest data), thereby ensuring that the model is optimal for the latest data.

[0137] S208, the integrated model uses algorithms such as voting to integrate the sub-model results to obtain the final inference result.

[0138] Real-time integrated reasoning of models: The reasoning results of each sub-model are integrated using algorithms such as voting to obtain the final reasoning result. The specific algorithm that makes the final reasoning result more accurate needs to be tested and adjusted according to the specific business.

[0139] S209: Whether to continue to detect data and update the model.

[0140] Continue to detect the KL divergence of the key features, and determine whether to update the model based on the KL divergence of the key features.

[0141] This solution adopts an integrated learning approach. When real-time data drift is detected, the latest data is collected and a sub-model is trained instead of the entire model. In this way, targeted training of a small model not only adapts to the latest data environment, but also saves the amount of training calculations (the amount of data is not as large as the full amount of data, and the computing power required for training is less than that required for training a large model). Provide algorithm and engineering personnel with a complete set of full pipeline solutions from data concept drift detection to integrated model management, training, and reasoning to solve and alleviate the problem of model performance degradation caused by data concept drift. The entire solution is released in the form of a toolkit of a python library, which integrates various relevant functional modules and APIs for algorithm engineers to call. Provide a complete set of pipeline sample codes for real scenarios, which makes it easy for algorithm engineers to understand and the entire system easier to operate and promote.

[0142] The purpose of the present invention is to use the data concept drift detection algorithm to perform drift detection on real-time data while inferring the model online, and trigger the integrated training process of the model when the data concept drift is detected, so as to achieve not blindly retraining the model, but training the model on demand to adapt to the data drift situation, and timely adjusting the model to prevent the model performance from decaying. In the integrated model, a first-in-first-out method is used to manage several integrated sub-models, so that the entire integrated model can be retrained on demand, and the model that best conforms to the most recent data distribution can be retained and adopted in the reasoning process, which is conducive to ensuring that the performance of the entire model is optimized in the latest data distribution.

[0143] See also Fig.11 , Fig.11 A schematic diagram of the structure of a model performance improvement device provided in an embodiment of the present application.

[0144] The embodiment of the present application also provides a model performance improvement device 800, including: a feature detection unit 801, a model training unit 802 and a model determination unit 803.

[0145] The feature detection unit 801 is used to determine the feature detection result based on the probability density distribution change between each key feature and the application features of the same type in the application data; wherein the application data is the data obtained after the model is launched; the key features are the M features with the greatest feature importance in the sample data; M is an integer greater than 0;

[0146] A model training unit 802 is used to train a corresponding sub-model based on new sample data acquired in time series when the feature detection result indicates the drift of the key feature;

[0147] The model determination unit 803 is used to determine a new model based on the sub-model and the model.

[0148] In an embodiment of the present application, the feature detection unit 801 in the model performance improvement device 800 is used to determine the relative entropy based on the probability density changes corresponding to each of the key features and the corresponding application features; wherein the relative entropy is used to characterize the probability density distribution changes between each key feature and the corresponding application feature; and the feature detection result is determined based on the relative entropy corresponding to each of the key features.

[0149] In the embodiment of the present application, the feature detection unit 801 in the model performance improvement device 800 is used to determine the feature detection result characterizing the drift of the key feature if the relative entropy corresponding to any of the key features is outside a preset range;

[0150] If the relative entropy corresponding to each of the key features is within the preset range, the feature detection result indicating that the key feature is normal is determined.

[0151] In an embodiment of the present application, the model training unit 802 in the model performance improvement device 800 is used to obtain the new sample data within each predetermined time period in a time sequence; train the initial sub-model based on the new sample data until the training conditions are met, and obtain the sub-model of the new sample data.

[0152] In an embodiment of the present application, the model determination unit 803 in the model performance improvement device 800 is used to integrate the sub-models obtained in time sequence into the model to form an intermediate model; and determine the new model based on at least one of the number and performance information of the sub-models in the intermediate model.

[0153] In an embodiment of the present application, the intermediate model includes: N sub-models; N is an integer greater than 0; the model determination unit 803 in the model performance improvement device 800 is used to delete the sub-model with the earliest time sequence in the intermediate model if the N sub-models reach a predetermined number; train the initial sub-model based on new sample data within the next predetermined time period to obtain a first new sub-model; integrate the first new sub-model into the intermediate model to obtain the new model.

[0154] In an embodiment of the present application, the intermediate model includes: N sub-models; N is an integer greater than 0; the model determination unit 803 in the model performance improvement device 800 is used to determine the performance information corresponding to the N sub-models if the N sub-models do not reach a predetermined number; based on the performance information corresponding to the N sub-models, and integrating the second new sub-model to obtain the new model; wherein the second new sub-model is based on the new sample data within the next predetermined time period to train the initial sub-model.

[0155] In an embodiment of the present application, the model performance improvement device 800 is used to calculate the importance of multiple features included in the sample data through a preset model to obtain the importance of each feature; determine the features corresponding to the top M largest importances as M key features; M is an integer greater than 0.

[0156] It should be noted that in the embodiment of the present application, if the above-mentioned model performance improvement method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a model performance improvement device (which can be a personal computer, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific combination of hardware and software.

[0157] Correspondingly, an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the steps in the above method are implemented.

[0158] Correspondingly, an embodiment of the present application provides an electronic device 900, including a memory 902 and a processor 901, wherein the memory 902 stores a computer program that can be executed on the processor 901, and the processor 901 implements the steps in the above method when executing the program.

[0159] It should be noted here that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.

[0160] It should be noted that Fig.12A hardware entity diagram of an electronic device provided in an embodiment of the present application, such as Fig.12 As shown, the hardware entity of the electronic device 900 includes: a processor 901 and a memory 902, wherein;

[0161] The processor 901 generally controls the overall operation of the electronic device 900 .

[0162] The memory 902 is configured to store instructions and applications executable by the processor 901, and can also cache data to be processed or processed by the processor 901 and various modules in the electronic device 900 (for example, image data, audio data, voice communication data, and video communication data), which can be implemented through flash memory (FLASH) or random access memory (Random Access Memory, RAM).

[0163] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned sequence numbers of the embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.

[0164] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0165] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be electrical, mechanical or other forms.

[0166] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0167] In addition, all functional units in the embodiments of the present application may be integrated into one processing unit, or each unit may be a separate unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0168] A person skilled in the art can understand that all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: a mobile storage device, a read-only memory (ROM), a magnetic disk or an optical disk, and other media that can store program codes.

[0169] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application can be essentially or partly embodied in the form of a software product that contributes to the relevant technology. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a mobile storage device, a ROM, a magnetic disk, or an optical disk.

[0170] The above is only an implementation method of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A method for improving model performance, It is characterized in that include: Determine the feature detection result based on the change in probability density distribution between each key feature and similar application features in the application data; wherein the application data is data acquired after the model is launched; the key features are the M features with the greatest feature importance in the sample data; M is an integer greater than 0; In the case where the feature detection result represents the drift of the key feature, a corresponding sub-model is obtained by training based on new sample data acquired in time series; A new model is determined based on the sub-models and the model.

2. The model performance improvement method according to claim 1, It is characterized in that The determining of the feature detection result based on the probability density distribution change between each key feature and similar application features in the application data includes: Determine the relative entropy based on the probability density changes corresponding to each of the key features and the corresponding application features; wherein the relative entropy is used to characterize the probability density distribution changes between each key feature and the corresponding application feature; The feature detection result is determined based on the relative entropy corresponding to each of the key features.

3. The model performance improvement method according to claim 2, It is characterized in that The determining the feature detection result based on the relative entropy corresponding to each of the key features includes one of the following: If the relative entropy corresponding to any of the key features is outside a preset range, determining the feature detection result characterizing the drift of the key feature; If the relative entropy corresponding to each of the key features is within the preset range, the feature detection result indicating that the key feature is normal is determined.

4. The method for improving model performance according to claim 1, It is characterized in that The corresponding sub-model is obtained by training the new sample data acquired in time series, including: Acquire the new sample data within each predetermined time period in a time sequence; The initial sub-model is trained based on the new sample data until the training condition is met, thereby obtaining the sub-model of the new sample data.

5. The method for improving model performance according to any one of claims 1 to 4, It is characterized in that The determining of a new model based on the sub-model and the model comprises: Integrating the sub-models obtained in time sequence into the model to form an intermediate model; The new model is determined based on at least one of the number and performance information of the sub-models in the intermediate model.

6. The method for improving model performance according to claim 5, It is characterized in that The intermediate model includes: N sub-models; N is an integer greater than 0; The determining of the new model based on at least one of the number and performance information of the sub-models in the intermediate model comprises: If the N sub-models reach a predetermined number, deleting the sub-model with the earliest time sequence in the intermediate model; Training the initial sub-model based on new sample data within a next predetermined period of time to obtain a first new sub-model; The first new sub-model is integrated into the intermediate model to obtain the new model.

7. The method for improving model performance according to claim 5, It is characterized in that The intermediate model includes: N sub-models; N is an integer greater than 0; The determining of the new model based on at least one of the number and performance information of the sub-models in the intermediate model comprises: If the N sub-models do not reach a predetermined number, determining the performance information respectively corresponding to the N sub-models; The new model is obtained based on the performance information respectively corresponding to the N sub-models and the integration of the second new sub-model; wherein the second new sub-model is trained on the initial sub-model based on the new sample data within the next predetermined time period.

8. The method for improving model performance according to any one of claims 1 to 4, It is characterized in that Before determining the feature detection result based on the probability density distribution change between each key feature and similar application features in the application data, the method further includes: The importance of each feature in the sample data is calculated by using a preset model to obtain the importance of each feature; The features corresponding to the top M greatest importances are determined as the M key features.

9. A model performance improvement device, It is characterized in that include: A feature detection unit, used to determine a feature detection result based on a change in probability density distribution between each key feature and similar application features in the application data; wherein the application data is data acquired after the model is launched; the key features are M features with the greatest feature importance in the sample data; M is an integer greater than 0; A model training unit, configured to obtain a corresponding sub-model by training based on new sample data acquired in time series when the feature detection result represents the drift of the key feature; A model determination unit is used to determine a new model based on the sub-model and the model.

10. An electronic device, It is characterized in that The method comprises a memory and a processor, wherein the memory stores a computer program executable on the processor, and the processor implements the steps of the method according to any one of claims 1 to 8 when executing the computer program.

11. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the steps in the method according to any one of claims 1 to 8 are implemented.