Model Adaptive Training Method, Device, Equipment, Medium and Program Product
The self-adaptive training method addresses AI model performance degradation by using value models and semi-supervised learning to update models with minimal human intervention, enhancing stability and reducing annotation needs.
Patent Information
- Application Number
- CN202110945461.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-17
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-08-17
AI Technical Summary
When existing AI models face concept drift caused by environmental changes, their performance gradually decreases and require frequent human intervention to retrain, resulting in large workloads and difficulty in adapting to different types of concept drifts.
By obtaining the concept drift value of the running data detection, using the active learning mechanism and the semi-supervised learning method, the running data is allocated to the labeled and labelless data sets, and adaptive training is performed, and the model weight is adjusted in combination with the integrated inference mechanism to reduce the amount of manual annotations and improve the model stability.
With as little human intervention as possible, adaptive training of AI models is implemented, reducing the amount of manual annotations and improving the performance stability of the model, and being able to adapt to various scenario changes.
Smart Images

Figure CN113657501B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer data processing, and in particular, to a method, device, equipment, medium, and program product for adaptive training of a digital model. Background Art
[0002] Currently, for an AI (Artificial Intelligence) model, historical data that already exists is generally used for training and then the model is put into production service to make predictions on new actual operation data.
[0003] However, over time, it is inevitable that the sample distribution changes due to environmental changes, and this phenomenon is called concept drift. At this time, the performance of the AI model will gradually decrease. Therefore, we need to regularly retrain the AI model with the latest data to continuously update the AI model and monitor the model performance in real time to ensure the stable performance of the AI model put into production.
[0004] Therefore, the subsequent performance maintenance of the AI model put into production brings a long-term and heavy workload to algorithm engineers and operation and maintenance personnel. Therefore, how to enable the AI model to perform adaptive training with as little human intervention as possible has become a technical problem to be solved urgently. Summary of the Invention
[0005] The present application provides a method, device, equipment, medium, and program product for adaptive training of a model, which solves the technical problem of how to enable the AI model to perform adaptive training with as little human intervention as possible.
[0006] In a first aspect, the present application provides a method for adaptive training of a model, including:
[0007] Obtain each operation data of the original model during actual operation, and detect a first concept drift value of the original model during actual operation according to the operation data;
[0008] According to the value model and the first concept drift value, allocate each operation data to a labeled data set and / or an unlabeled data set respectively;
[0009] Judge whether the data volume of the labeled data set is greater than or equal to a preset threshold;
[0010] If so, use the adaptive training model to adaptively train the original model according to the labeled data set and the unlabeled data set to determine a new model after training, and the second concept drift value of the new model is less than the first concept drift value.
[0011] In a possible design, each piece of operation data is respectively allocated to a labeled data set and / or an unlabeled data set according to a value model and a first concept drift value, including:
[0012] Determine the comprehensive value of each piece of operation data according to the value model;
[0013] Adjust the data accumulation speed of the labeled data set according to the first concept drift value and the comprehensive value;
[0014] Allocate each piece of operation data to a labeled data set and / or an unlabeled data set according to the data accumulation speed and the comprehensive value.
[0015] In a possible design, each piece of operation data is respectively allocated to a labeled data set and / or an unlabeled data set according to the data accumulation speed and the comprehensive value, including:
[0016] Screen out the data to be labeled from each piece of operation data according to the data accumulation speed and the comprehensive value;
[0017] Send the data to be labeled to the user side for labeling to determine the labeled data, and add the remaining operation data to the unlabeled data set;
[0018] Receive the labeled data returned by the user side, and add the labeled data to the labeled data set.
[0019] In a possible design, adjusting the data accumulation speed of the labeled data set according to the first concept drift value and the comprehensive value includes:
[0020] Determine the sorting sequence of each piece of operation data according to the comprehensive value;
[0021] When the first concept drift value is less than or equal to the warning threshold value, select the first M pieces of operation data in the sorting sequence as the data to be labeled.
[0022] In a possible design, after determining the sorting sequence of each piece of operation data according to the comprehensive value, it further includes:
[0023] When the first concept drift value is greater than or equal to the trigger threshold value, select the first N pieces of operation data in the sorting sequence as the data to be labeled, and the warning threshold value is less than the trigger threshold value.
[0024] In a possible design, after determining the sorting sequence of each piece of operation data according to the comprehensive value, it further includes:
[0025] When the first concept drift value is greater than the warning threshold value and less than the trigger threshold value, select the first K pieces of operation data in the sorting sequence as the data to be labeled, and there is a preset corresponding relationship between K, M, and N.
[0026] In a possible design, an adaptive training model is used to adaptively train an original model according to a labeled data set and an unlabeled data set to determine a new trained model, including:
[0027] Sampling the labeled data set and the unlabeled data set through a preset sampling model to determine a training sample set;
[0028] Using a semi-supervised training model to perform semi-supervised training on the original model according to the training sample set and multiple learning rates, and then determining multiple sub-models, where the sub-models correspond to the learning rates;
[0029] Combining each sub-model into a new model through an integrated weight value.
[0030] In a possible design, before combining each sub-model into a new model through an integrated weight value, it further includes:
[0031] Using a dynamic update algorithm to determine an updated integrated weight value according to each learning rate and a first concept drift value.
[0032] In a possible design, using a dynamic update algorithm to determine an updated integrated weight value according to each learning rate and a first concept drift value, including:
[0033] Initializing the integrated weight value corresponding to each sub-model;
[0034] When the first concept drift value is less than or equal to a warning threshold value, determining an updated integrated weight value according to a first update model, a preset update factor, and each learning rate.
[0035] In a possible design, after initializing the integrated weight value corresponding to each sub-model, it further includes:
[0036] When the first concept drift value is greater than or equal to a trigger threshold value, determining an updated integrated weight value according to a second update model, a preset update factor, and each learning rate.
[0037] In a possible design, after initializing the integrated weight value corresponding to each sub-model, it further includes:
[0038] When the first concept drift value is greater than the warning threshold value and less than the trigger threshold value, determining an updated integrated weight value according to a third update model, a preset update factor, and each learning rate.
[0039] In a possible design, after determining the updated integrated weight value, it further includes:
[0040] Normalize the integrated weight values according to a preset normalization model to determine the respective target integrated weight values;
[0041] Correspondingly, combine each sub-model into a new model through the integrated weight values, including:
[0042] Combine each sub-model into a new model through the respective target integrated weight values.
[0043] In a second aspect, the present application provides a model adaptive training device, including:
[0044] An acquisition module, configured to acquire each operation data of the original model during actual operation;
[0045] A processing module, configured to:
[0046] Detect a first concept drift value of the original model during actual operation according to the operation data;
[0047] Allocate each operation data to a labeled data set and / or an unlabeled data set respectively according to the value model and the first concept drift value;
[0048] Judge whether the data volume of the labeled data set is greater than or equal to a preset threshold;
[0049] If so, use the adaptive training model to adaptively train the original model according to the labeled data set and the unlabeled data set to determine a new trained model, and the second concept drift value of the new model is less than the first concept drift value.
[0050] In a possible design, the processing module is configured to:
[0051] Determine the comprehensive value of each operation data according to the value model;
[0052] Adjust the data accumulation speed of the labeled data set according to the first concept drift value and the comprehensive value;
[0053] Allocate each operation data to a labeled data set and / or an unlabeled data set respectively according to the data accumulation speed and the comprehensive value.
[0054] In a possible design, the processing module is configured to:
[0055] Screen out the data to be labeled from each operation data according to the data accumulation speed and the comprehensive value;
[0056] Send the data to be labeled to the user side for labeling to determine the labeled data, and add the remaining operation data to the unlabeled data set;
[0057] The acquisition module is further configured to receive the labeled data returned by the user side;
[0058] The processing module is also used to add the labeled data to the labeled dataset.
[0059] In a possible design, the processing module is used to:
[0060] Determine the sorting sequence of each piece of operation data according to the comprehensive value;
[0061] When the first concept drift value is less than or equal to the warning threshold value, select the first M pieces of operation data in the sorting sequence as the data to be labeled.
[0062] In a possible design, the processing module is used to:
[0063] When the first concept drift value is greater than or equal to the trigger threshold value, select the first N pieces of operation data in the sorting sequence as the data to be labeled, and the warning threshold value is less than the trigger threshold value.
[0064] In a possible design, the processing module is used to:
[0065] When the first concept drift value is greater than the warning threshold value and less than the trigger threshold value, select the first K pieces of operation data in the sorting sequence as the data to be labeled, and there is a preset corresponding relationship between K, M, and N.
[0066] In a possible design, the processing module is used to:
[0067] Sample the labeled dataset and the unlabeled dataset through a preset sampling model to determine the training sample set;
[0068] Use a semi-supervised training model to perform semi-supervised training on the original model according to the training sample set and multiple learning rates, and then determine multiple sub-models, and the sub-models correspond to the learning rates;
[0069] Combine each sub-model into a new model through the integrated weight value.
[0070] In a possible design, the processing module is used to:
[0071] Use a dynamic update algorithm to determine the updated integrated weight value according to each learning rate and the first concept drift value.
[0072] In a possible design, the processing module is used to:
[0073] Initialize the integrated weight value corresponding to each sub-model;
[0074] When the first concept drift value is less than or equal to the warning threshold value, determine the updated integrated weight value according to the first update model, the preset update factor, and each learning rate.
[0075] In a possible design, a processing module is configured to:
[0076] When the first concept drift value is greater than or equal to the trigger threshold value, determine the updated integrated weight value according to the second update model, the preset update factor, and each learning rate.
[0077] In a possible design, a processing module is configured to:
[0078] When the first concept drift value is greater than the warning threshold value and less than the trigger threshold value, determine the updated integrated weight value according to the third update model, the preset update factor, and each learning rate.
[0079] In a possible design, the processing module is further configured to:
[0080] Normalize the integrated weight value according to the preset normalization model to determine each target integrated weight value; and combine each sub-model into a new model through each target integrated weight value.
[0081] In a third aspect, the present application provides an electronic device, including:
[0082] A memory for storing program instructions;
[0083] A processor for calling and executing the program instructions in the memory and performing any possible model adaptive training method provided in the first aspect.
[0084] In a fourth aspect, the present application provides a storage medium, in which a computer program is stored, and the computer program is used to perform any possible model adaptive training method provided in the first aspect.
[0085] In a fifth aspect, the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements any possible model adaptive training method provided in the first aspect.
[0086] The present application provides a model adaptive training method, apparatus, device, medium and program product. By obtaining various operation data of the original model during actual operation and detecting the first concept drift value of the original model during actual operation according to the operation data; then, according to the value model and the first concept drift value, each operation data is respectively assigned to a labeled data set and / or an unlabeled data set, and then it is determined whether the data volume of the labeled data set is greater than or equal to a preset threshold. If so, an adaptive training model is used to adaptively train the original model according to the labeled data set and the unlabeled data set to determine a new trained model, and the second concept drift value of the new model is less than the first concept drift value. The technical problem of how to enable an AI model to perform adaptive training with as little human intervention as possible is solved. The technical effect of reducing the amount of manual annotation required by developers when updating the model and improving the performance stability of the model by integrating multiple models is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0087] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the prior art description. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0088] Figure 1 It is a schematic diagram of the scenario of model adaptive training provided by the present application;
[0089] Figure 2 It is a schematic flowchart of a model adaptive training method provided by the present application;
[0090] Figure 3 It is a schematic flowchart of another model adaptive training method provided by the present application;
[0091] Figure 4 It is a schematic flowchart of yet another model adaptive training method provided by the present application;
[0092] Figure 5 It is a schematic diagram of updating the integrated weight value provided by an embodiment of the present application;
[0093] Figure 6 It is a graph of the experimental results of online model update for online learning during sudden concept drift provided by an embodiment of the present application;
[0094] Figure 7 It is a graph of the experimental results of model adaptive update for active learning of sudden concept drift provided by an embodiment of the present application;
[0095] Figure 8Experimental result graph of online model update for incremental concept drift in the embodiments of this application;
[0096] Figure 9 Experimental result graph of model adaptive update for active learning of incremental concept drift in the embodiments of this application;
[0097] Figure 10 Structural schematic diagram of a model adaptive training device provided by this application;
[0098] Figure 11 Structural schematic diagram of an electronic device provided by this application. Detailed implementation manners
[0099] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only some of the embodiments of this application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application without creative efforts, including but not limited to combinations of multiple embodiments, fall within the scope of protection of this application.
[0100] Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification and claims of this application and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these process, method, product, or device.
[0101] In current artificial intelligence learning systems, the general training method for a model is: training with a preset machine learning algorithm on a given data set to construct a model applicable to a certain scenario, and putting this model online for application in actual tasks.
[0102] Obviously, this machine learning model training method is essentially not suitable for processing concept drift data.
[0103] To ensure the prediction ability of online models (i.e., models that have been launched and applied), it is necessary to correct for the phenomenon of concept drift in order to improve or enhance the online performance of the models (i.e., reduce the difference between the results of the model in actual application and those during modeling and training). Usually, the model needs to be updated (fine-tuned) at regular intervals using the latest collected data, and there are various methods available, such as retraining, online training, incremental learning, etc.
[0104] However, all of the above-mentioned methods require obtaining labeled data (i.e., data that has been manually annotated) to perform model training and updates. This presents the following drawbacks:
[0105] 1. In an actual online environment, it is often very difficult to obtain labeled data, which requires the online system to provide a timely feedback mechanism or requires manual data annotation. Either way, it requires a large amount of additional economic and time costs. If the latest labeled data cannot be obtained in a timely manner, the online model may face a continuous decline in prediction ability; and even if labeled data is obtained, if the data volume is small, the existing model update methods have limited improvement on the model's prediction ability, and there is also a risk of a decline in model performance.
[0106] 2. Since concept drift can be divided into: incremental drift, gradual drift, sudden drift, cyclic drift, and their combinations. Therefore, newly obtained models through methods such as retraining, online training, and incremental learning are difficult to adapt to different types of concept drift.
[0107] To solve the above problems, the inventive concept of this application is:
[0108] First, in combination with an active learning mechanism: select the most valuable actual operation data for annotation, and use semi-supervised learning methods to make full use of a large amount of unlabeled data to automatically update the online model and ensure the application effect of the online model.
[0109] Then, in combination with an ensemble inference mechanism: adopt different learning rates, online learn multiple models, and control the weights of each model using the degree of data concept drift to achieve self-adaptation to concept drift and improve the stability and performance of the model.
[0110] Compared with traditional methods, this unlabeled data is widespread and easy to collect, saving economic, human, and time costs; through a small amount of data annotation, the prediction effect of the model can be greatly improved; and the learning method of the ensemble inference mechanism can enhance the stability of the model in the face of various scenario changes.
[0111] Figure 1 This is a schematic diagram of the scenario for model adaptive training provided by this application. As Figure 1As shown, the running data generated during the actual task or business operation, i.e., the online data stream 101, is transmitted to the concept drift module 102 and the unlabeled data pool 103 respectively. The concept drift value calculated by the concept drift module 102 is given to the active learning module 104 as a reference index for the amount of manually labeled data.
[0112] The active learning model 104 detects the comprehensive value of each data from the unlabeled data pool 103, and evaluates the comprehensive value of each data by indicators such as representativeness, diversity, and uncertainty. The top N data with larger comprehensive value are extracted as manually labeled data for labeling. The specific size of N is determined with reference to the concept drift value of the current model. The data after manual labeling is put into the labeled data pool 105. After the labeled data pool 105 is full, the update training of the current model is triggered.
[0113] The data sampling module 106 collects training data from the unlabeled data pool 103 and the labeled data pool 105, and sends it to the model online update module 107 for semi-supervised learning training. And each model obtained after training with different learning rates is stored in the model library 108.
[0114] Then, the integration weights of each model in the model library are updated through the concept drift value to obtain the updated integration weights. Then, multiple models are integrated through each integration weight to obtain the finally trained new model. The new model is used to process each running data in the online data stream 101 to obtain the inference result or prediction result for output.
[0115] To facilitate the understanding of the specific implementation steps in the above scenario, the following will introduce the detailed steps of the model adaptive training method provided in this application in combination with the accompanying drawings.
[0116] Figure 2 It is a schematic flowchart of a model adaptive training method provided in this application. As Figure 2 shown, the specific steps of this model adaptive training method include:
[0117] S201. Obtain each running data of the original model during actual operation, and detect the first concept drift value of the original model during actual operation according to the running data.
[0118] In this step, the first concept drift value is used for the magnitude of the distribution difference between the first processing result and the second processing result. Among them, the first processing result represents the processing result obtained by the original model when processing the running data of each actual business or service during actual operation; the second processing result represents the processing result of the element model on the training sample data during modeling training.
[0119] In this embodiment, the actual operation of the original model is also referred to as model go-live, and the operation data that needs to be processed by the actual task or business is also referred to as online data. The original model can also be called the online model.
[0120] It should be noted that there are many calculation methods for the concept drift value, and those skilled in the art can select according to the needs of specific scenarios, and this application does not make limitations.
[0121] S202. According to the value model and the first concept drift value, allocate each piece of operation data to the labeled data set and / or the unlabeled data set respectively.
[0122] In this step, first, determine the comprehensive value of each piece of operation data according to the value model; then, adjust the data accumulation speed of the labeled data set according to the first concept drift value and the comprehensive value; next, according to the data accumulation speed and the comprehensive value, allocate each piece of operation data to the labeled data set and / or the unlabeled data set respectively.
[0123] Specifically, through the active learning mechanism, use the value model to calculate the comprehensive value of each piece of operation data (or called online data). The composition indicators of the comprehensive value include: indicators such as the representativeness, diversity, and uncertainty of the operation data. Then, synthesize each indicator according to the preset joint model, such as multiplying each indicator by the corresponding weight and then summing, to calculate the comprehensive value of each piece of operation data.
[0124] It should be noted that indicators such as representativeness, diversity, and uncertainty can refer to the data in the existing unlabeled data set and the labeled data set that has been manually labeled.
[0125] For example, if the operation data is similar to or the same as the labeled data that has been manually labeled, then its representativeness, diversity, and uncertainty will decrease, and at the same time, it means that this piece of operation data does not need to be re-allocated to the labeled data set, that is, there is no need to manually label this data again, so that the workload of manual labeling can be reduced.
[0126] In this embodiment, the data accumulation speed of the labeled data set (such as Figure 1 the labeled data pool 105 therein) can be adjusted by selecting different numbers of operation data for manual labeling. That is, when selecting the operation data ranked in the top N in terms of comprehensive value for manual labeling, different concept drift values correspond to different N values, that is, N is a function value of the concept drift value.
[0127] After selecting the top N operation data that need to be manually labeled, allocate the remaining operation data to the unlabeled data set.
[0128] After manually annotating N pieces of operation data, the labeled data is added to the labeled dataset. When the amount of data in the labeled dataset reaches a preset threshold, an update training of the original model is triggered.
[0129] It should be noted that after each update, the labeled dataset can be cleared, so that the preset threshold is the size of the labeled dataset.
[0130] It can be understood that if the labeled dataset is not cleared, the preset threshold represents the difference in the amount of data in the labeled dataset between the previous update training and before the current update training.
[0131] S203. Determine whether the amount of data in the labeled dataset is greater than or equal to the preset threshold.
[0132] In this step, the implementation method of the labeled dataset includes a data stack constructed according to the first-in, first-out principle. The size of this data stack is the preset threshold. When the data stack is full, an adaptive training will be triggered.
[0133] S204. If so, use the adaptive training model to adaptively train the original model according to the labeled dataset and the unlabeled dataset to determine the new trained model.
[0134] In this step, the second concept drift value of the new model is less than the first concept drift value.
[0135] In this embodiment, through a preset sampling model, the labeled dataset and the unlabeled dataset are sampled to determine the training sample set;
[0136] Use the semi-supervised training model to semi-supervise the original model according to the training sample set and multiple learning rates to determine multiple sub-models, and the sub-models correspond to the learning rates;
[0137] Combine each sub-model into a new model through the integrated weight value.
[0138] It should be noted that the semi-supervised training mode can greatly reduce the amount of data that needs to be manually annotated. Moreover, by setting different learning rates, it can be ensured that the trained models can output a relatively reasonable prediction result or inference result under different actual environmental changes, that is, through the integrated inference method, the stability of the model is guaranteed, enabling it to cope with a wide range of data fluctuations.
[0139] This embodiment provides a model adaptive training method. By obtaining various running data of the original model during actual operation and detecting the first concept drift value of the original model during actual operation according to the running data; then, according to the value model and the first concept drift value, each piece of running data is respectively assigned to a labeled data set and / or an unlabeled data set, and then it is determined whether the data volume of the labeled data set is greater than or equal to a preset threshold. If so, an adaptive training model is used to adaptively train the original model according to the labeled data set and the unlabeled data set to determine the trained new model, and the second concept drift value of the new model is less than the first concept drift value. This solves the technical problem of how to enable an AI model to perform adaptive training with as little human intervention as possible. It achieves the technical effects of reducing the amount of manual annotation required by developers when updating the model and improving the performance stability of the model by integrating multiple models.
[0140] For ease of understanding, the following further illustrates various implementation manners of S202 and S204 through Figure 3 and Figure 5 the embodiments shown.
[0141] Figure 3 It is a flowchart of another model adaptive training method provided by this application. As Figure 3 shown, the specific steps of this model adaptive training method include:
[0142] S301. Obtain various running data of the original model during actual operation.
[0143] In this embodiment, the running data of the online data stream 101 as Figure 1 shown is collected and stored in the temporary data pool corresponding to the unlabeled data set. The temporary data pool is set to a fixed size. When the temporary data pool is full of data, it enters S303; otherwise, it continues to collect online data, that is, running data.
[0144] S302. Detect the first concept drift value of the original model during actual operation according to the running data.
[0145] In this step, the first concept drift value is determined by the distribution difference between the prediction result or inference result when the original model processes the actual running data and the result of processing the training data. For example, the energy distance is used as the value of the first concept drift value.
[0146] It should be noted that S301 and S302 can be executed in parallel without a requirement for a sequence.
[0147] S303. Determine the comprehensive value of each piece of running data according to the value model.
[0148] In this step, the comprehensive value of the temporary samples is calculated by using the active learning mechanism, i.e., the value model.
[0149] In this embodiment, specifically, by combining the data in the unlabeled dataset, the data in the labeled dataset, and the data in the temporary data pool, the indicators such as representativeness, diversity, and uncertainty of the data in the temporary data pool are calculated, and multiple indicators are combined to calculate the annotation value, i.e., the comprehensive value, of each data.
[0150] It should be noted that for the specific implementation manner of the value model of the active learning mechanism, those skilled in the art can select according to the actual situation, and this application does not make a limitation.
[0151] S304. Adjust the data accumulation speed of the labeled dataset according to the first concept drift value and the comprehensive value.
[0152] In this step, first, the sorting sequence of each running data is determined according to the comprehensive value. For example, each running data is arranged in descending order of the comprehensive value.
[0153] When the first concept drift value is less than or equal to the warning threshold value, the first M running data in the sorting sequence are selected as the data to be labeled.
[0154] In this embodiment, the value of M is shown in formula (1):
[0155] M = Min label , (β ≤ β warning ) (1)
[0156] where, Min label is the preset minimum data volume, β is the first concept drift value, and β warning is the warning threshold value.
[0157] In a possible design, when the first concept drift value is greater than or equal to the trigger threshold value, the first N running data in the sorting sequence are selected as the data to be labeled, and the warning threshold value is less than the trigger threshold value.
[0158] Specifically, the value of N is shown in formula (2):
[0159] N = Max label , (β ≥ β detected ) (2)
[0160] where, Max label is the preset maximum data volume, and β detected is the trigger threshold value.
[0161] In a possible design, when the first concept drift value is greater than the warning threshold and less than the trigger threshold, the top K pieces of operation data in the sorting sequence are selected as the data to be labeled, and there is a preset corresponding relationship between K, M, and N.
[0162] Specifically, the value of K is shown in formula (3):
[0163] K = Min label +(β - β warning / β detected -β warning ) * Max label , (β warning <β<β detected ) (3)
[0164] S305. Screen out the data to be labeled from each piece of operation data according to the data accumulation speed and comprehensive value.
[0165] S306. Send the data to be labeled to the user side for labeling to determine the labeled data, and add the remaining operation data to the unlabeled data set.
[0166] S307. Receive the labeled data returned by the user side and add the labeled data to the labeled data set.
[0167] For steps S305 to S307, in this embodiment, the data not selected in the temporary data pool is pushed to the unlabeled data set; the selected data, after being labeled, that is, manually labeled, is pushed to the labeled data set. If the data volume of the labeled data set is greater than the trigger threshold, that is, the preset threshold, then a model adaptive update training is triggered. After the training is completed, the labeled data set is cleared.
[0168] S308. When the data volume of the labeled data set is greater than or equal to the preset threshold, use the adaptive training model to adaptively train the original model according to the labeled data set and the unlabeled data set to determine the trained new model.
[0169] In this step, the second concept drift value of the new model is less than the first concept drift value of the original model.
[0170] For the specific implementation method of this step, refer to the relevant step introductions in the embodiments shown in S204 and Figure 4 the relevant steps in the embodiments shown.
[0171] This embodiment provides a model adaptive training method. By obtaining various running data of the original model during actual operation and detecting the first concept drift value of the original model during actual operation according to the running data; then, according to the value model and the first concept drift value, each piece of running data is respectively assigned to a labeled data set and / or an unlabeled data set, and then it is judged whether the data volume of the labeled data set is greater than or equal to a preset threshold. If so, an adaptive training model is used to adaptively train the original model according to the labeled data set and the unlabeled data set to determine a new trained model, and the second concept drift value of this new model is less than the first concept drift value. This solves the technical problem of how to enable an AI model to perform adaptive training with as little human intervention as possible. It achieves the technical effects of reducing the amount of manual annotation required by developers when updating the model and improving the performance stability of the model by integrating multiple models.
[0172] For ease of understanding, the specific implementation manners of S204 and S308 are further explained below.
[0173] Figure 4 It is a schematic flowchart of another model adaptive training method provided by this application. As Figure 4 shown, the specific steps of this model adaptive training method include:
[0174] S401. Obtain various running data of the original model during actual operation and detect the first concept drift value of the original model during actual operation according to the running data.
[0175] S402. According to the value model and the first concept drift value, respectively assign each piece of running data to a labeled data set and / or an unlabeled data set.
[0176] For the detailed implementation manners and principles of steps S401 to S402, reference can be made to the relevant steps of the embodiments shown in Figure 2 and Figure 3 and will not be elaborated here.
[0177] S403. Sample the labeled data set and the unlabeled data set through a preset sampling model to determine a training sample set.
[0178] In this embodiment, the bootstrap sampling model is used as the preset sampling model to sample the labeled data set and the unlabeled data set to obtain a training sample set for adaptive training.
[0179] In this embodiment, in order to adapt to the impacts brought by various concept drifts such as adaptive progressive concept drift and sudden concept drift, the present application introduces integrated reasoning (i.e., combining multiple models according to preset weights), and dynamically updates the integrated weights. Generally, the integrated weights are updated once every preset time.
[0180] S404. Use the semi-supervised training model to perform semi-supervised training on the original model according to the training sample set and multiple learning rates, and then determine multiple sub-models.
[0181] In this embodiment, k different learning rates are adopted, which are respectively denoted as λ1, λ2…λ k , and λ1 < λ2… < λ k . Through the preset semi-supervised training model, semi-supervised training is performed according to the data in the training sample set to obtain k different models, namely the above-mentioned sub-models. The integrated weight of each sub-model for integrated reasoning is W i , then the inference result of the input data X is shown in formula (4):
[0182]
[0183] where k is an integer greater than 2 to utilize the advantages of integrated learning.
[0184] In a possible design, k is less than 10 to avoid the problem of excessive resource consumption in the inference process.
[0185] In a possible design, λ1, λ2…λ k can be set as a geometric sequence. The user first sets the minimum learning rate λ1 and the maximum learning rate λ k values according to their own experience, and then calculates other learning rates according to the rules of the geometric sequence.
[0186] S405. Use the dynamic update algorithm to determine the updated integrated weight value according to each learning rate and the first concept drift value.
[0187] Figure 5 is a schematic diagram for updating the integrated weight value provided by the embodiment of the present application. As Figure 5 shown, after initializing the weight W, that is, initializing the integrated weight value, according to the concept drift value, that is, the drift monitoring calculation result, it is judged whether the drift warning condition is satisfied. For example, if the concept drift value is greater than or equal to the warning threshold value, if not, the weight of the slow learning model is increased. If so, it is continued to judge whether the drift trigger condition is satisfied. For example, if the concept drift value is greater than or equal to the trigger threshold value, if not, weight balancing is performed, such as setting each integrated weight value to the same value. If so, the weight of the fast learning model is increased. Finally, the normalization operation of the integrated weight value, that is, weight normalization, is performed.
[0188] Specifically:
[0189] First, initialize the integrated weight values corresponding to each sub-model.
[0190] For example, initialize the integrated weight of each sub-model as: W i = W init = 1 / k.
[0191] When the first concept drift value is less than or equal to the warning threshold value, determine the updated integrated weight value according to the first update model, the preset update factor, and each learning rate.
[0192] Specifically, let the update factor be p, and p ∈ [0, 1), and the mathematical expression of the first update model is as shown in formula (5):
[0193]
[0194] where, W n is the updated integrated weight value, β warning is the warning threshold value, β is the first concept drift value, W i0 is the initial integrated weight value.
[0195] log(λ i / λ1) represents calculating the logarithm value of λ i / λ1 with base 10.
[0196] In a possible design, when the first concept drift value is greater than or equal to the trigger threshold value, determine the updated integrated weight value according to the second update model, the preset update factor, and each learning rate.
[0197] Specifically, the mathematical expression of the second update model is as shown in formula (6):
[0198]
[0199] where, β detected is the trigger threshold value.
[0200] In a possible design, when the first concept drift value is greater than the warning threshold value and less than the trigger threshold value, determine the updated integrated weight value according to the third update model, the preset update factor, and each learning rate.
[0201] Specifically, the mathematical expression of the third update model is as shown in formula (7):
[0202]
[0203] When concept drift occurs, increase the weights of the fast learning model to cope with the impacts brought by sudden concept drift and the like; when in the concept drift warning state, gradually move the model weights closer to the initialization state to cope with the impacts brought by gradual concept drift and the like; when the concept drift value is less than the warning threshold, gradually increase the weights of the slow learning rate model.
[0204] S406. According to a preset normalization model, perform a normalization operation on the integrated weight values to determine each target integrated weight value.
[0205] In this step, the preset normalization model is shown in formula (8):
[0206]
[0207] S407. Combine each sub-model into a new model through each target integrated weight value.
[0208] This embodiment provides a model adaptive training method. By obtaining each running data of the original model during actual operation and detecting the first concept drift value of the original model during actual operation according to the running data; then, according to the value model and the first concept drift value, respectively allocate each running data to the labeled data set and / or the unlabeled data set, and then determine whether the data volume of the labeled data set is greater than or equal to a preset threshold. If so, use the adaptive training model to adaptively train the original model according to the labeled data set and the unlabeled data set to determine the trained new model, and the second concept drift value of the new model is less than the first concept drift value. This solves the technical problem of how to enable the AI model to perform adaptive training with as little human intervention as possible. It achieves the technical effects of reducing the amount of manual annotation required by developers when updating the model and improving the performance stability of the model by integrating multiple models.
[0209] The following shows the technical effects of the method of this embodiment with specific data verification:
[0210] Taking the open-source data working income prediction data (‘http: / / mlr.cs.umass.edu / ml / machine-learning-databases / adult / adult.data’) as an example, a 2-layer deep neural network is used to model and predict the model. The data set has a total of 32,561 data. In the overall data set, the sample occupancy rate of incomes greater than 50K is 24%. The first 50% of the data is the initial environment, and the last 50% of the data is the environment after concept drift. For the last 50% of the data, the sample occupancy rate of incomes greater than 50K suddenly rises (concept drift) to 35%. The initial model uses the first 10,000 data for model initialization training to adaptively update the original model that has undergone concept drift.
[0211] The data is input into the model in a streaming manner. The size of the zero-time unlabeled data pool is N = 400. Every 400 data points form a block module, and the calculation of the model prediction accuracy rate is performed once. The experimental parameters are selected as shown in Table 1:
[0212]
[0213]
[0214] Table 1
[0215] It should be noted that the static model is a model that has just been established and has not been run in the actual environment.
[0216] Then, the concept drift detection, active learning, and automatic update mechanism of the model integration weights are connected. The experimental parameter settings are shown in Table 2:
[0217]
[0218] Table 2
[0219] Among them, the KS algorithm is used for drift detection, and the size of the historical data sample is 5000; drift detection is performed once for each block module, and the drift degree of each dimension feature is 1 - P_value, and the early warning threshold β warning = 0.8, and the trigger threshold β detected = 0.95, and the preset threshold Train_Num_Th for triggering adaptive update training is 64.
[0220] Figure 6 This is the experimental result graph of the online model update for online learning during sudden concept drift provided by the embodiment of the present application. As Figure 6 shown:
[0221] 1) Continuous online learning can enable the model to automatically adapt to the impact brought by concept drift, and improve the model's adaptive performance from 0.74 of the static model to about 0.80;
[0222] 2) For different learning rates, after concept drift, the speed of adapting to the environment is different, and the adaptation speed is positively correlated with the learning rate;
[0223] 3) Through ensemble learning, the variance of the inference results is reduced, and the online stability of the model is increased.
[0224] Figure 7 This is the experimental result graph of the model adaptive update for active learning of sudden concept drift provided by the embodiment of the present application. As Figure 7 shown:
[0225] 1) Through active learning, the amount of data for standard data is reduced. It is reduced from the original 20,000 pieces to about 3,520 pieces, and the performance is similar to that of full - volume data annotation. Compared with fixing the number of sampled labels for active learning, by introducing drift detection to control the number of samples for active learning, the number of labeled samples is reduced from 3,520 to 1,380, a 63% reduction in the number of labels, and the online performance is comparable.
[0226] 2) Compared with fixing the weights of the integrated model, the automatic update of the weights enables the model to adapt to the new environment faster after concept drift occurs in the online data.
[0227] Figure 7 The experimental results of the adaptive performance of each model in [it] are shown in Table III:
[0228]
[0229] Table III
[0230] Figure 8 This is the experimental result graph of online model update for online learning during incremental concept drift provided by the embodiment of the present application. As Figure 8 shown:
[0231] 1) Through online learning, adaptation to incremental drift can be achieved. Similar to a static model, at the end of the drift, the adaptive performance is improved by 3.5%.
[0232] 2) The effect of the integrated model is roughly equivalent to that of other online learning models during the data drift stage. After the drift ends, its model effect is slightly better than that of other online models.
[0233] Figure 9 This is the experimental result graph of model adaptive update for active learning of incremental concept drift provided by the embodiment of the present application. As Figure 9 shown:
[0234] 1) After the environmental change stops, the integrated model with automatically updated weights or the integrated model with the maximum number of samples can adapt to the new environment faster; while the other two methods require a certain adaptation time;
[0235] 2) Compared with the integrated model with the maximum number of samples, the integrated model with automatically updated weights reduces the number of labels by 63%, and the effect is slightly improved.
[0236] Figure 9 The experimental results of the adaptive performance of each model in [it] are shown in Table IV:
[0237]
[0238] Table IV
[0239] Figure 10 Schematic structural diagram of a model adaptive training device provided for this application. The model adaptive training device 1000 can be implemented by software, hardware, or a combination of both.
[0240] As Figure 10 shown, the model adaptive training device 1000 includes:
[0241] An acquisition module 1001, configured to acquire various operation data of the original model during actual operation;
[0242] A processing module 1002, configured to:
[0243] Detect a first concept drift value of the original model during actual operation according to the operation data;
[0244] Allocate each piece of operation data to a labeled data set and / or an unlabeled data set according to the value model and the first concept drift value;
[0245] Judge whether the data volume of the labeled data set is greater than or equal to a preset threshold;
[0246] If so, use the adaptive training model to adaptively train the original model according to the labeled data set and the unlabeled data set to determine a new trained model, and the second concept drift value of the new model is less than the first concept drift value.
[0247] In a possible design, the processing module 1002 is configured to:
[0248] Determine the comprehensive value of each piece of operation data according to the value model;
[0249] Adjust the data accumulation speed of the labeled data set according to the first concept drift value and the comprehensive value;
[0250] Allocate each piece of operation data to a labeled data set and / or an unlabeled data set according to the data accumulation speed and the comprehensive value.
[0251] In a possible design, the processing module 1002 is configured to:
[0252] Screen out data to be labeled from each piece of operation data according to the data accumulation speed and the comprehensive value;
[0253] Send the data to be labeled to the user side for labeling to determine the labeled data, and add the remaining operation data to the unlabeled data set;
[0254] The acquisition module 1001 is further configured to receive the labeled data returned by the user side;
[0255] The processing module 1002 is further configured to add the labeled data to the labeled dataset.
[0256] In a possible design, the processing module 1002 is configured to:
[0257] Determine the sorting sequence of each piece of operation data according to the comprehensive value;
[0258] When the first concept drift value is less than or equal to the warning threshold value, select the first M pieces of operation data in the sorting sequence as the data to be labeled.
[0259] In a possible design, the processing module 1002 is configured to:
[0260] When the first concept drift value is greater than or equal to the trigger threshold value, select the first N pieces of operation data in the sorting sequence as the data to be labeled, where the warning threshold value is less than the trigger threshold value.
[0261] In a possible design, the processing module 1002 is configured to:
[0262] When the first concept drift value is greater than the warning threshold value and less than the trigger threshold value, select the first K pieces of operation data in the sorting sequence as the data to be labeled, where K has a preset corresponding relationship with M and N.
[0263] In a possible design, the processing module 1002 is configured to:
[0264] Sample the labeled dataset and the unlabeled dataset through a preset sampling model to determine the training sample set;
[0265] Use a semi-supervised training model to perform semi-supervised training on the original model according to the training sample set and multiple learning rates, and then determine multiple sub-models, where the sub-models correspond to the learning rates;
[0266] Combine each sub-model into a new model through the integration weight value.
[0267] In a possible design, the processing module 1002 is configured to:
[0268] Use a dynamic update algorithm to determine the updated integration weight value according to each learning rate and the first concept drift value.
[0269] In a possible design, the processing module 1002 is configured to:
[0270] Initialize the integration weight value corresponding to each sub-model;
[0271] When the first concept drift value is less than or equal to the warning threshold value, determine the updated integration weight value according to the first update model, the preset update factor, and each learning rate.
[0272] In a possible design, the processing module 1002 is configured to:
[0273] When the first concept drift value is greater than or equal to the trigger threshold value, determine the updated integrated weight value according to the second update model, the preset update factor, and each learning rate.
[0274] In a possible design, the processing module 1002 is configured to:
[0275] When the first concept drift value is greater than the warning threshold value and less than the trigger threshold value, determine the updated integrated weight value according to the third update model, the preset update factor, and each learning rate.
[0276] In a possible design, the processing module 1002 is further configured to:
[0277] Perform a normalization operation on the integrated weight value according to the preset normalization model to determine each target integrated weight value; combine each sub-model into a new model through each target integrated weight value.
[0278] It should be noted that Figure 10 The model adaptive training device provided by the illustrated embodiment can execute the methods provided in any of the above method embodiments. The specific implementation principles, technical features, explanations of professional terms, and technical effects are similar and will not be elaborated here.
[0279] Figure 11 This is a schematic structural diagram of an electronic device provided by this application. As Figure 11 shown, the electronic device 1100 may include: at least one processor 1101 and a memory 1102. Figure 11 The illustrated electronic device takes one processor as an example.
[0280] The memory 1102 is used to store programs. Specifically, the program may include program code, and the program code includes computer operation instructions.
[0281] The memory 1102 may include a high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.
[0282] The processor 1101 is configured to execute the computer execution instructions stored in the memory 1102 to implement the methods described in the above method embodiments.
[0283] Among them, the processor 1101 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0284] Optionally, the memory 1102 can be either independent or integrated with the processor 1101. When the memory 1102 is a device independent of the processor 1101, the electronic device 1100 may further include:
[0285] A bus 1103 for connecting the processor 1101 and the memory 1102. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc., but it does not mean that there is only one bus or one type of bus.
[0286] Optionally, in specific implementation, if the memory 1102 and the processor 1101 are integrated on a chip, the memory 1102 and the processor 1101 can communicate through an internal interface.
[0287] The present application also provides a computer-readable storage medium, which may include various media capable of storing program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc. Specifically, the computer-readable storage medium stores program instructions for the methods in the above method embodiments.
[0288] The present application also provides a computer program product, including a computer program, which implements the methods in the above method embodiments when executed by a processor.
[0289] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. An electronic device, characterized in that, including a processor; and, a memory for storing a computer program of the processor; wherein the processor is configured to: obtain respective operation data of the original model during actual operation, and detect a first concept drift value of the original model during actual operation according to the operation data; allocate the respective operation data to a labeled data set and / or an unlabeled data set according to a value model and the first concept drift value; judge whether the data volume of the labeled data set is greater than or equal to a preset threshold; if so, adaptively train the original model according to the labeled data set and the unlabeled data set by using an adaptive training model to determine a trained new model, and a second concept drift value of the new model is less than the first concept drift value; the allocating the respective operation data to the labeled data set and / or the unlabeled data set according to the value model and the first concept drift value includes: screening out data to be labeled from the respective operation data according to the value model and the first concept drift value; sending the data to be labeled to a client for labeling to determine labeled data, and adding the remaining operation data to the unlabeled data set; receiving the labeled data returned by the client and adding the labeled data to the labeled data set.
2. The electronic device according to claim 1, wherein The processor is specifically configured to: determine the comprehensive value of the respective operation data according to the value model; adjust the data accumulation speed of the labeled data set according to the first concept drift value and the comprehensive value; screen out data to be labeled from the respective operation data according to the data accumulation speed and the comprehensive value.
3. The electronic device according to claim 2, characterized in that, The processor is specifically configured to: determine a sorting sequence of the respective operation data according to the comprehensive value; when the first concept drift value is less than or equal to a warning threshold value, select the first M pieces of the operation data in the sorting sequence as the data to be labeled.
4. The electronic device according to claim 3, characterized in that, When the processor is further configured to: when the first concept drift value is greater than or equal to a trigger threshold value, select the first N pieces of the operation data in the sorting sequence as the data to be labeled, and the warning threshold value is less than the trigger threshold value.
5. The electronic device according to claim 4, wherein The processor is further configured to: when the first concept drift value is greater than the warning threshold value and less than the trigger threshold value, select the first K pieces of the operation data in the sorting sequence as the data to be labeled, where K has a preset corresponding relationship with M and N.
6. The electronic device according to claim 1, wherein The processor is specifically configured to: sample the labeled data set and the unlabeled data set through a preset sampling model to determine a training sample set; perform semi-supervised training on the original model according to the training sample set and a plurality of learning rates by using a semi-supervised training model to determine a plurality of sub-models, and the sub-models correspond to the learning rates; combine the respective sub-models into the new model through an integrated weight value.
7. The electronic device according to claim 6, wherein The processor is further configured to: Using a dynamic update algorithm, determine the updated integrated weight value according to each of the learning rates and the first concept drift value.
8. The electronic device according to claim 7, wherein The processor is specifically configured to: Initialize the integrated weight value corresponding to each of the sub-models; When the first concept drift value is less than or equal to the warning threshold value, determine the updated integrated weight value according to a first update model, a preset update factor, and each of the learning rates.
9. The electronic device according to claim 8, wherein, The processor is further configured to: When the first concept drift value is greater than or equal to the trigger threshold value, determine the updated integrated weight value according to a second update model, a preset update factor, and each of the learning rates.
10. The electronic device according to claim 9, wherein, The processor is further configured to: When the first concept drift value is greater than the warning threshold value and less than the trigger threshold value, determine the updated integrated weight value according to a third update model, a preset update factor, and each of the learning rates.
11. The electronic device according to any one of claims 8-10, characterized in that, The processor is further configured to: Perform a normalization operation on the integrated weight value according to a preset normalization model to determine each target integrated weight value; Correspondingly, combining each of the sub-models into the new model by the integrated weight value includes: Combining each of the sub-models into the new model by each of the target integrated weight values.
12. A model adaptive training device, characterized in that, Includes: An acquisition module, configured to acquire each running data of the original model during actual operation; A processing module, configured to: Detect a first concept drift value of the original model during actual operation according to the running data; According to a value model and the first concept drift value, respectively allocate each of the running data to a labeled data set and / or an unlabeled data set; Judge whether the data volume of the labeled data set is greater than or equal to a preset threshold; If so, use an adaptive training model to adaptively train the original model according to the labeled data set and the unlabeled data set to determine a trained new model, and a second concept drift value of the new model is less than the first concept drift value; The processing module is specifically configured to: Screen out data to be labeled from each of the running data according to the value model and the first concept drift value; Send the data to be labeled to a user terminal for labeling to determine labeled data, and add the remaining running data to the unlabeled data set; Receive the labeled data returned by the user terminal and add the labeled data to the labeled data set.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the model adaptive training method according to any one of claims 1 to 10.
14. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the model adaptive training method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Exception detection method based on data flow concept drift
CN111143413A
Information processing program, information processing method, and information processing apparatus
JP2021051576A