Concept drift adaptive method and system based on label-free data, and terminal
By using KL distance detection concept drift in machine learning models and optimizing the random forest model with Gini coefficient, weight and density difference, the problem of manually labeling data updating the model in the existing technology is solved, and real-time concept drift adaptation with labelless data is realized.
Patent Information
- Application Number
- CN202510106895.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art requires the use of manual labeled data to update the model when dealing with concept drift, which is time-consuming and labor-intensive, and it is impossible to directly use real-time data without labels for timely update.
By collecting old data as training sets and new data without labels as test sets, KL distance is calculated to detect concept drift, and when drift is detected, the test set data is added to the training set. Based on the original random forest model, a new criterion is established to optimize the model to adapt to data changes.
The concept drift adaptation based on labelless data is realized, and the new labelless data is directly updated, saving time and effort, and can adapt to data distribution changes in time.
Smart Images

Figure CN120030349A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning and data mining, and in particular to a concept drift self-adaptation method based on unlabeled data, a system, a terminal and a computer-readable storage medium. Background Art
[0002] Concept drift is an important concept in the field of machine learning and data mining. It refers to the phenomenon that the statistical characteristics of data (such as distribution, average, etc.) change over time, causing the originally established data model to no longer be applicable. This phenomenon is particularly common in real-time data stream environments, such as financial market analysis, network security monitoring, and weather forecasting. Concept drift will cause the mapping relationship between some old data and new data to be inconsistent, making these old data noise relative to the new data, so the model trained with these old data cannot be applied to the new data. In order to solve the problems caused by concept drift, concept drift adaptation is a good way to deal with it after detecting the occurrence of concept drift. Concept drift adaptation can adapt to changes in data distribution by autonomously and continuously iteratively updating the model.
[0003] However, in existing concept drift adaptive methods, model updates require a large amount of new unlabeled data, but these new unlabeled data need to be manually labeled before they can be used to update the model. This process is time-consuming and labor-intensive. In real-time data streams, existing technologies cannot directly use unlabeled real-time data to update models in a timely manner, and still need to use labeled data to update the model.
[0004] Therefore, the prior art still needs to be improved and developed. Summary of the invention
[0005] The main purpose of the present invention is to provide a concept drift adaptive method, system, terminal and computer-readable storage medium based on unlabeled data, aiming to solve the problem that the existing concept drift adaptive method needs to use manually labeled data to update the model, which is time-consuming and labor-intensive.
[0006] To achieve the above-mentioned object of the invention, the present invention provides a concept drift adaptive method based on unlabeled data, and the concept drift adaptive method based on unlabeled data includes:
[0007] Collect old data as training set and collect new unlabeled data as test set;
[0008] Calculate the KL distance based on the training set and the test set, compare the KL distance with a set threshold, and determine whether concept drift occurs in the new data;
[0009] When it is determined that the new data has a concept drift, based on the training set, some data is selected from the test set and added to the training set to obtain a new training set;
[0010] Based on the original random forest model, a new tree construction criterion is established by combining the Gini coefficient, weight and density difference to obtain the RFKL model.
[0011] Based on the new training set, a feature corresponding to the minimum value of the new criterion for establishing the new tree is selected as a split feature, and the RFKL model is trained by using the split feature to obtain a trained RFKL model;
[0012] The performance of the trained RFKL model is evaluated using the accuracy as an evaluation index, and the trained RFKL model whose performance meets the requirements is used as the RFKL model after concept drift adaptation.
[0013] Optionally, the collecting of old data as a training set and collecting of unlabeled new data as a test set may also include:
[0014] When the number of samples of the new data in the real-time data stream reaches the set window length, the concept drift detection process is started.
[0015] Optionally, the calculating the KL distance according to the training set and the test set, comparing the KL distance with a set threshold, and determining whether concept drift occurs in the new data specifically includes:
[0016] The KL distance is calculated based on the training set and the test set, wherein the KL distance is used to measure the difference between the distribution of the training set on the original random forest model and the distribution of the test set on the original random forest model:
[0017]
[0018] Among them, D(P i ||Q i ) represents the training set P i With the test set Q i The KL distance between v,i represents the number of samples in the training set that fall on the vth leaf node of the i-th tree, Q v,i represents the number of samples in the test set that fall on the vth leaf node of the i-th tree, W 1 is the total number of samples in the training set, W 2 is the total number of samples in the test set, and L represents the number of leaf nodes in the i-th tree;
[0019] Compare the KL distance with a set threshold;
[0020] When the KL distance is greater than the set threshold, it is determined that concept drift occurs in the new data;
[0021] When the KL distance is not greater than the set threshold, it is determined that no concept drift occurs in the new data.
[0022] Optionally, the new training set includes labeled training set data and unlabeled test set data.
[0023] Optionally, the new tree construction criteria are established based on the original random forest model in combination with the Gini coefficient, weight and density difference to obtain the RFKL model, which specifically includes:
[0024] The original random forest model is used as the basic model, wherein the decision tree type of the original random forest model is a classification tree;
[0025] Based on the original random forest model using the Gini coefficient Gini as the tree building criterion, the weight w and density difference D are added to obtain the new tree building criterion C:
[0026] C = Gini + w * D;
[0027] The density difference D refers to the density difference between the labeled training set data and the unlabeled test set data in the left and right branches when constructing the decision tree:
[0028] D=|l left / l full -u left / u full |;
[0029] Among them, l left is the number of samples in the labeled training set data that are less than the split threshold and distributed in the left subtree; l full is the total number of samples of labeled training set data at the current node; u left is the number of samples in the unlabeled test set data that are less than the split threshold and distributed in the left subtree; u full is the total number of samples of unlabeled test set data at the current node;
[0030] The original random forest model and the new tree building criterion are combined to obtain the RFKL model.
[0031] Optionally, based on the new training set, selecting a feature corresponding to the minimum value of the new criterion for establishing the new criterion as a split feature, and using the split feature to train the RFKL model to obtain a trained RFKL model specifically includes:
[0032] Extracting features of the new training set to obtain multiple feature columns;
[0033] Establishing an optimal gain value, and setting an initial value of the optimal gain value as a set value;
[0034] Traversing all feature columns, and calculating the midpoint value between the current feature and the adjacent features in each feature column while traversing the feature column, and forming a threshold column of the current feature with multiple midpoint values;
[0035] Traverse the threshold column of the current feature and calculate the value of the current new criterion C;
[0036] Compare the value of the current new criterion C with the current optimal gain value, and if the value of the current new criterion C is less than the current optimal gain value, update the current optimal gain value to the value of the current criterion C;
[0037] Repeat the calculation to obtain the value of the new criterion C corresponding to each feature, and use the feature corresponding to the smallest value of the new criterion C as the split feature;
[0038] The RFKL model is trained using the split features to obtain a trained RFKL model.
[0039] Optionally, the using accuracy as an evaluation index to evaluate the performance of the trained RFKL model, and using the trained RFKL model that meets the performance requirements as the RFKL model after concept drift adaptation, specifically includes:
[0040] Calculate the accuracy rate according to the prediction results of the trained RFKL model, wherein the accuracy rate is the ratio of the number of correct prediction results to the total number of prediction results;
[0041] Using the accuracy rate as an evaluation index to evaluate the performance of the trained RFKL model, if the accuracy rate is greater than a performance threshold, the performance of the trained RFKL model is determined to meet the requirements;
[0042] The trained RFKL model whose performance meets the requirements is used as the RFKL model after concept drift adaptation.
[0043] To achieve the above-mentioned object of the invention, the present invention further provides a concept drift adaptive system based on unlabeled data, the concept drift adaptive system based on unlabeled data comprising:
[0044] Data collection module: used to collect old data as training sets and new unlabeled data as test sets;
[0045] Concept drift detection module: used to calculate the KL distance based on the training set and the test set, compare the KL distance with a set threshold, and determine whether the new data has concept drift;
[0046] New training set acquisition module: used for selecting part of the data in the test set and adding it to the training set to obtain a new training set when it is determined that the new data has concept drift, based on the training set;
[0047] RFKL model acquisition module: used to establish a new tree construction criterion based on the original random forest model, combining the Gini coefficient, weight and density difference to obtain the RFKL model;
[0048] RFKL model training module: used for selecting the feature corresponding to the minimum value of the new criterion for establishing the new criterion as the split feature based on the new training set, and training the RFKL model by using the split feature to obtain a trained RFKL model;
[0049] RFKL model evaluation module: used to evaluate the performance of the trained RFKL model using accuracy as an evaluation indicator, and use the trained RFKL model that meets the performance requirements as the RFKL model after concept drift adaptation.
[0050] To achieve the above-mentioned purpose of the invention, the present invention also provides a terminal, which includes: a memory, a processor, and a concept drift adaptation program based on unlabeled data stored in the memory and executable on the processor, wherein the concept drift adaptation program based on unlabeled data implements the steps of the concept drift adaptation method based on unlabeled data as described above when executed by the processor.
[0051] To achieve the above-mentioned purpose of the invention, the present invention also provides a computer-readable storage medium, which stores a concept drift adaptation program based on unlabeled data. When the concept drift adaptation program based on unlabeled data is executed by a processor, the steps of the concept drift adaptation method based on unlabeled data as described above are implemented.
[0052] In the present invention, old data is collected as a training set, and unlabeled new data is collected as a test set; the KL distance is calculated according to the training set and the test set, and the KL distance is compared with a set threshold to determine whether the new data has undergone concept drift; when it is determined that the new data has undergone concept drift, on the basis of the training set, part of the data is selected from the test set and added to the training set to obtain a new training set; based on the original random forest model, a new tree-building criterion is established in combination with the Gini coefficient, weight and density difference to obtain an RFKL model; based on the new training set, a feature corresponding to the minimum value of the new tree-building criterion is selected as a split feature, and the RFKL model is trained using the split feature to obtain a trained RFKL model; the accuracy is used as an evaluation index to evaluate the performance of the trained RFKL model, and the trained RFKL model whose performance meets the requirements is used as the RFKL model after concept drift adaptation. The present invention uses KL distance to measure the difference between the new data distribution and the original data distribution, and then detects whether the data has undergone concept drift; unlabeled test set data is used to participate in model training to achieve self-adaptation based on concept drift of unlabeled data; a new tree building criterion C is used to build a tree, and feature values with a smaller Gini coefficient and little change in feature distribution of new and old data are preferentially selected as split nodes for building a classification tree. The present invention directly uses unlabeled new data to update the model, achieves self-adaptation based on concept drift of unlabeled data, and saves time and effort. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 It is a flow chart of a preferred embodiment of the concept drift adaptive method based on unlabeled data of the present invention;
[0054] Figure 2 is another flow chart of a preferred embodiment of the concept drift adaptive method based on unlabeled data of the present invention;
[0055] Figure 3 It is a structural diagram of a preferred embodiment of the concept drift adaptive system based on unlabeled data of the present invention;
[0056] Figure 4 It is a structural diagram of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION
[0057] In order to make the purpose, technical solution and advantages of the present invention clearer and more specific, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0058] Concept drift is an important concept in the field of machine learning and data mining. It refers to the phenomenon that the statistical characteristics of data (such as distribution, average, etc.) change over time, causing the originally established data model to no longer be applicable. This phenomenon is particularly common in real-time data stream environments, such as financial market analysis, network security monitoring, and weather forecasting. Concept drift will cause the mapping relationship between some old data and new data to be inconsistent, making these old data noise relative to the new data, so the model trained with these old data cannot be applied to the new data. Concept drift can be caused by many reasons, including but not limited to changes in the data generation process, environmental changes, changes in data collection methods, etc. The types of concept drift usually include: sudden drift, gradual drift, incremental drift, and recurrent drift.
[0059] The emergence of concept drift will cause the previously established model to no longer be applicable or become less effective. In order to solve the problems caused by concept drift, concept drift adaptation is a good way to deal with it after the concept drift is detected. Concept drift adaptation can adapt to changes in data distribution by autonomously and continuously iteratively updating the model.
[0060] The existing concept drift adaptive methods can be roughly divided into two perspectives: single model and integration. From the integration perspective, Ryan et al. designed a streaming prediction model based on online ensemble learning, Learn++NSE (Lear++Non-Stationary Environments, a multi-classifier integration algorithm for non-stable environments), which creates classifiers to learn new knowledge, combines dynamically updated voting weights, and selects current and historical classifiers for prediction; the Adaptive Random Forest (ARF) algorithm designed by Heitor designs a drift monitor for each Hofting tree in the forest, trains a new tree when an early warning is detected, and replaces the tree when a warning is detected. Single models can generally be divided into two aspects: retraining and incremental learning. Adaptive methods of incremental learning, such as the SAM-KNN algorithm (Self-Adjusting Memory K-NearestNeighbors), combine the current concept and all old concept-specific models to handle concept drifts of different types and rates, while minimizing errors and building short-term memory of current concepts and long-term memory of old concepts; adaptive methods of retraining, such as sample filtering methods based on drift regions, delete pre-drift data, use new data to retrain the base classifier to effectively identify drift regions, and use drift region information to filter data samples of the training model. In the above existing technologies, the update of the model requires a large amount of unlabeled new data, but these unlabeled new data need to be manually labeled before they can be used to update the model. This process is time-consuming and labor-intensive. In real-time data streams, existing technologies cannot directly use unlabeled real-time data to update the model in a timely manner, and still need to use labeled data to update the model.
[0061] In order to solve the above technical problems, the present invention provides a concept drift adaptive method based on unlabeled data, which collects old data as a training set and collects unlabeled new data as a test set; calculates the KL distance according to the training set and the test set, compares the KL distance with a set threshold, and determines whether the new data has concept drift; when it is determined that the new data has concept drift, on the basis of the training set, selects part of the data in the test set to add to the training set to obtain a new training set; based on the original random forest model, a new tree building criterion is established in combination with the Gini coefficient, weight and density difference to obtain a RFKL (Random Forest with KL divergence minimization, random forest based on minimizing KL divergence) model; based on the new training set, selects the feature corresponding to the minimum value of the new tree building criterion as a split feature, and uses the split feature to train the RFKL model to obtain a trained RFKL model; uses the accuracy as an evaluation index to evaluate the performance of the trained RFKL model, and uses the trained RFKL model that meets the performance requirements as the RFKL model after concept drift adaptation. The present invention uses KL distance to measure the difference between the new data distribution and the original data (i.e., old data) distribution, and then detects whether the data has undergone concept drift; unlabeled test set data is used to participate in model training to achieve self-adaptation based on concept drift of unlabeled data; a new tree building criterion C is used to build a tree, and feature values with a smaller Gini coefficient and little change in feature distribution of new and old data are preferentially selected as split nodes for building a classification tree. The present invention directly uses unlabeled new data to update the model, achieves self-adaptation based on concept drift of unlabeled data, and saves time and effort.
[0062] The application content is further explained below through the description of embodiments in conjunction with the accompanying drawings.
[0063] A preferred embodiment of the concept drift adaptive method based on unlabeled data of the present invention is as follows: Figure 1 and Figure 2 As shown, specifically including:
[0064] S1. Collect old data as training set and collect new unlabeled data as test set.
[0065] In an implementation of this embodiment, the collecting of old data as a training set and the collecting of unlabeled new data as a test set also include:
[0066] When the number of samples of the new data in the real-time data stream reaches the set window length, the concept drift detection process is started.
[0067] S2. Calculate the KL distance based on the training set and the test set, compare the KL distance with a set threshold, and determine whether concept drift occurs in the new data.
[0068] In an implementation of this embodiment, calculating the KL distance according to the training set and the test set, comparing the KL distance with a set threshold, and determining whether concept drift occurs in the new data specifically includes:
[0069] The KL distance is calculated based on the training set and the test set, wherein the KL distance is used to measure the difference between the distribution of the training set on the original random forest model and the distribution of the test set on the original random forest model:
[0070]
[0071] Among them, D(P i ||Q i ) represents the training set P i With the test set Q i The KL distance between v,i represents the number of samples in the training set that fall on the vth leaf node of the i-th tree, Q v,i represents the number of samples in the test set that fall on the vth leaf node of the i-th tree, W 1 is the total number of samples in the training set, W 2 is the total number of samples in the test set, and L represents the number of leaf nodes in the i-th tree;
[0072] Compare the KL distance with a set threshold;
[0073] When the KL distance is greater than the set threshold, it is determined that concept drift occurs in the new data;
[0074] When the KL distance is not greater than the set threshold, it is determined that no concept drift occurs in the new data.
[0075] Specifically, when analyzing whether concept drift has occurred in new data, the present invention uses KL distance as an indicator to detect whether concept drift has occurred. KL distance measures the difference between the distribution of the training set on the original random forest model and the distribution of the test set on the original random forest model. The distribution of the training set represents the distribution of the old data, and the distribution of the test set represents the distribution of the new data. The size of the KL distance reflects the distribution difference between the new and old data. Once the value of the KL distance exceeds the set threshold, it can be determined that concept drift has occurred in the new data.
[0076] S3. When it is determined that the new data has undergone concept drift, based on the training set, some data is selected from the test set and added to the training set to obtain a new training set.
[0077] In one implementation of this embodiment, the new training set includes labeled training set data and unlabeled test set data.
[0078] Specifically, when new data is detected to have concept drift, the concept drift adaptive process is started. First, based on the original training set, a part of unlabeled data extracted from the real-time data stream (i.e., part of the test set data) is added to participate in the model training. The existing training sets are all labeled data, representing normal data before the concept drift occurs, while the newly added part of the test set data is unlabeled data. The present invention combines these two parts of data to participate in the tree building process.
[0079] S4. Based on the original random forest model, a new tree construction criterion is established by combining the Gini coefficient, weight and density difference to obtain the RFKL model.
[0080] In one implementation of this embodiment, the original random forest model is based on a new tree construction criterion combined with the Gini coefficient, weight and density difference to obtain the RFKL model, which specifically includes:
[0081] The original random forest model is used as the basic model, wherein the decision tree type of the original random forest model is a classification tree;
[0082] Based on the original random forest model using the Gini coefficient Gini as the tree building criterion, the weight w and density difference D are added to obtain the new tree building criterion C:
[0083] C = Gini + w * D;
[0084] The density difference D refers to the density difference between the labeled training set data and the unlabeled test set data in the left and right branches when constructing the decision tree:
[0085] D=|l left / l full -u left / u full ;
[0086] Among them, l left is the number of samples in the labeled training set data that are less than the split threshold and distributed in the left subtree; l full is the total number of samples of labeled training set data at the current node; u left is the number of samples in the unlabeled test set data that are less than the split threshold and distributed in the left subtree; u full is the total number of samples of unlabeled test set data at the current node;
[0087] The original random forest model and the new tree building criterion are combined to obtain the RFKL model.
[0088] Specifically, the present invention uses the original random forest model (i.e., random forest ensemble model) as the basic model, in which the decision tree type is a classification tree. Based on the original use of the Gini coefficient as the tree building criterion, the present invention adds two new parameters, weight w and density difference D, wherein the weight w is the weight occupied by the density difference D, which is an adjustable parameter; the density difference D refers to the density difference between the training set data and the test set data in the left and right branches when constructing a decision tree classifier, and the density difference D can reflect the distribution difference between the labeled training set data and the unlabeled test set data at the same node. When selecting the split feature, the smaller the density difference D, the more similar the split direction of the labeled training set data and the unlabeled test set data at the same node, indicating that the model has a good adaptability to new data with concept drift, and can make its splitting effect on the current split feature node roughly consistent with the old data.
[0089] S5. Based on the new training set, select the feature corresponding to the minimum value of the new criterion for establishing the new criterion as the split feature, and use the split feature to train the RFKL model to obtain a trained RFKL model.
[0090] In an implementation of this embodiment, based on the new training set, selecting a feature corresponding to the minimum value of the new criterion for establishing the new criterion as a split feature, and using the split feature to train the RFKL model to obtain a trained RFKL model specifically includes:
[0091] Extracting features of the new training set to obtain multiple feature columns;
[0092] Establishing an optimal gain value, and setting an initial value of the optimal gain value as a set value;
[0093] Traversing all feature columns, and calculating the midpoint value between the current feature and the adjacent features in the feature column while traversing each feature column, and forming a threshold column of the current feature with multiple midpoint values;
[0094] Traverse the threshold column of the current feature and calculate the value of the current new criterion C;
[0095] Compare the value of the current new criterion C with the current optimal gain value, and if the value of the current new criterion C is less than the current optimal gain value, update the current optimal gain value to the value of the current criterion C;
[0096] Repeat the calculation to obtain the value of the new criterion C corresponding to each feature, and use the feature corresponding to the smallest value of the new criterion C as the split feature;
[0097] The RFKL model is trained using the split features to obtain a trained RFKL model.
[0098] Specifically, the minimum value of the new criterion, Min(C), is considered comprehensively to select the appropriate split feature. The detailed steps for selecting the appropriate split feature are as follows:
[0099] (1) Establish an optimal gain value best_gain, with an initial value of 1;
[0100] (2) Traverse all feature columns and calculate the midpoint value (i.e., threshold) of adjacent elements in each feature column while traversing the feature column to obtain the threshold column of the current feature;
[0101] (3) Traverse the threshold column of the current feature, calculate the value of the current new criterion C, and then compare it with the current optimal gain value. If the value of the current new criterion C is smaller than the current optimal gain value, update the current optimal gain value to the value of the current new criterion C;
[0102] (4) Repeat the above steps until the most suitable splitting feature for the entire new training set is found.
[0103] After completing the data preprocessing, we start training the random forest model, which is called the RFKL model. The training process of the RFKL model is the same as that of the traditional random forest model, except that the criterion for constructing each tree is changed from minimizing the Gini coefficient to minimizing the C value.
[0104] S6. Use accuracy as an evaluation index to evaluate the performance of the trained RFKL model, and use the trained RFKL model that meets the performance requirements as the RFKL model after concept drift adaptation.
[0105] In an implementation of this embodiment, the use of accuracy as an evaluation index to evaluate the performance of the trained RFKL model, and using the trained RFKL model that meets the performance requirements as the RFKL model after concept drift adaptation, specifically includes:
[0106] Calculate the accuracy rate according to the prediction results of the trained RFKL model, wherein the accuracy rate is the ratio of the number of correct prediction results to the total number of prediction results;
[0107] Using the accuracy rate as an evaluation index to evaluate the performance of the trained RFKL model, if the accuracy rate is greater than a performance threshold, the performance of the trained RFKL model is determined to meet the requirements;
[0108] The trained RFKL model whose performance meets the requirements is used as the RFKL model after concept drift adaptation.
[0109] Specifically, in the prediction results of the RFKL model, accuracy is used as the evaluation indicator, and accuracy is the proportion of correct prediction results to all results.
[0110] When dealing with the problem of concept drift, since the model trained with old data is no longer sufficient to cope with the changes in new data, the simplest method is generally to retrain a model with new data that has experienced concept drift. However, this sometimes costs a lot of time and manpower (because labeling new data takes time and manpower). Therefore, the present invention considers optimizing the model to achieve the effect of being able to adapt to changes in new data and continue to be used. The two most important issues in model optimization are detecting whether concept drift has occurred and how to promptly perform concept drift adaptation after concept drift has occurred.
[0111] The present invention designs corresponding solutions based on the random forest model for these two problems: when detecting whether concept drift occurs, the KL distance is used to measure the difference between the distribution of the newly added data and the distribution of the original data, so that whether concept drift occurs can be determined without data labels; after detecting that concept drift has occurred, the goal of the present invention is to use unlabeled test data (i.e., test set data) plus the original labeled training data (training set data) to construct a new decision tree classifier to adapt to data changes and improve the accuracy of the decision tree on data with concept drift. The specific method is that when constructing the decision tree classifier, the standard for selecting split features and values is not only the Gini index, but also the distribution difference between the training data and the test data in the left and right branches is considered.
[0112] In another implementation of this embodiment, Figure 2 As shown in the figure, KL distance is first used to detect whether concept drift has occurred. If the value of KL distance exceeds the set threshold, it means that the distribution of new and old data in the original random forest model is relatively different, that is, concept drift has occurred; then the existing labeled training set and the unlabeled test set are merged into a new training set, and Bootstrap sampling is performed to obtain the training set of each decision tree in the new random forest model; then the new tree criterion Min(C) is used to build the tree, where C=Gini+w*D; finally, the prediction results of all decision trees are integrated, and the accuracy is used to evaluate the results.
[0113] It should be noted that the unlabeled concept drift adaptive method proposed in the present invention is applicable to the case where some data features drift or the overall drift is small. If the data features drift greatly as a whole, the present invention will not improve the model much, and it is recommended to collect new labeled data to update the model. However, the present invention can be used as a transition model in the data preparation stage of collecting data and labeling data labels to reduce the degree of reduction in classification accuracy.
[0114] In addition, based on the above-mentioned concept drift adaptive method based on unlabeled data, the present invention also provides a concept drift adaptive system based on unlabeled data, wherein a preferred embodiment of the concept drift adaptive system based on unlabeled data is as follows: Figure 3 As shown, specifically including:
[0115] Data collection module 01: used to collect old data as training sets and new unlabeled data as test sets;
[0116] Concept drift detection module 02: used to calculate the KL distance according to the training set and the test set, compare the KL distance with a set threshold, and determine whether the new data has concept drift;
[0117] New training set acquisition module 03: for selecting part of the data in the test set and adding it to the training set to obtain a new training set when it is determined that the new data has concept drift;
[0118] RFKL model acquisition module 04: used to establish a new tree construction criterion based on the original random forest model, combining the Gini coefficient, weight and density difference, and obtain the RFKL model;
[0119] RFKL model training module 05: used for selecting, based on the new training set, a feature corresponding to the minimum value of the new criterion for establishing the new criterion as a split feature, and using the split feature to train the RFKL model to obtain a trained RFKL model;
[0120] RFKL model evaluation module 06: used to evaluate the performance of the trained RFKL model using accuracy as an evaluation indicator, and use the trained RFKL model that meets the performance requirements as the RFKL model after concept drift adaptation.
[0121] In addition, based on the above-mentioned concept drift adaptive method and system based on unlabeled data, the present invention also provides a terminal accordingly, wherein a preferred embodiment of the terminal is as follows: Figure 4 As shown, it specifically includes a processor 10, a memory 20 and a display 30. Figure 4 Only some components of the terminal are shown, but it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0122] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, and a flash card (Flash Card) equipped on the terminal. Further, the memory 20 may also include both an internal storage unit of the terminal and an external storage device. The memory 20 is used to store application software and various types of data installed on the terminal, such as program codes of the storage terminal. The memory 20 may also be used to temporarily store data that has been output or is to be output. In one embodiment, a concept drift adaptive program 40 based on unlabeled data is stored on the memory 20, and the concept drift adaptive program 40 based on unlabeled data can be executed by the processor 10, thereby implementing the steps of the concept drift adaptive method based on unlabeled data in the present application.
[0123] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor or other data processing chip, used to run program codes stored in the memory 20 or process data, such as executing a concept drift adaptation program 40 based on unlabeled data.
[0124] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an OLED (Organic Light-Emitting Diode) touch device, etc. The display 30 is used to display information on the terminal and to display a visual user interface.
[0125] In one embodiment, when the processor 10 executes the concept drift adaptation program 40 based on unlabeled data in the memory 20 , the steps of the concept drift adaptation method based on unlabeled data as described above are implemented.
[0126] The present invention also provides a computer-readable storage medium accordingly, wherein the computer-readable storage medium stores a concept drift adaptation program based on unlabeled data, and when the concept drift adaptation program based on unlabeled data is executed by a processor, the steps of the concept drift adaptation method based on unlabeled data as described above are implemented.
[0127] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal. In the absence of further restrictions, an element defined by the sentence "includes a ..." does not exclude the existence of other identical elements in the process, method, article or terminal including the element.
[0128] Of course, those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing related hardware (such as a processor, a controller, etc.) through a computer program, and the program can be stored in a computer-readable storage medium that can be read by a computer, and the program can include the processes of the above-mentioned method embodiments when executed. The computer-readable storage medium can be a memory, a disk, an optical disk, etc.
[0129] It should be understood that the application of the present invention is not limited to the above examples. For ordinary technicians in this field, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.
Claims
1. A concept drift adaptive method based on unlabeled data, characterized in that: The concept drift adaptive method based on unlabeled data includes: Collect old data as training set and collect new unlabeled data as test set; Calculate the KL distance based on the training set and the test set, compare the KL distance with a set threshold, and determine whether concept drift occurs in the new data; When it is determined that the new data has a concept drift, based on the training set, some data is selected from the test set and added to the training set to obtain a new training set; Based on the original random forest model, a new tree construction criterion is established by combining the Gini coefficient, weight and density difference to obtain the RFKL model. Based on the new training set, a feature corresponding to the minimum value of the new criterion for establishing the new tree is selected as a split feature, and the RFKL model is trained by using the split feature to obtain a trained RFKL model; The performance of the trained RFKL model is evaluated using the accuracy as an evaluation index, and the trained RFKL model whose performance meets the requirements is used as the RFKL model after concept drift adaptation.
2. The concept drift adaptive method based on unlabeled data according to claim 1, characterized in that: The old data is collected as a training set, and the new data without labels is collected as a test set. Previously, the following was also included: When the number of samples of the new data in the real-time data stream reaches the set window length, the concept drift detection process is started.
3. The concept drift adaptive method based on unlabeled data according to claim 1, characterized in that: The calculating of the KL distance according to the training set and the test set, comparing the KL distance with a set threshold, and determining whether concept drift occurs in the new data specifically includes: The KL distance is calculated based on the training set and the test set, wherein the KL distance is used to measure the difference between the distribution of the training set on the original random forest model and the distribution of the test set on the original random forest model: Among them, D(P i ||Q i ) represents the training set P i With the test set Q i The KL distance between v,i represents the number of samples in the training set that fall on the vth leaf node of the i-th tree, Q v,i represents the number of samples in the test set that fall on the vth leaf node of the i-th tree, W1 is the total number of samples in the training set, W2 is the total number of samples in the test set, and L represents the number of leaf nodes in the i-th tree; Compare the KL distance with a set threshold; When the KL distance is greater than the set threshold, it is determined that concept drift occurs in the new data; When the KL distance is not greater than the set threshold, it is determined that no concept drift occurs in the new data.
4. The concept drift adaptive method based on unlabeled data according to claim 1, characterized in that: The new training set includes labeled training set data and unlabeled test set data.
5. The concept drift adaptive method based on unlabeled data according to claim 4 is characterized in that: The original random forest model is based on the Gini coefficient, weight and density difference to establish a new tree construction criterion, and the RFKL model is obtained, which specifically includes: The original random forest model is used as the basic model, wherein the decision tree type of the original random forest model is a classification tree; Based on the original random forest model using the Gini coefficient Gini as the tree building criterion, the weight w and density difference D are added to obtain the new tree building criterion C: C = Gini + w * D; The density difference D refers to the density difference between the labeled training set data and the unlabeled test set data in the left and right branches when constructing the decision tree: D=|l left / l full -u left / u full |; Among them, l left is the number of samples in the labeled training set data that are less than the split threshold and distributed in the left subtree; l full is the total number of samples of labeled training set data at the current node; u left is the number of samples in the unlabeled test set data that are less than the split threshold and distributed in the left subtree; u full is the total number of samples of unlabeled test set data at the current node; The original random forest model and the new tree building criterion are combined to obtain the RFKL model.
6. The concept drift adaptive method based on unlabeled data according to claim 4, characterized in that: The step of selecting, based on the new training set, a feature corresponding to the minimum value of the new criterion for establishing the new tree as a split feature, and using the split feature to train the RFKL model to obtain a trained RFKL model specifically includes: Extracting features of the new training set to obtain multiple feature columns; Establishing an optimal gain value, and setting an initial value of the optimal gain value as a set value; Traversing all feature columns, and calculating the midpoint value between the current feature and the adjacent features in the feature column while traversing each feature column, and forming a threshold column of the current feature with multiple midpoint values; Traverse the threshold column of the current feature and calculate the value of the current new criterion C; Compare the value of the current new criterion C with the current optimal gain value, and if the value of the current new criterion C is less than the current optimal gain value, update the current optimal gain value to the value of the current criterion C; Repeat the calculation to obtain the value of the new criterion C corresponding to each feature, and use the feature corresponding to the smallest value of the new criterion C as the split feature; The RFKL model is trained using the split features to obtain a trained RFKL model.
7. The concept drift adaptive method based on unlabeled data according to claim 1, characterized in that: The method of using the accuracy as an evaluation index to evaluate the performance of the trained RFKL model and using the trained RFKL model that meets the performance requirements as the RFKL model after concept drift adaptation specifically includes: Calculate the accuracy rate according to the prediction results of the trained RFKL model, wherein the accuracy rate is the ratio of the number of correct prediction results to the total number of prediction results; Using the accuracy rate as an evaluation index to evaluate the performance of the trained RFKL model, if the accuracy rate is greater than a performance threshold, the performance of the trained RFKL model is determined to meet the requirements; The trained RFKL model whose performance meets the requirements is used as the RFKL model after concept drift adaptation.
8. A concept drift adaptive system based on unlabeled data, characterized in that: The concept drift adaptive system based on unlabeled data includes: Data collection module: used to collect old data as training sets and new unlabeled data as test sets; Concept drift detection module: used to calculate the KL distance based on the training set and the test set, compare the KL distance with a set threshold, and determine whether the new data has concept drift; New training set acquisition module: used for selecting part of the data in the test set and adding it to the training set to obtain a new training set when it is determined that the new data has concept drift, based on the training set; RFKL model acquisition module: used to establish a new tree construction criterion based on the original random forest model, combining the Gini coefficient, weight and density difference to obtain the RFKL model; RFKL model training module: used for selecting the feature corresponding to the minimum value of the new criterion for establishing the new criterion as the split feature based on the new training set, and training the RFKL model by using the split feature to obtain a trained RFKL model; RFKL model evaluation module: used to evaluate the performance of the trained RFKL model using accuracy as an evaluation indicator, and use the trained RFKL model that meets the performance requirements as the RFKL model after concept drift adaptation.
9. A terminal, characterized in that: The terminal includes: a memory, a processor, and a concept drift adaptation program based on unlabeled data stored in the memory and executable on the processor. When the concept drift adaptation program based on unlabeled data is executed by the processor, the steps of the concept drift adaptation method based on unlabeled data as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a concept drift adaptation program based on unlabeled data, and when the concept drift adaptation program based on unlabeled data is executed by a processor, the steps of the concept drift adaptation method based on unlabeled data according to any one of claims 1 to 7 are implemented.