Data stream dynamic feature selection method, electronic equipment and medium

Through the method of combining Gaussian Copula model and decision grove, missing feature values are dynamically filled and subsets of feature are divided, and an integrated learning framework is constructed, which solves the feature drift and missing problems in data flow analysis, and achieves efficient and accurate feature selection and prediction.

CN120387009AActive Publication Date: 2025-07-29NAT UNIV OF DEFENSE TECH
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510873006.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-29
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

In data flow analysis in open environments, feature drift and missing feature values lead to dynamic changes in the subset of features, and the existing technology has problems of loss of knowledge and high computing costs.

Method used

The Gaussian Copula model is used to fill in the missing feature values, and the target feature subset is filtered through multiple decision groves in the sliding window, and the instances are divided into core, auxiliary and complete feature subsets. An integrated learning framework is built, and the classifier weights are dynamically adjusted to achieve weighted prediction.

Benefits of technology

It effectively solves the problem of feature drift and missing, reduces calculation costs, avoids knowledge loss, realizes continuous and accurate predictions, and reduces dependence on real labels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387009A_ABST
    Figure CN120387009A_ABST
Patent Text Reader

Abstract

The invention relates to a data stream dynamic feature selection method, electronic equipment and a medium. The method comprises the following steps: filling missing feature values in an instance by utilizing a Gaussian Copula model; a target feature subset is screened based on the filled data, and an instance is divided into a core feature subset, an auxiliary feature subset and a complete feature subset. Distributing weights for the subsets according to the feature values, and generating corresponding training samples; initializing classifiers, training by using the training samples, and distributing initial prediction weights for the classifiers; and adding the trained classifier into an integrated learning framework, and adjusting an integrated prediction weight according to the type of the classifier. And when a new instance is received, complementing missing features through a Gaussian Copula model, performing weighted prediction by using a classifier, and outputting a final classification result. According to the method, the limitations of knowledge loss, high calculation cost and dependence on real tags in traditional feature selection are overcome, and continuous and accurate prediction is realized by dynamically screening related features and retaining multi-dimensional knowledge.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of dynamic feature selection learning of streaming data, and specifically relates to a method for dynamically selecting data flow features, an electronic device, and a medium. Background Art

[0002] In the problem of data flow analysis in an open environment, the data captured by data flow sensors faces challenges such as incomplete feature values and changing feature correlations. The existence of feature drift and missing feature values complicates the problem because it may cause changes in the feature subset related to the decision-making problem. This change poses a severe challenge to the continuous use and accurate prediction of learning models.

[0003] Existing research work mainly focuses on two types of technologies: feature selection and variable feature space learning. Feature selection technology trains a classification model by screening features most relevant to the category to reduce the impact of redundant features on classification performance. However, this technology may lead to discontinuous feature spaces of training samples and may cause potential knowledge loss. For this reason, researchers have also proposed variable feature space learning technology. Variable feature space learning technology selects features to cope with feature drift and alleviates knowledge loss by mining the relationships between features. However, due to the need to establish a feature mapping model, when the number of features in the data stream is large, this method incurs high computational costs. To process changing features, the above methods also need to obtain the true labels of data stream samples to detect drift, which brings expensive label costs in practical applications. Summary of the Invention

[0004] The present invention provides a method for dynamically selecting data flow features, an electronic device, and a medium, aiming to solve the problem of dynamic changes in feature subsets caused by feature drift and missing feature values in data flow analysis in an open environment.

[0005] To achieve the above object, the first aspect of the present invention provides a method for dynamically selecting data flow features, including the following steps: Receiving the currently arriving instance in the data stream, and using the Gaussian Copula model to fill in the missing feature values in the currently arriving instance to obtain processed instance data; Based on the processed instance data, screening the target feature subset related to the classification task at the current moment through multiple decision trees in a sliding window; According to the target feature subset, dividing the instances into a core feature subset, an auxiliary feature subset, and a complete feature subset, and assigning weighted training weights to different feature subsets based on feature value, and generating corresponding training samples respectively; Initialize the core feature classifier, the auxiliary feature classifier, and the complete feature classifier. Respectively, use the corresponding training samples as training data for training, and assign initial prediction weights to each classifier to obtain each trained classifier; Add each trained classifier to the ensemble learning framework, and dynamically adjust the ensemble prediction weights according to the classifier type; In the ensemble learning framework, when a classifier correctly predicts a labeled sample, dynamically increase its ensemble prediction weight according to a preset step size; Receive newly arrived unlabeled instances. After completing the missing features through the Gaussian Copula model, use each classifier in the ensemble learning framework for weighted prediction, and perform final classification output on the weighted prediction results.

[0006] Furthermore, the method for filling the missing feature values in the currently arrived instances using the Gaussian Copula model is to iteratively optimize the correlation matrix through the online EM algorithm and reconstruct the missing feature values.

[0007] Furthermore, the online EM algorithm includes: Buffer the newly arrived instances into a data window with a fixed size; Estimate the monotonic function to map the mixed-type observation data to the latent normal distribution space; Calculate the conditional expectation value of the latent representation of the missing features in the E step; Update the correlation matrix by maximizing the log-likelihood function in the M step; Repeat the above steps until the correlation matrix converges; Reconstruct the missing feature values based on the converged correlation matrix.

[0008] Furthermore, the method for screening the target feature subset related to the classification task at the current moment from multiple decision trees in the sliding window includes: Construct multiple decision tree groves based on Hoeffding trees. Each Hoeffding tree continuously collects statistical information of data samples within a time period and calculates the information gain of each feature; When the gain difference between the two features with the maximum information gain exceeds the Hoeffding bound, generate a split node with the feature having the maximum information gain. The calculation formula of the Hoeffding bound is: where ϵ is the Hoeffding bound, is the possible range of the attribute information gain, is the confidence level, represents the number of samples collected within the time period; Integrate the feature selection results of multiple decision tree groves to determine the target feature subset at the current moment.

[0009] Furthermore, the update of the decision tree cluster is achieved through a sliding window: when a newly generated decision tree cluster is added to the sliding window, it replaces the earliest generated decision tree cluster in the window, and each decision tree cluster is generated based on training samples in different time periods.

[0010] Furthermore, the allocation rule of the weighted training weights is as follows: The weight of the core feature subset samples is: ; The weight of the auxiliary feature subset samples is 1; The weight of the complete feature subset samples is: ; where is the number of features in the target feature subset, represents the current moment, is at the moment the target feature subset screened out by the decision tree cluster sliding window, is the base of the natural logarithm.

[0011] Furthermore, the allocation rule of the initial prediction weights is as follows: The initial prediction weight of the core feature classifier: The initial prediction weight of the auxiliary feature classifier is fixed as: The initial prediction weight of the complete feature classifier: where is the base of the natural logarithm.

[0012] Furthermore, the method of weighted prediction using each classifier in the ensemble learning framework includes: where is the prediction result of the ensemble learning framework for the input sample , represents the ensemble learning framework, represents three classifiers, is the input instance to be predicted, represents the sample index in the core feature classifier, represents the sample index in the auxiliary feature classifier, represents the sample index in the complete feature classifier, 、 、 correspond to the prediction weights of the three classifiers respectively; , , respectively represent samples of three types of feature subsets, denotes the th core feature classifier's prediction result for the core feature subset , is the set of core feature classifiers, including classifiers, denotes the th auxiliary feature classifier's prediction result for the auxiliary feature subset , is the set of auxiliary feature classifiers, including classifiers; denotes the th complete feature classifier's prediction result for the complete feature subset , is the set of complete feature classifiers, including classifiers.

[0013] To achieve the above object, the second aspect of the present invention provides an electronic device, including a memory and a processor. The memory is used to store a program for supporting the processor to execute the data flow dynamic feature selection method, and the processor is configured to execute the program stored in the memory.

[0014] To achieve the above object, the third aspect of the present invention provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, it executes the steps of the data flow dynamic feature selection method.

[0015] Advantages of the present invention: Compared with the prior art, a method, an electronic device and a medium for dynamically selecting data flow features provided by the present invention effectively solve the dynamic feature challenges in data flow analysis in an open environment through multi-dimensional technology integration. First, the Gaussian Copula model is used to model the mixed data, and the missing features are reconstructed as the expected values in the latent normal distribution space, solving the problems of missing and incomplete feature values and capturing the correlation of dynamically changing features. Secondly, a feature selection mechanism based on a decision tree sliding window is adopted to avoid the interference of historical feature drift. Subsequently, the instances are divided into core, auxiliary and complete feature subsets through feature value evaluation, higher training weights are assigned to the core features, and an integrated framework E3C composed of three types of classifiers is constructed. This framework introduces a dynamic adjustment mechanism for classifier weights: the initial weights are exponentially distributed based on the dimension of the feature subset, and the weights are adjusted in real time through the correct prediction feedback of the labeled samples, enabling the integrated model to adaptively strengthen the decision-making influence of high-contribution classifiers. In addition, the three types of classifiers respectively retain the discriminant knowledge of the core features, the potential associations of the auxiliary features and the global information of the complete features, and realize multi-dimensional knowledge complementarity through weighted integrated prediction, avoiding both the information loss of traditional feature selection and the high computational cost of variable feature space methods. Finally, this method gets rid of the strong dependence on real labels and can autonomously adapt to feature drift through the online EM and sliding window mechanisms with only a small amount of labeled data, significantly reducing the label cost in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments.

[0017] Figure 1 It is a flowchart of a method for dynamically selecting data flow features disclosed in an embodiment of the present invention.

[0018] Figure 2 It is a schematic diagram of the framework principle of a method for dynamically selecting data flow features disclosed in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] In order to enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0020] According to an embodiment of the present invention, it should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the following manufacturing method, in some cases, the steps shown or described can be executed in a different order than here.

[0021] As Figure 1 , Figure 2 shown, the present invention provides a method for dynamically selecting data flow characteristics, including the following steps: Step S100: Receive the currently arriving instance in the data flow, and use the Gaussian Copula model to fill in the missing feature values in the currently arriving instance to obtain the processed instance data; During the data flow learning process, the data instances may have missing feature values due to unstable environment or sensor problems. To ensure that subsequent feature selection and classifier training can be processed based on complete feature data, this step uses the Gaussian Copula model to fill in the missing feature values. Through this method, the integrity of the data can be restored and the correlation information between features can be retained in the case of missing features.

[0022] Specifically, the operation of this step can be divided into the following sub-steps: In the data flow scenario, the instances arrive one by one in chronological order. Assume that each instance in the data flow contains features, that is,

[0023] , where some of the feature values may be missing, and the missing features will affect the accuracy of subsequent feature selection and model training. To address the missing value problem, the present invention processes these data through Gaussian Copula modeling technology. Gaussian Copula is an effective tool for modeling multivariate distributions. In Gaussian Copula modeling, first, the mixed data (instances containing missing feature values) is mapped to the latent standard normal distribution space. The mapping process is through a monotonic function to map the normal distribution in the latent space to the data in the observation space. The specific mapping relationship is: where, represents the observed feature data vector, is the latent normal distribution vector, is the correlation matrix, which describes the correlation between features; represents a d-dimensional multivariate normal distribution.

[0024] Gaussian Copula modeling lies in the correlation matrix , which is used to describe the dependencies between different features. In the latent space, the correlations of the features are modeled through the covariance matrix, thus capturing the correlation structure between the features. This enables the estimation of the missing feature values from the known feature values and the correlation matrix when some feature values are missing.

[0025] By Gaussian Copula modeling of the data, the missing feature values can be estimated in the latent space using the correlation matrix and the observed feature values. Specifically, the online EM algorithm is adopted to continuously optimize the correlation matrix , thereby gradually improving the estimation accuracy of the missing features. The EM algorithm consists of two steps: E-step (expectation step): Calculate the conditional expectation of the latent representation of the missing features, given the observed features and the correlation matrix at the previous time .

[0026] Among them, is the estimated value of the missing feature, is the combination of two expectation calculations. The first expectation represents estimating the value of the missing feature given the observed latent representation and the correlation matrix, and the second expectation is further estimated based on the observed features and the correlation matrix to obtain the final estimated value of the missing feature, is the observed latent representation, is the missing latent representation, is the correlation matrix, describing the correlations between the features, is the observed feature values (features available in the data stream), represents the expected value, used to estimate the missing value.

[0027] This step infers the possible values of the missing features through conditional expectation.

[0028] M-step (maximization step): Based on the known observed data and latent representation, update the correlation matrix by maximizing the likelihood function.

[0029] The likelihood function is defined as: Among them, is the updated correlation matrix, is the correlation matrix calculated at the previous moment t - 1 and is used as the initial value at the current moment, represents the observed data set, which is a set of instances containing all features at the current moment; represents solving to make the objective function maximized , that is, finding the optimal correlation matrix , is a constant, is the correlation matrix 's logarithmic determinant, is the trace of the matrix, which represents the trace of the product of the inverse of the correlation matrix and a certain matrix . Here, the trace is the sum of the diagonal elements of the matrix, is the conditional expectation based on the observed data, is the inverse matrix of the correlation matrix.

[0030] This step optimizes the correlation matrix between features, thereby improving the estimation ability of missing features.

[0031] Once the data is processed through Gaussian Copula modeling, the missing feature values will be filled, and the processed data obtained will be complete feature data. At this time, each instance in the data stream is filled as a complete data vector containing all feature values, ready for subsequent steps of feature selection and model training.

[0032] Step S200: Based on the processed instance data, screen out the target feature subset related to the classification task at the current moment through multiple decision trees in a sliding window; In this step, multiple decision trees in a sliding window are used to screen out the target feature subset related to the classification task at the current moment. The core purpose of this step is to dynamically select the most relevant feature subset according to the instances in the data stream, thereby avoiding the negative impact of redundant features on the model performance. For this purpose, the present invention proposes a feature selection method based on a decision tree forest sliding window, which can efficiently perform online feature selection in a data stream environment and continuously update the feature selection results as the data stream continues to arrive.

[0033] Specifically, the operations of this step can be divided into the following sub-steps: To adapt to the dynamic characteristics of the data stream, a decision tree forest containing multiple decision trees is first designed. The core unit of the decision tree forest is multiple decision trees, which process and make decisions on the data within the same time period. Each decision tree divides and models the data by selecting different feature subsets, having high diversity. In the present invention, the Hoeffding tree is adopted as the basic model, which can continuously update the model through incremental learning without storing a large amount of data.

[0034] To ensure that the model always focuses on the latest concepts in the data stream, a sliding window mechanism is adopted. The size of the sliding window is fixed. Whenever a new instance arrives, the data within the window is updated, and the oldest sample will be replaced out of the window. The data within the window will be used to train the decision tree forest. This mechanism ensures that the model can continuously adapt to the feature drift in the data stream, avoids the model over-relying on historical data, and maintains a high sensitivity to the data at the current moment.

[0035] During the training process of the decision tree forest, each decision tree will be constructed based on different features of the data samples, and the splitting feature will be selected by calculating the information gain. The information gain measures the contribution of a certain feature to the classification task, and the decision tree selects the feature with the largest gain for data splitting. For each decision tree, the selection of its splitting node is based on the feature distribution and gain value of the training samples at that node.

[0036] Specifically, each time a feature split occurs, calculate the information gain of all features, and then select the feature with the largest information gain and the most significant gain difference for splitting. If the gain difference of the feature exceeds the Hoeffding bound ϵ, then this feature is selected as the feature of the current splitting node. The calculation formula of the Hoeffding bound is as follows: where is the Hoeffding bound, is the possible range of the attribute information gain (usually taking log number of classes), is the confidence level (usually set to 0.05 or 0.1), represents the number of samples collected within the time period; Through the above method, the Hoeffding tree can ensure that even with a limited number of samples, it can make feature selections with a relatively high confidence level.

[0037] In the decision tree forest, each decision tree independently evaluates the importance of features and makes feature selections. During the training process, multiple decision trees will make different selections. Based on the selection results of each tree, a voting mechanism is used to determine which features are the most important for the current classification task. The voting mechanism can effectively reduce the overfitting phenomenon that may occur in a single decision tree.

[0038] Within the sliding window, all decision trees participate in feature selection together. Through the voting results, the target feature subset that is most relevant to the classification task at the current moment is selected. These feature subsets not only represent the features in the current data stream that have the greatest impact on the classification task but also enhance the stability of the feature selection results by integrating multiple decision trees. Finally, based on the voting-based decision results, a target feature subset is selected. For subsequent steps.

[0039] This feature selection method can handle feature drift and missing features while enhancing the diversity and stability of the model by integrating the selection results of multiple decision trees, avoiding potential biases that may be brought by a single feature selection strategy.

[0040] Step S300: According to the target feature subset, divide the instances into a core feature subset, an auxiliary feature subset, and a complete feature subset, assign weighted training weights to different feature subsets based on feature value, and generate corresponding training samples respectively. According to the target feature subset selected in step S200, each instance is divided into three different feature subsets: the core feature subset, the auxiliary feature subset, and the complete feature subset. The purpose is to assign appropriate weighted training weights to each feature subset based on the evaluation of feature value and generate corresponding training samples so that the subsequent classifier can better handle different types of feature data. This step is to solve the problem of redundant features in feature selection and ensure that the classification model can highlight those features that are most valuable for the classification task during training. Specifically: In step S200, the most relevant feature subset at the current moment has been selected through the sliding window mechanism of the decision tree cluster. In this step, based on this target feature subset, each instance in the data stream is divided into the following three parts: Core feature subset : Core features are those features that are determined to have the greatest influence on the classification task during the feature selection process. Through these features, the classifier can make decisions efficiently.

[0041] Auxiliary feature subset : Features that have a relatively low correlation with the classification task but may still provide auxiliary information for the model. The addition of auxiliary features can help the model handle more situations and avoid overfitting problems caused by overly single information.

[0042] Complete feature subset : The complete data set containing all features, including core features and auxiliary features. The complete feature set contains all the information of the instance and can provide all available input data for the classification model.

[0043] In feature selection, different feature subsets have different importance. To enable the classifier to learn effectively based on the value of features, weighted training weights are assigned to each feature subset. The basis for weighting is the value of the features, that is, the contribution degree of each feature to the subset. The value of features is evaluated based on their impact on the classification task. For example, the weighted training weights of core features are higher because they directly determine the classification accuracy, while the weights of auxiliary features are lower. Specifically, the evaluation methods of feature value include the following aspects: Information gain: Evaluate the information gain brought by features in the classification task. The greater the information gain, the greater the contribution of the feature to the classification result, and higher weights should be assigned.

[0044] Association measure: Measure the correlation between features and the target variable (i.e., class label). Features with higher correlation are regarded as more important.

[0045] Prediction ability: Evaluate the prediction ability of features in the existing model. Features with higher importance can significantly improve the prediction accuracy of the classification model.

[0046] Based on the above evaluation results, core features will obtain higher weights, and auxiliary features will obtain lower weights, thus forming different training samples, which are respectively used for subsequent classifier training.

[0047] After dividing the instances in the data stream into different feature subsets, corresponding training samples will be generated for each subset. These training samples will be used to train different classifiers: Core feature training samples : The core feature subset contains the features most relevant to the classification task. The training samples constructed based on these features will be used to train the core feature classifier .

[0048] Auxiliary feature training samples : The auxiliary feature subset contains those features that have relatively little impact on the classification task but may still provide useful information. Therefore, the auxiliary feature classifier will be trained based on these features to help the model handle other complex situations.

[0049] Full feature training samples : The full feature subset contains all features and is used to generate the full feature classifier , and the full feature classifier will utilize all the feature information to provide the most comprehensive classification ability.

[0050] During the generation of training samples, weighted training weights are assigned to each training sample according to the importance evaluation results of feature subsets. The specific weight setting can be carried out in the following way: for the core feature subset, a higher training weight is assigned so that the classifier can pay more attention to the features that have the greatest impact on the classification task. For the auxiliary feature subset, a lower training weight is assigned to reflect its relatively minor role in the classification task. For the complete feature subset, the weight may be moderate to ensure that important core information is not lost when considering all feature information.

[0051] Among them, 、 、 represent samples of three types of feature subsets respectively, represents the current moment, is the target feature subset filtered by the decision tree sliding window at time , is the base of the natural logarithm; represents the data set except feature subset, represents the data set.

[0052] Through the above steps, the instances in the data stream are reasonably divided into core features, auxiliary features and complete feature subsets, and appropriate weighted training weights are assigned according to the value of the features. The goal of this process is to optimize the learning process of the model, ensure that the classifier can be efficiently trained according to the actual importance of the features, avoid the interference of redundant or irrelevant features at the same time, and improve the accuracy and robustness of the classification model.

[0053] Step S400: Initialize the core feature classifier, auxiliary feature classifier and complete feature classifier, train them respectively using the corresponding training samples as training data, and assign initial prediction weights to each classifier to obtain the trained classifiers; According to the three types of feature subsets (core features, auxiliary features and complete features) divided in step S300, three classifiers are initialized. When initializing the classifiers, the three classifiers will be constructed according to their corresponding feature subsets and training samples, and different training strategies will be specified for each classifier. After initializing the classifiers, each classifier will be trained using the training samples of different feature subsets generated in step S300 respectively. During the training process, each classifier learns the classification rules according to the data of its feature subset, and these rules will determine how the classifier makes predictions when facing new instances. In this way, each classifier can be effectively optimized according to the characteristics of its feature subset.

[0054] After initializing and training each classifier, initial prediction weights are assigned to each classifier. The weights reflect the relative importance of each classifier in the ensemble learning framework. The initial prediction weights are set according to the capabilities and importance of the classifiers: The core feature classifier has a higher weight because it is trained based on the most important features and can have the greatest impact on the classification results.

[0055] The auxiliary feature classifier has a lower weight because it is trained based on the most important features and can have the greatest impact on the classification results.

[0056] The complete feature classifier has a moderate weight. Since it is trained using all features, it has strong prediction ability, but it is not necessarily stronger than the core feature classifier.

[0057] Specifically, the weights can be adjusted according to the performance of the classifiers. For example, the core feature classifier may be assigned a relatively large initial weight (such as by calculating information gain or classification accuracy), the initial weight of the auxiliary feature classifier is smaller, and the weight of the complete feature classifier is set at a medium level. This assignment of weights will provide a basis for the subsequent ensemble learning framework.

[0058] After the above training process, the three classifiers (core feature classifier, auxiliary feature classifier, and complete feature classifier) will complete training respectively and obtain training results. These trained classifiers will be able to classify according to different types of features and output corresponding prediction results. Finally, these classifiers will be incorporated into the ensemble learning framework and perform weighted predictions within the framework according to their prediction capabilities.

[0059] Step S500: Add the trained classifiers to the ensemble learning framework and dynamically adjust the ensemble prediction weights according to the classifier types; Incorporate the three trained classifiers into the ensemble learning framework E3C to form an integrated classification system. The ensemble learning framework can reduce the risks of overfitting and underfitting within a larger range and improve the generalization ability of the model by combining the prediction results of multiple classifiers. Specifically, each classifier is responsible for predicting the input instances, but they use different subsets of data features. By adding these three classifiers to the framework, it is ensured that each type of feature can obtain corresponding weights in the final classification decision.

[0060] In the ensemble learning framework, "dynamic adjustment" means that as data streams continue to arrive, the prediction weights of each classifier are adjusted according to their actual prediction performance. This process ensures that in the ensemble learning framework, classifiers with better performance will be assigned more weight, while those with poorer performance will be given less weight.

[0061] The principle of weight adjustment is based on the prediction accuracy, information gain, or other performance evaluation criteria of the classifier. At each classification, the classifier makes a prediction based on the characteristics and context of the current data. The ensemble learning framework then performs a weighted average based on the prediction results of each classifier according to their weights, and finally outputs a comprehensive prediction result.

[0062] In the ensemble learning framework, the ensemble prediction weights determine the contribution of each classifier to the final prediction result. For each newly arrived instance, the ensemble framework uses three classifiers to make predictions respectively, and fuses these prediction results according to their weights. The formula for weighted fusion is: where, is the prediction result of the ensemble learning framework for the input sample , represents the ensemble learning framework, represents three classifiers, is the input instance to be predicted, represents the sample index in the core feature classifier, represents the sample index in the auxiliary feature classifier, represents the sample index in the complete feature classifier, , , correspond to the prediction weights of the three classifiers respectively; , , represent the samples of three types of feature subsets respectively, represents the th prediction result of the core feature classifier for the core feature subset , is the set of core feature classifiers, including classifiers, represents the th prediction result of the auxiliary feature classifier for the auxiliary feature subset , is the set of auxiliary feature classifiers, including classifiers; represents the th prediction result of the complete feature classifier for the complete feature subset The prediction results is a complete set of feature classifiers, including classifiers.

[0063] Through this weighted fusion method, the ensemble learning framework can comprehensively consider the prediction results of all classifiers, avoiding the influence of a single classifier due to overfitting or bias on the final decision. At the same time, this weighting mechanism also allows the framework to dynamically adjust the weights of each classifier during the prediction process according to their performance, thereby improving the adaptability and accuracy of the model.

[0064] To achieve dynamic adjustment of the prediction weights, the ensemble framework will regularly update the weights of each classifier according to their prediction performance. The weight adjustment strategy is based on the following factors: If a classifier has a high prediction accuracy within a certain time period, its weight will increase, indicating that the classifier performs well in the current data stream.

[0065] If the features used by a classifier provide more information gain in the classification task, its prediction results will also be given higher weights.

[0066] If a classifier can show stable prediction ability on multiple data stream instances, its weight will also be increased accordingly. Conversely, if there are large fluctuations in the prediction results of the classifier, the weight may be reduced.

[0067] The way of weight update can adopt a step-by-step adjustment strategy. Whenever a classifier correctly predicts a labeled instance, its weight will increase proportionally, while for a classifier with incorrect predictions, its weight will be appropriately reduced. This dynamic adjustment mechanism ensures that the ensemble learning framework always gives priority to the best-performing classifiers.

[0068] By dynamically adjusting the weights, the ensemble learning framework can continuously adapt to feature drift, feature loss, and changes in sample distribution in the data stream environment. As the data stream changes continuously, some classifiers may become more effective, while others may gradually lose their prediction ability. By continuously adjusting the prediction weights of the classifiers, the ensemble framework can ensure that the most appropriate classifier is used for prediction at each moment, thereby ensuring the accuracy and robustness of the final prediction results.

[0069] In addition, the ensemble learning framework can effectively fuse various types of features (such as core features, auxiliary features, and complete features), make full use of various types of information in the data stream, and avoid information loss that may be caused by a single feature selection method. In this way, the ensemble learning framework can provide more stable and accurate classification predictions when dealing with complex and dynamically changing data streams.

[0070] Step S600: In the ensemble learning framework, when a classifier correctly predicts a labeled sample, increment its ensemble prediction weight dynamically according to a preset step size. In the ensemble learning framework, each classifier makes predictions on the unlabeled samples received and compares them with the actual labels. The criterion for judging whether the prediction is correct is determined by calculating the consistency between the output result of the classifier and the actual label. If the classifier correctly predicts the label of the sample, it means that the performance of the classifier at the current moment is effective. On the contrary, incorrect prediction will lead to a decrease in its weight.

[0071] This judgment criterion ensures that the ensemble learning framework can accurately evaluate the performance of each prediction of the classifier. Classifiers with correct predictions will be rewarded (weight increased), while classifiers with incorrect predictions will be punished (weight decreased).

[0072] After each correct prediction by the classifier, the ensemble learning framework will increment the prediction weight of the classifier dynamically according to the preset step size. The preset step size (denoted as ) controls the amplitude of each weight increment. A smaller step size means a more stable weight increment, while a larger step size may cause a more drastic change in the weight. The formula for weight increment can be expressed as: where, is the current weight of the classifier, is the preset step size, is the updated weight.

[0073] By dynamically adjusting the weights of the classifiers, the ensemble learning framework can ensure that when facing new data, it can give priority to the classifiers with excellent performance. As the data stream changes continuously, the prediction accuracy of the classifiers may also fluctuate accordingly. Therefore, through the mechanism of incrementing weights, the influence of each classifier can be effectively balanced dynamically. For example, if a classifier makes accurate predictions continuously for multiple times, its weight will increase continuously, which means that the performance of this classifier in the current data stream is better and it should bear more prediction responsibilities.

[0074] As the weights of the classifiers are continuously adjusted, the classifiers with excellent performance will account for a larger proportion in the ensemble learning framework, thus ensuring the prediction accuracy of the framework. As the weights increase, these classifiers will have more say in future decisions, further improving the stability and robustness of the ensemble model. For example, when new concepts or feature drifts appear in the data stream, some classifiers may become no longer suitable for these changes, resulting in a decrease in their prediction accuracy. Through the mechanism of dynamically incrementing weights, the ensemble framework can automatically reduce the influence of these classifiers, so as to allocate more resources to the classifiers with better performance, and thus can quickly adapt to the new data pattern.

[0075] Although the increasing weights help enhance the influence of high-performance classifiers, over-reliance on a single classifier may lead to overfitting problems in the ensemble model. Therefore, the ensemble learning framework also needs to avoid the excessive inflation of the weights of a certain classifier through certain restrictive or regularization means. For example, setting an upper limit on the weights of each classifier to prevent the weight of a certain classifier from being too large, resulting in its overly strong dominant effect on the final prediction result.

[0076] Step S700: Receive the newly arrived unlabeled instances. After completing the missing feature imputation through the Gaussian Copula model, use each classifier in the ensemble learning framework for weighted prediction, and perform final classification output on the weighted prediction results.

[0077] Through the mapping process of the Gaussian Copula model, map the missing features to the latent normal distribution space, and infer the values of the missing features based on the known features and the correlation matrix. Once the missing features are imputed, the complete feature data can be input into the ensemble learning framework. In the ensemble learning framework, there are multiple classifiers (such as the core feature classifier, the auxiliary feature classifier, and the complete feature classifier). Each classifier will predict the newly arrived instances respectively according to the feature subsets in its training process. Each classifier generates a prediction result.

[0078] Once all classifiers have completed their predictions, the ensemble learning framework will combine the prediction results of each classifier for weighted fusion. The weighted fusion process ensures that the contributions of different classifiers are combined according to their prediction weights. Through weighted fusion, the ensemble framework can integrate the advantages of each classifier, reduce the possible biases of a certain classifier, and make the final prediction result more stable and accurate.

[0079] After completing the weighted prediction, the ensemble learning framework will output the final classification result, which is the synthesis of the weighted prediction results of all classifiers. The final predicted class is generated through weighted voting, weighted averaging, or other appropriate strategies, and the finally output predicted class is the classification result of the new instance.

[0080] In this way, the ensemble learning framework can not only utilize the prediction capabilities of all classifiers but also dynamically adjust according to the weights of each classifier to ensure that the framework can make optimal classification decisions when dealing with different types of feature data.

[0081] In this embodiment, the weight design has an important impact on the algorithm performance proposed by the present invention. The classifiers in E3C mainly have two weights. The first is the integrated prediction weight of the classifier, and the second is the weighted training weight based on different sample types when training the classifier. When designing the integrated prediction weight of the classifier, on the one hand, appropriate weights need to be assigned according to the nature of different classifiers. For example, classifiers trained based on core features usually have higher classification ability than those trained based on auxiliary features. On the other hand, the weights need to be dynamically adjusted during the algorithm operation, giving higher weights to classifiers that contribute significantly to the performance to improve the overall prediction ability of the integrated model. Whenever the classifier is initialized, the initial weight values of different types of classifiers are set as follows: Among them, is the number of features of the target feature subset, represents the current moment, is at the moment the target feature subset screened by the decision tree sliding window, is the base of the natural logarithm.

[0082] To balance the prediction weight ratio between different classifiers, an upper limit of the integrated prediction weight of the classifier containing core features is set. In addition, whenever the classifier correctly predicts a labeled instance, its weight will increase proportionally by (1 + s). Here, s is the preset weight adjustment step size. This mechanism ensures that classifiers with excellent prediction performance obtain more prediction weights.

[0083] The training weight is crucial for the learning performance of the algorithm. Some ensemble learning methods enhance model diversity by increasing the sample training weight and the weight difference. In the training weight design of the present invention, samples containing the core feature subset are given higher training weights, while samples of auxiliary features are set with lower weights. The specific settings are as follows: As can be seen from the above, the present invention solves the problems of feature drift, feature loss, and knowledge retention through multi-dimensional collaborative innovation. First, the hybrid data modeling based on Gaussian Copula dynamically estimates the correlation matrix of the latent normal space through the online EM algorithm, and reconstructs the missing values by using the dependence relationship between features, overcoming the limitation of the traditional imputation method's independent assumption of features, and effectively dealing with mixed data types and missing problems. Second, the decision tree grove sliding window feature selection mechanism constructs multiple Hoeffding trees in parallel and adopts a sliding window update strategy, improving diversity while ensuring the rigor of feature selection: the grove structure covers more candidate features through multiple splits, and the sliding window dynamically replaces the old model to ensure adaptation to the latest data distribution, avoiding the knowledge fragmentation caused by the traditional feature selection method due to static feature subsets. In addition, the feature value-driven instance weighting framework innovatively divides the feature space into three types of subsets: core, auxiliary, and complete, trains classifiers respectively and assigns dynamic weights: the core feature classifier amplifies the contribution of important features through squared weights, the auxiliary classifier retains potential association information, and the complete feature classifier captures global dependencies. The three form a synergistic effect through the adaptive adjustment of the integrated prediction weights (such as incrementing by step size when predicting correctly), alleviating both the knowledge loss caused by training with fixed feature subsets and avoiding the high computational cost of variable feature space methods. At the same time, the unsupervised feature drift detection mechanism evaluates the change of feature importance through the online updated decision tree grove, replacing the traditional drift detection method that relies on true labels, and significantly reducing the annotation cost. Finally, the method achieves a balance among feature dynamics, computational efficiency, and knowledge retention, improving the robustness and accuracy of data stream classification in an open environment.

[0084] Taking the abnormal behavior recognition in the Internet of Things scenario as an example. First, a set of abnormal behavior data sets in the Internet of Things scenario is collected, and each instance of the abnormal behavior data set contains 30 features.

[0085] Whenever an instance arrives, first fill in the missing feature values of the instance through Gaussian Copula modeling. Then, perform feature importance evaluation on the decision tree grove sliding window for feature importance evaluation, and output the important feature subset of this instance. . Based on the important feature subset at the current moment , divide the training data set into different parts. For example, for the sample , its complete feature space is Use the important feature subset to form the core feature training samples , and use the remaining features to form the auxiliary feature training samples . After that, the core feature classifier uses the core feature training samples for training, and the auxiliary feature classifier Auxiliary feature-based training samples are trained, while the complete feature classifier is trained based on the training samples in the complete feature space. After that, the three classifiers are added to the ensemble learning framework E3C. After the ensemble framework is trained, the current ensemble learning framework E3C is used to predict unknown instances to determine whether there is abnormal behavior in the recognized data instance.

[0086] Through the data flow dynamic feature selection method proposed by the present invention, abnormal behavior recognition and classification in the Internet of Things scenario can be performed to identify the data with abnormal behavior in the Internet of Things scenario. In addition, it can be extended to detect network attack data with probability drift problems.

[0087] According to another aspect of the embodiments of the present application, an electronic device is further provided, including a processor and a memory. The processor is used to implement the steps of the method when executing the computer program stored in the memory.

[0088] In the above embodiments of the present invention, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0089] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0090] In addition, each functional unit in the various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0091] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0092] The foregoing are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for dynamically selecting data flow characteristics, characterized in that It includes the following steps: Receive the currently arrived instance in the data stream, fill in the missing feature values in the currently arrived instance using the Gaussian Copula model to obtain the processed instance data; Based on the processed instance data, screen the target feature subset related to the classification task at the current moment through multiple decision trees in the sliding window; According to the target feature subset, divide the instances into a core feature subset, an auxiliary feature subset, and a complete feature subset, assign weighted training weights to different feature subsets based on feature value, and generate corresponding training samples respectively; Initialize the core feature classifier, the auxiliary feature classifier, and the complete feature classifier, respectively use the corresponding training samples as training data for training, and assign initial prediction weights to each classifier to obtain the trained classifiers; Add the trained classifiers to the ensemble learning framework and dynamically adjust the ensemble prediction weights according to the classifier type; In the ensemble learning framework, when the classifier correctly predicts the labeled sample, dynamically increase its ensemble prediction weight according to the preset step size; Receive the newly arrived unlabeled instance, after filling in the missing features through the Gaussian Copula model, use each classifier in the ensemble learning framework for weighted prediction, and perform final classification output on the weighted prediction results.

2. The data flow dynamic feature selection method according to claim 1, wherein The method of filling in the missing feature values in the currently arrived instance using the Gaussian Copula model is to iteratively optimize the correlation matrix through the online EM algorithm and reconstruct the missing feature values.

3. The data flow dynamic feature selection method according to claim 2, wherein The online EM algorithm includes: Buffer the newly arrived instance into a data window with a fixed size; Estimate the monotonic function to map the mixed observation data into the latent normal distribution space; Calculate the conditional expectation of the latent representation of the missing feature in the E step; Update the correlation matrix by maximizing the log-likelihood function in the M step; Repeat the above steps until the correlation matrix converges; Reconstruct the missing feature values based on the converged correlation matrix.

4. The data flow dynamic feature selection method according to claim 1, wherein The method of screening the target feature subset related to the classification task at the current moment through multiple decision trees in the sliding window includes: Construct multiple decision tree groves based on Hoeffding trees. Each Hoeffding tree continuously collects statistical information of data samples within the time period and calculates the information gain of each feature; When the gain difference between the two features with the maximum information gain exceeds the Hoeffding bound, generate a split node with the feature with the maximum information gain. The calculation formula of the Hoeffding bound is: wherein, is the Hoeffding bound, is the possible range of the attribute information gain, is the confidence level, represents the number of samples collected within the time period; Integrate the feature selection results of multiple decision tree groves to determine the target feature subset at the current moment.

5. The data flow dynamic feature selection method according to claim 4, wherein The update of the decision tree grove is realized through the sliding window: when the newly generated decision tree grove is added to the sliding window, replace the earliest generated decision tree grove in the window, and each decision tree grove is generated based on the training samples in different time periods.

6. The data flow dynamic feature selection method according to claim 1, characterized in that The allocation rule of the weighted training weight is: The weights of the samples in the core feature subset are as follows: ; The weight of the auxiliary feature subset sample is 1; The weight of the complete feature subset sample is as follows: ; Among them, is the number of features of the target feature subset, represents the current moment, at the moment is the target feature subset screened by sliding the decision tree through the window, is the base of the natural logarithm.

7. The method for selecting dynamic characteristics of data flow according to claim 6, wherein The allocation rule of the initial prediction weight is: The initial prediction weight of the core feature classifier: The initial prediction weight of the auxiliary feature classifier is fixed as: The initial prediction weight of the complete feature classifier: wherein, is the base of the natural logarithm.

8. The data flow dynamic feature selection method according to claim 1, wherein The method of using each classifier in the ensemble learning framework for weighted prediction includes: Among them, is an ensemble learning framework for the prediction result of the input sample . represents the ensemble learning framework, representing three classifiers, is the input instance to be predicted, represents the sample index in the core feature classifier, represents the sample index in the auxiliary feature classifier, represents the sample index in the complete feature classifier, , , correspond to the prediction weights of the three classifiers respectively; , , represent the samples of three types of feature subsets respectively, represents the th core feature classifier's prediction result for the core feature subset , is the set of core feature classifiers, including classifiers, represents the th auxiliary feature classifier's prediction result for the auxiliary feature subset , is the set of auxiliary feature classifiers, including classifiers; represents the th complete feature classifier's prediction result for the complete feature subset , is the set of complete feature classifiers, including classifiers.

9. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program for supporting the processor to execute the data flow dynamic feature selection method according to any one of claims 1-8, and the processor is configured to execute the program stored in the memory.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run by the processor, it executes the steps of the data flow dynamic feature selection method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Adaptive network flow concept drift detection method based on information entropy

    CN110445726A

  • Data flow classification algorithm based on AAE-DWMAL-LearnNSE (Automatic Assisted Engineering-Discrete Wavelength Multiple Input Multiple Output-LearnNSE)

    CN110647671A

  • Semi-supervised algorithm in mixed online data stream scene

    CN115796301A

  • Multi-feature fusion financial user portrait classification method based on ensemble learning

    CN117992819A

  • Conceptual drift detection and adaptation method based on sub-feature selection

    CN118260687A