An integrated learning classification method and system with concept drift detection function
Through the integrated learning classification method and adaptive window Kolmogorov-Smirnov early warning detection, real-time monitoring and response to concept drift in network traffic data, the problem of model failure and accuracy is solved, and higher adaptability and accuracy are achieved.
Patent Information
- Application Number
- CN202510362980.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-26
AI Technical Summary
The prior art is difficult to effectively detect and cope with concept drift in network traffic data, resulting in model failure and reduced accuracy.
The integrated learning classification method is adopted, combining principal component analysis (PCA) and adaptive window Kolmogorov-Smirnov early warning detection method, and real-time classification and model update are performed through the learning model based on the integrated Hoeffding tree.
Real-time detection and response to concept drift is realized, the adaptability and accuracy of the model is improved, and the risk of false positives caused by short-term fluctuations is reduced.
Smart Images

Figure CN119884959B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of machine learning and data stream processing, and specifically, to an ensemble learning classification method and system with a concept drift detection function. Background Art
[0002] In practical applications, network traffic data changes continuously over time, and new changes may occur in the attack patterns, resulting in the invalidation or obsolescence of the original model. This phenomenon is called concept drift, that is, as time goes by, the data distribution has changed significantly, and the model needs to be updated in real time to maintain accuracy. For the problem of concept drift, traditional static models cannot respond in time, and new methods are needed to dynamically adjust the learning process of the model. Concept drift in network traffic data is particularly obvious, especially in some large network environments where attackers constantly innovate their attack means, resulting in frequent invalidation of the model. Therefore, researching how to effectively detect and respond to concept drift has become a key issue in the field of network security today. Summary of the Invention
[0003] To solve the deficiencies mentioned in the above background art, the purpose of the present invention is to provide an ensemble learning classification method and system with a concept drift detection function.
[0004] In a first aspect, the purpose of the present invention can be achieved through the following technical solutions: An ensemble learning classification method with a concept drift detection function, the method comprising the following steps:
[0005] Obtain high-dimensional data of an extensible sliding window, and perform dimensionality reduction on the high-dimensional data of the extensible sliding window by using principal component analysis PCA to extract key feature data;
[0006] Based on the adaptive window Kolmogorov-Smirnov early warning detection method, perform real-time monitoring on the key feature data to determine whether concept drift occurs. If concept drift occurs, obtain the original data of the adaptive window when concept drift occurs, and automatically expand the adaptive window. If concept drift does not occur, the adaptive window slides normally;
[0007] Input the original data of the adaptive window when concept drift occurs into a pre-established learning model based on an ensemble Hoeffding tree to obtain a trained learning model based on an ensemble Hoeffding tree, and perform real-time classification of the data stream based on the trained learning model based on an ensemble Hoeffding tree.
[0008] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: the pre-established learning model based on the integrated Hoeffding tree is integrated through a master tree and multi-subtree architecture, where the master tree uses the complete metric set for training to ensure global feature coverage, the subtrees generate differential classifiers by randomly selecting some metric dimensions, and both the master tree and the multi-subtrees dynamically determine node splitting based on the Gini gain difference and the Hoeffding bound, and update the tree structure to capture changes in the feature distribution.
[0009] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: the initialization process of the integrated Hoeffding tree:
[0010] Build a model consisting of k empty decision trees T1, T2, …, T k where T1 is the whole tree trained with all metrics, and T2~T k are subtrees trained with partial metrics. The complete tree T1 is trained using the complete metric set ; T2~T k The subtrees randomly select metrics. To ensure that each metric can be selected at least once in T2~T k , set the target metric coverage probability , and the required minimum number of subtrees can be derived through the metric coverage probability formula as follows:
[0011] The metric coverage of a single subtree. The probability that a certain metric is not selected in a single subtree , as shown in formula (1):
[0012] (1)
[0013] The coverage of multiple subtrees. Suppose there are subtrees. Among subtrees, the probability that a certain metric is not selected at all , as shown in formula (2):
[0014] (2)
[0015] Among subtrees, the probability that a certain metric is selected at least once is expressed as formula (3). Substituting formula (2), we get formula (4) as follows:
[0016] (3)
[0017] (4)
[0018] When the coverage probability of the target metric takes a preset value, calculate the condition that a certain metric is selected at least once in T2~T k , and at the same time minimize the number of subtrees required for computing resources .
[0019] Combined with the first aspect, in some implementations of the first aspect, the method further includes: for the pre-constructed learning model based on the integrated Hoeffding tree, determine the optimal region partitioning method for the leaf nodes of each tree, and the steps are as follows:
[0020] Traverse the metric variables: Let the current node sample be S, and the metric variable set be , and traverse each metric variable in turn , where ;
[0021] Generate candidate split points: For each metric variable , sort all the metric values included in the metric variable from small to large, denoted as , the candidate split point set is the median of all adjacent metric values, and the candidate split point is represented as , as shown in formula (5):
[0022] (5)
[0023] where represents the th metric value in the sorted set of the metric variable , represents the number of all metric values of this metric variable , is the index of the candidate split point, and the value range is from 1 to ;
[0024] Calculate the partitioning effect corresponding to the split point: For each candidate split point , split the current node sample S into two sub-regions, as shown in formula (6):
[0025] (6)
[0026] where , are the left and right subsets after partitioning, s is a data instance in the sample set S, is the specific value of the sample instance s on the metric variable ;
[0027] In combination with the first aspect, in some implementations of the first aspect, the method further includes: the optimal selection process of the candidate segmentation points:
[0028] Quantify the purity of the node through the Gini coefficient. For a node, its Gini coefficient is defined as shown in formula (7):
[0029] (7)
[0030] where is the total number of classification categories, is the sample proportion of category in the node, is the Gini coefficient;
[0031] Traverse all candidate segmentation points of each metric variable and calculate the Gini gain after splitting under each segmentation point. The calculation formula of the Gini gain is as shown in (8):
[0032] (8)
[0033] where is the Gini coefficient of the parent node before splitting, representing the sample purity of the parent node, , are the Gini coefficients of the left and right child nodes after splitting, representing the sample purities of the left and right child nodes respectively, , are the sample numbers of the left and right child nodes, is the total sample number of the parent node;
[0034] Select the candidate segmentation point with the largest Gini gain value as the current optimal segmentation point, record the Gini gain value of the sub-optimal segmentation point, and calculate the Gini gain difference between the two to verify the significance of the segmentation, as shown in formula (9):
[0035] (9)
[0036] where is the Gini gain of the optimal segmentation point, is the Gini gain of the sub-optimal segmentation point, is the Gini gain difference between the two segmentation points;
[0037] To ensure the statistical significance of the Gini gain difference, introduce the Hoeffding bound as the confidence constraint condition, which can be expressed as (10):
[0038] (10)
[0039] where is the number of samples, is the detection threshold, is the variable value range length;
[0040] If the Gini gain difference , it is considered that the optimal splitting point is significantly better than the sub - optimal splitting point, and the optimal splitting point is accepted for tree splitting; if the Gini gain difference , it means that the current sample is not sufficient to prove the significance of the splitting point, and it needs to be re - evaluated after data accumulation;
[0041] For the newly generated child nodes, the optimal splitting point selection process is repeatedly executed until all nodes meet the conditions for stopping splitting.
[0042] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: the integrated Hoeffding tree joint decision - making process:
[0043] Based on the pre - established integrated Hoeffding tree model, when the data stream is input into the model for prediction, the prediction results of the main tree T1 and the sub - trees T2~T k are integrated, and the final classification result is generated through a dynamic weighted voting mechanism;
[0044] Decision window initialization: A sliding window Q is used to quantify the historical decision reliability of the sub - trees T2~T k , the sliding window capacity is set to , and the initial state is filled with all 1 values, assuming that the initial decisions of the sub - trees T2~T k are correct;
[0045] Window update rule: Whenever a new sample is input, the sub - trees T2~T k perform weighted voting prediction on the sample. If the prediction is correct, 1 is filled in the latest position of the sliding window, otherwise 0 is filled, where the correct prediction of the sub - trees T2~T k is marked as 1, and the wrong prediction is marked as 0;
[0046] Confidence calculation: When the data stream arrives, the sum of the correct decisions made by the sub - trees T2~T k within the window is counted , which can be expressed by formula (11):
[0047] (11)
[0048] where represents the th decision value stored in the sliding window Q;
[0049] Joint decision: Compare with the preset threshold ;
[0050] If , it is determined that the subtree T2~T k has a reliable decision, and the prediction results of the main tree T1 and the subtrees T2~T are fused using the majority voting strategy; k
[0051] If , only the output of the main tree T1 is used as the final classification result.
[0052] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: performing a dimensionality reduction process on the high-dimensional data of the scalable sliding window by using the principal component analysis PCA:
[0053] The proposed principal component analysis PCA data dimensionality reduction operation requires creating two sliding windows: a fixed-capacity historical window L and a dynamically expandable adaptive window R, where the initial capacity of L is C L , the initial capacity of R is C R and the maximum expansion upper limit R is set MAX , and dimensionality reduction operations are respectively performed on the historical window L and the adaptive window R;
[0054] Normalize the data. First, perform normalization processing on the window data as shown in formula (12):
[0055] (12)
[0056] where is the original data, is the mean of the index, is the standard deviation of the index, is the normalized data;
[0057] Calculate the covariance matrix, calculate the covariance matrix of the normalized data , the covariance matrix represents the correlation between the various indexes in the data, as shown in formula (13):
[0058] (13)
[0059] where is the number of samples in the window, is the transposed matrix of;
[0060] By solving the eigenvalues and corresponding eigenvectors of the covariance matrix, the principal component direction of the data is obtained. The eigenvalue represents the variance of the principal component, and the eigenvector represents the direction of the principal component. According to the eigenvalues arranged in descending order, the first The eigenvectors corresponding to the largest eigenvalues, where the principal components with larger eigenvalues represent a larger proportion of the original information volume of the data, and calculate the variance contribution rate of each principal component , as shown in formula (14):
[0061] (14)
[0062] Wherein, is the eigenvalue of the th principal component after sorting, is the number of selected principal components;
[0063] Project the data. Arrange the selected eigenvectors in columns to form a transformation matrix , and project the standardized data onto the selected principal components, as shown in formula (15):
[0064] (15)
[0065] Wherein, is the data after dimensionality reduction.
[0066] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: The usage process of the adaptive window Kolmogorov-Smirnov early warning detection method is as follows:
[0067] For the data after dimensionality reduction, extract the data of the th dimension of the historical window and the adaptive window respectively for Kolmogorov-Smirnov test, as shown in formula (16):
[0068] (16)
[0069] Wherein, is the data matrix after dimensionality reduction of the historical window, is the data matrix after dimensionality reduction of the adaptive window, is the data of the historical window on the th principal component, is the data of the adaptive window on the th principal component;
[0070] The Kolmogorov-Smirnov test process is as follows:
[0071] Establish the null hypothesis H0: Assume that there is no difference between the two samples and they come from the same distribution;
[0072] Calculate the cumulative distribution function: For each dimension , calculate respectively , 's cumulative distribution function , . The cumulative distribution function describes the probability that a random variable is less than or equal to a preset specific value, that is . Let the sample data set in which the samples are independent and random, then the cumulative distribution function is expressed as formula (17):
[0073] (17)
[0074] where is the indicator function, as shown in formula (18):
[0075] (18)
[0076] Calculate the Kolmogorov - Smirnov statistic , expressed as and the maximum vertical distance between the cumulative distribution functions, which is the maximum difference between the two distributions, as shown in formula (19):
[0077] (19)
[0078] Define the significance level and calculate the value, convert the maximum difference between the data distributions into value. Equation (20) is the calculation expression:
[0079] (20)
[0080] where is the number of effective samples, that is , is the number of samples in is the number of samples in;
[0081] Weighted value: According to the variance contribution rate of each principal component, value is weighted and calculated to obtain the global drift significance score , as shown in formula (21).
[0082] (21)
[0083] In combination with the first aspect, in certain implementations of the first aspect, the method further includes: The process of real-time monitoring of key feature data based on the adaptive window Kolmogorov-Smirnov early warning detection method is as follows:
[0084] For the dimensions of the extracted key feature data, calculate the statistic of the cumulative distribution function difference respectively and the corresponding value, and then obtain the global drift significance score by weighting according to the variance contribution rate of each principal component . Set the significance level as the concept drift determination threshold, and define the early warning threshold ; ;
[0085] If , it is considered that data concept drift has occurred, and the system will clear the historical window data and retrain the ensemble Hoeffding tree model using the adaptive window data;
[0086] If , it is considered that a drift early warning has occurred, and the system automatically enters the elastic buffer stage to dynamically expand the adaptive window to accumulate potential drift samples;
[0087] If , it is considered that no drift has occurred, and the data is continuously updated by sliding the window.
[0088] In the second aspect, in order to achieve the above object, the present invention discloses an ensemble learning classification system with a concept drift detection function, including:
[0089] A data dimensionality reduction module, configured to obtain high-dimensional data of an expandable sliding window, perform dimensionality reduction on the high-dimensional data of the expandable sliding window by using principal component analysis PCA, and extract key feature data;
[0090] A data expansion module, configured to perform real-time monitoring of key feature data based on the adaptive window Kolmogorov-Smirnov early warning detection method, determine whether concept drift has occurred, if concept drift has occurred, obtain the original data of the adaptive window when the concept drift occurs, and automatically expand the adaptive window, if no concept drift has occurred, the adaptive window slides normally;
[0091] A data classification module, configured to input the original data of the adaptive window when the concept drift occurs into a pre-established learning model based on an ensemble Hoeffding tree, obtain a trained learning model based on the ensemble Hoeffding tree, and perform real-time classification of the data stream based on the trained learning model based on the ensemble Hoeffding tree.
[0092] In another aspect of the present invention, in order to achieve the above object, a terminal device is disclosed, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The memory stores a computer program capable of running on the processor. When the processor loads and executes the computer program, an ensemble learning classification method with concept drift detection function as described above is adopted.
[0093] Advantages of the present invention:
[0094] In terms of data dimensionality reduction and feature extraction, the present invention uses principal component analysis (PCA) to perform dimensionality reduction on high-dimensional data within an expandable sliding window and extract key feature data. This method not only effectively reduces the computational overhead, but also retains the main information of the data, improving the efficiency and accuracy of data stream classification.
[0095] In terms of concept drift detection, the present invention is based on the adaptive window Kolmogorov-Smirnov early warning detection method to achieve real-time monitoring of key feature data. When the distribution difference between windows reaches the early warning threshold but has not crossed the drift determination boundary, the system enters the elastic buffer stage, dynamically expands the adaptive window to accumulate more data for further detection. This mechanism can reduce false alarms caused by short-term fluctuations and improve the robustness of drift detection. At the same time, to prevent the problem of computational resource consumption caused by unlimited window growth, a maximum expansion length is set to ensure that the window is adaptively adjusted within a reasonable range and resumes the normal sliding update mechanism after reaching the upper limit.
[0096] In terms of classification model design, the present invention adopts a learning model based on an ensemble Hoeffding tree, and when concept drift occurs, the model is incrementally trained using the original data of the adaptive window. This method combines the collaborative learning mechanism of the main tree and multiple subtrees, not only retaining the key information of the global data features, but also enhancing the adaptability to local data distribution changes. By dynamically adjusting the weights of the subtrees participating in the prediction and combining the joint voting decision of the main tree and subtrees, the influence of abnormal data or local drift on a single model is reduced, improving the stability and accuracy of data stream classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0097] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts;
[0098] Figure 1 It is a schematic flowchart of the method of the present invention;
[0099] Figure 2 This is the structural diagram of the adaptive window Kolmogorov - Smirnov early warning detection window in the present invention;
[0100] Figure 3 This is the decision - making structural diagram of the integrated Hoeffding tree in the present invention;
[0101] Figure 4 This is the schematic diagram of the system structure of the present invention. Specific implementation manners
[0102] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.
[0103] Embodiment 1:
[0104] As Figure 1 shown, an integrated learning classification method with a concept drift detection function includes the following steps:
[0105] S101: Obtain high - dimensional data of an extensible sliding window, and use principal component analysis (PCA) to reduce the dimension of the high - dimensional data of the extensible sliding window, and extract key feature data;
[0106] As Figure 2 shown, Figure 2 is the structural diagram of the adaptive window Kolmogorov - Smirnov early warning detection window. Establish an adaptive window Kolmogorov - Smirnov early warning detection window model, and input the data learned by the integrated Hoeffding tree into the window for dimension reduction operation. The method includes the following steps:
[0107] Specifically, creating two sliding windows is required to reduce the dimension of the high - dimensional data of the extensible sliding window by using principal component analysis (PCA): a fixed - capacity historical window L and a dynamically extensible adaptive window R. The initial capacity of L is C L , and the window index range is expressed as 1, 2,..., n - r; the initial capacity of R is C R , and the window index range is expressed as n - r + 1,..., n - 1, n, and R is provided with a maximum expansion upper limit R MAX . The data stream is preferentially stored in L. When L is full, it is transferred to R. If R fills its initial capacity C RThen the dimensionality reduction and drift detection process is triggered. At the same time, R is allowed to add raw data to the adaptive window expansion bit in the warning state to accumulate potential drift samples. In this example, n is set to 800, r is set to 300, and the maximum expansion limit R of R MAX is set to 600;
[0108] Normalize the data. First, normalize the window data as shown in formula (12):
[0109] (12)
[0110] where is the raw data, is the mean of the index, is the standard deviation of the index, is the data after normalization;
[0111] Calculate the covariance matrix. Calculate the covariance matrix of the normalized data The covariance matrix represents the correlation between the various indicators in the data, as shown in formula (13):
[0112] (13)
[0113] where, is the number of samples in the window, is The transpose matrix of;
[0114] By solving the eigenvalues and corresponding eigenvectors of the covariance matrix, obtain the principal component direction of the data. The eigenvalue represents the variance of the principal component, and the eigenvector represents the direction of the principal component. According to the descending order of the eigenvalues, select the first The eigenvectors corresponding to the k largest eigenvalues, where the principal component with a large eigenvalue represents a large proportion of the original information volume of the data, and calculate the variance contribution rate of each principal component as shown in formula (14):
[0115] (14)
[0116] where, is the eigenvalue of the i-th principal component after sorting, is the number of selected principal components. In this example, k takes the value of 3;
[0117] Project the data. Arrange the selected k eigenvectors in columns to form a transformation matrix and project the normalized data onto the selected principal components, as shown in formula (15):
[0118] (15)
[0119] Among them, is the data after dimensionality reduction.
[0120] S102: Based on the adaptive window Kolmogorov - Smirnov early warning detection method, real - time monitoring is carried out on the key feature data to judge whether concept drift occurs. If concept drift occurs, the original data of the adaptive window at the time of concept drift is obtained, and the adaptive window is automatically expanded. If concept drift does not occur, the adaptive window slides normally;
[0121] Specifically, in this embodiment, the drift detection of sample data is realized through the adaptive window Kolmogorov - Smirnov early warning detection method.
[0122] For the data after dimensionality reduction , the data of the th dimension of the historical window and the adaptive window are respectively extracted for Kolmogorov - Smirnov test, as shown in formula (16):
[0123] (16)
[0124] Among them, is the data matrix after dimensionality reduction of the historical window, is the data matrix after dimensionality reduction of the adaptive window, is the data of the historical window on the th principal component, is the data of the adaptive window on the th principal component;
[0125] The Kolmogorov - Smirnov test process is as follows:
[0126] Establish the null hypothesis H0: Assume that there is no difference between the two samples and they come from the same distribution;
[0127] Calculate the cumulative distribution function: For each dimension , calculate , 's cumulative distribution function , . The cumulative distribution function describes the probability that a random variable is less than or equal to a preset specific value, that is, . Assume the sample data set 's samples are all independent and random, then the cumulative distribution function is expressed as formula (17):
[0128] (17)
[0129] Among them, is an indicator function, as shown in formula (18):
[0130] (18)
[0131] Calculate the Kolmogorov - Smirnov statistic , expressed as and the maximum vertical distance between the cumulative distribution functions, which is the maximum difference between the two distributions, as shown in formula (19):
[0132] (19)
[0133] Define the significance level and calculate the value, and convert the maximum difference between the data distributions into the value. Equation (20) is the calculation expression:
[0134] (20)
[0135] Where is the effective sample size, that is , is the number of samples in is the number of samples in ;
[0136] Weighted value: According to the variance contribution rate of each principal component, the value is weighted and calculated to obtain the global drift significance score, as shown in formula (21).
[0137] (21)
[0138] In this embodiment, a double - threshold progressive drift detection mechanism is used to determine whether the sample data drifts. Set the significance level as the concept drift determination threshold, and define the warning threshold . In this example, is set to 0.05, is set to 0.1.
[0139] If , it is considered that data concept drift has occurred. The system will clear the historical window data and retrain the integrated Hoeffding tree model using the adaptive window data;
[0140] If , it is considered that a drift warning has occurred. The system automatically enters the elastic buffer stage and dynamically expands the adaptive window to accumulate potential drift samples;
[0141] If , it is considered that no drift has occurred, and the window is continued to slide to update the data.
[0142] S103: Input the original data of the adaptive window when concept drift occurs into a pre-established learning model based on the integrated Hoeffding tree to obtain a trained learning model based on the integrated Hoeffding tree, and perform real-time classification of the data stream based on the trained learning model based on the integrated Hoeffding tree;
[0143] The pre-established learning model based on the integrated Hoeffding tree improves the classification performance through an ensemble learning strategy. Among them, the Hoeffding tree is an incremental learning model suitable for stream data classification, which can efficiently approximate the optimal decision tree with limited samples without storing historical data. On this basis, the present invention introduces an ensemble learning strategy to construct a main tree and a multi-subtree architecture to improve the model's adaptability to concept drift. Among them, the main tree is trained using the complete index set to ensure global feature coverage, and the subtrees generate differentiated classifiers by randomly selecting some index dimensions. Both the main tree and the multi-subtrees dynamically judge node splitting based on the Gini gain difference and the Hoeffding bound, and update the tree structure to capture the change of the feature distribution;
[0144] Specifically, in this embodiment, experiments are carried out through an intrusion detection data set. The intrusion detection data set contains benign and the latest common attacks, and is collected from a real network environment. For the intrusion detection data set, it constructs the abstract behaviors of 25 users based on the HTTP, HTTPS, FTP, SSH, and email protocols. The data collection ends at a certain moment on a certain day. Since the attack patterns in the data set change over time, multiple concept drifts occur in the data set. After adopting the k-means clustering sampling method, 28,303 representative records in it are used to evaluate the model. The selected key traffic indicators are feature_list[Flow Duration, Total Length of Fwd Packets, Fwd Packet LengthMax,..., Flow IAT Min], and the attribute column is the classification label of the flow data, with values of "normal traffic" or "abnormal traffic". The network traffic data stream In the input integrated Hoeffding tree, the streaming data sample is represented as , where is the key metric vector of the classification model, and is the classification attribute of the streaming data. The integrated Hoeffding tree classification module realizes real-time and efficient classification of streaming data, and the steps are as follows:
[0145] Initialize the integrated Hoeffding tree and establish a model composed of k empty decision trees T1, T2, …, T k , where T1 is the complete tree trained with all metrics, and T2~T k are subtrees trained with partial metrics. The complete tree T1 is trained using the complete metric set ; T2~T k subtrees randomly select metrics. This design enhances the diversity of the model and reduces the risk of overfitting. To minimize the computational cost, the number of subtrees needs to meet the following conditions to ensure that each metric can be selected at least once in T2~T k , and at the same time minimize the selection of duplicate metrics. A target metric coverage probability needs to be set, with a value of 99%. The required number of subtrees can be derived through the metric coverage probability formula as follows:
[0146] The metric coverage of a single subtree. The probability that a certain metric is not selected in a single subtree , as shown in formula (1):
[0147] (1)
[0148] The coverage of multiple subtrees. Suppose there are subtrees. In subtrees, the probability that a certain metric is not selected at all , as shown in formula (2):
[0149] (2)
[0150] In subtrees, the probability that a certain metric is selected at least once is expressed as formula (3). Substituting formula (2), we get formula (4) as follows:
[0151] (3)
[0152] (4)
[0153] At the target metric coverage probability When the value is 99%, it is possible to calculate the condition that a certain index is selected at least once in T2~T k while minimizing the number of subtrees required for computing resources ;
[0154] Partition the index and the splitting point. As the data stream samples continue to increase, for the pre-constructed learning model based on the integrated Hoeffding tree, determine the best regional partitioning method for the leaf nodes of each tree. The specific steps are as follows:
[0155] Traverse the index variables: Let the current node sample be S, and the index variable set be , and traverse each index variable in turn , where ;
[0156] Generate candidate splitting points: For each index variable , sort all the index values included in the index variable from smallest to largest, denoted as . The candidate splitting point set is the median of all adjacent index values, and the candidate splitting point is represented as , as shown in formula (5):
[0157] (5)
[0158] where represents the th index value in the sorted set of the index variable , represents the number of all index values of this index variable , is the index of the candidate splitting point, and the value range is from 1 to ;
[0159] Calculate the partitioning effect corresponding to the splitting point: For each candidate splitting point , split the current node sample S into two sub-regions, as shown in formula (6):
[0160] (6)
[0161] where , are the left and right subsets after partitioning, s is a data instance in the sample set S, is the specific value of the sample instance s on the index variable ;
[0162] Select the optimal splitting point. To ensure that most of the samples in the splitting node come from the same category and to make the sample attributes in the splitting node less diverse, the node purity needs to be improved. We use the Gini coefficient to quantify the node purity. For a node, its Gini coefficient is defined as shown in formula (7):
[0163] (7)
[0164] where is the total number of classification categories, is the proportion of samples of category in the node, is the Gini coefficient. The smaller the value, the more concentrated the samples are, the larger the proportion of samples belonging to the same category, that is, the higher the node purity;
[0165] To measure the degree of improvement in purity by the splitting operation, the Gini gain index is introduced. The Gini gain reflects the degree of improvement in classification purity after splitting the current node. The larger the Gini gain value, the more significant the improvement in the purity of the child nodes after splitting and the better the splitting effect. During splitting, calculate the Gini gain after splitting at each candidate splitting point in each feature set. The calculation formula of the Gini gain is shown in (8):
[0166] (8)
[0167] where is the Gini coefficient of the parent node before splitting, representing the sample purity of the parent node, , are the Gini coefficients of the left and right child nodes after splitting, representing the sample purities of the left and right child nodes respectively, , are the sample numbers of the left and right child nodes, is the total sample number of the parent node;
[0168] To ensure that the currently selected splitting point is statistically significant and avoid making wrong decisions due to insufficient samples. The Hoeffding bound provides a way based on the sample size and confidence level to judge whether the difference in Gini gain is large enough to ensure the superiority of the current splitting point;
[0169] First, calculate the difference in Gini gain, as shown in formula (9):
[0170] (9)
[0171] where is the Gini gain of the optimal splitting point, is the Gini gain of the sub - optimal splitting point, is the difference in Gini gain between the two splitting points;
[0172] The Hoeffding bound formula can be expressed as (10):
[0173] (10)
[0174] Where is the number of samples, is the detection threshold, is the variable the length of the value range;
[0175] If the Gini gain difference , it is considered that the current split point is significantly better than the sub-optimal split point, and the candidate split point is accepted for tree splitting; if the Gini gain difference , it means that the current sample is not sufficient to prove the significance of the split point, and it needs to be re-evaluated after data accumulation;
[0176] Recursively split the node, and for the newly generated left and right child nodes, repeat the above process of traversing the indicators, calculating the candidate split points, Gini gain evaluation, and Hoeffding bound judgment;
[0177] If the child node satisfies the Gini gain difference or the sample size of the child node is less than the preset minimum split threshold, stop splitting the child node and mark it as a leaf node;
[0178] If the child node can continue to split, recursively execute the above splitting operation until all nodes reach the stop splitting condition;
[0179] When all nodes of the entire tree reach a state where they cannot be further split, the tree construction process is completed;
[0180] As new data continuously flows in, repeat the above steps to achieve learning of incremental data;
[0181] Integrate the Hoeffding tree joint decision-making process. Based on the pre-established integrated Hoeffding tree model, when the data stream is input into the model for prediction, synthesize the prediction results of the main tree T1 and the sub-trees T2~T k to generate the final classification result through a dynamic weighted voting mechanism;
[0182] Decision window initialization: Use a sliding window Q to quantify the historical decision reliability of the sub-trees T2~T k , set the sliding window capacity to , fill it with all 1 values in the initial state, and assume that the initial decisions of the sub-trees T2~T k are correct;
[0183] Window update rule: Whenever a new sample is input, the sub-trees T2~Tk Perform weighted voting prediction on the sample. If the prediction is correct, fill 1 in the latest position of the sliding window; otherwise, fill 0. Among them, subtrees T2 to T k Correct predictions are marked as 1, and incorrect predictions are marked as 0;
[0184] Confidence calculation: When the data stream arrives, count the sum of correct decisions of subtrees T2 to T within the window, which can be expressed by formula (11): k Make the sum of correct decisions, which can be expressed as formula (11):
[0185] (11)
[0186] Where Represents the th decision value stored in the sliding window Q;
[0187] Combined decision: Compare with the preset threshold In this example, Is set to 0.9, and the sliding window capacity Is set to 40;
[0188] If Determine that the decisions of subtrees T2 to T k Are credible, and adopt the majority voting strategy to fuse the prediction results of the main tree T1 and subtrees T2 to T k ;
[0189] If Only use the output of the main tree T1 as the final classification result.
[0190] As Figure 3 Shown, a dynamic decision-making method based on an ensemble Hoeffding tree includes the following steps:
[0191] Initialize the entire tree. The entire tree T1 is trained using all the features of the dataset to ensure global feature coverage to guarantee that key information is not lost;
[0192] Initialize the subtrees. Subtrees T2, …, subtree T k Are trained using partial features of the dataset to effectively capture data diversity features and enhance adaptability to local concept drift;
[0193] Build an ensemble model by combining the initialized entire tree T1 and subtrees T2, …, subtree T k Together to form an ensemble model, which combines the different characteristics and advantages of the entire tree and subtrees.
[0194] Combined decision. As new data is continuously input into the ensemble Hoeffding tree model, during the prediction process, the entire tree T1 and subtrees T2, …, subtree Tk In the process of jointly participating in the decision-making, based on their own model structures and training results, they each make predictions on the input data, and finally output the results through joint voting decisions, reducing the impact of individual models on abnormal data or local drifts and improving the overall classification accuracy.
[0195] Specifically, the following further elaborates on the solution of the present invention through embodiments:
[0196] Through experimental comparison and verification, the integrated learning classification algorithm with concept drift detection function proposed by the present invention has better performance. In this paper, the traffic data in the dataset is classified and detected, and the goal is to distinguish "normal traffic" from "abnormal traffic" to ensure the correct classification of traffic data. For the classification task, the following four evaluation metrics based on the confusion matrix are selected in this paper: accuracy, precision, recall, and .
[0197] Table 1 shows the confusion matrix, where and respectively represent the total number of samples correctly predicted as normal traffic and abnormal traffic. and respectively represent the total number of samples mispredicted in normal traffic and abnormal traffic. Accuracy represents the ratio of the number of correctly predicted samples to the total number of samples, precision represents the ratio of the samples correctly predicted as normal traffic to the samples predicted as normal traffic, recall represents the ratio of the samples correctly predicted as normal traffic to the samples actually being normal traffic, and combines the characteristics of precision and recall, which is their harmonic mean, and the calculation formula is as follows:
[0198] (22)
[0199] (23)
[0200] (24)
[0201] (25)
[0202] Table 1 Confusion Matrix
[0203]
[0204] Table 2 shows the comparison of the proposed method KS-DB-IOS with ELM, DDM, ADWIN, OCDD and these four methods.
[0205] Table 2 Performance Comparison of Five Algorithms
[0206]
[0207] Example 2: Second, as Figure 4 shown, to achieve the above object, the present invention discloses an ensemble learning classification system with concept drift detection function, including:
[0208] A data dimensionality reduction module 11, configured to obtain high-dimensional data of an extensible sliding window, perform dimensionality reduction on the high-dimensional data of the extensible sliding window by using principal component analysis PCA, and extract key feature data;
[0209] A data expansion module 12, configured to perform real-time monitoring on the key feature data based on an adaptive window Kolmogorov-Smirnov early warning detection method, determine whether concept drift occurs, if concept drift occurs, obtain the original data of the adaptive window when concept drift occurs, and automatically expand the adaptive window, if concept drift does not occur, the adaptive window slides normally;
[0210] A data classification module 13, configured to input the original data of the adaptive window when concept drift occurs into a pre-established learning model based on an ensemble Hoeffding tree, obtain a trained learning model based on the ensemble Hoeffding tree, and perform real-time classification of the data stream based on the trained learning model based on the ensemble Hoeffding tree.
[0211] Based on the same inventive concept, the present invention also provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is configured to execute the program instructions stored in the memory. The processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is used to implement one or more instructions, specifically for loading and executing one or more instructions in the computer storage medium to implement the above method.
[0212] It should be further noted that, based on the same inventive concept, the present invention also provides a computer storage medium, on which a computer program is stored, and when the computer program is run by a processor, it executes the above method. The storage medium may adopt any combination of one or more computer-readable media. The computer-readable media may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electro-magnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or combined with an instruction execution system, apparatus, or device.
[0213] In the description of this specification, the description with reference to the terms "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0214] The above shows and describes the basic principles, main features, and advantages of the present disclosure. Those skilled in the art of this industry should understand that the present disclosure is not limited by the above embodiments. The above embodiments and the descriptions in the specification only illustrate the principles of the present disclosure. Without departing from the spirit and scope of the present disclosure, the present disclosure will have various changes and improvements, and these changes and improvements all fall within the scope of the present disclosure claimed.
Claims
1. An ensemble learning classification method with concept drift detection function, characterized in that: The method comprises the following steps: Obtain high-dimensional data of an expandable sliding window, use principal component analysis (PCA) to reduce the dimension of the high-dimensional data of the expandable sliding window, and extract key feature data; Based on the adaptive window Kolmogorov-Smirnov early warning detection method, the key feature data is monitored in real time to determine whether concept drift occurs. If concept drift occurs, the original data of the adaptive window when the concept drift occurs is obtained, and the adaptive window is automatically expanded. If concept drift does not occur, the adaptive window slides normally. The original data of the adaptive window is network traffic data, and the network traffic data is divided into normal traffic and abnormal traffic; The process of real-time concept drift detection for key feature data based on the adaptive window Kolmogorov-Smirnov early warning detection method is as follows: For the k dimensions of the extracted key feature data, calculate the statistics D of the difference in cumulative distribution function respectively i and the corresponding p i value, and then according to the variance contribution rate V of each principal component i The global drift significance score p is obtained by weighting, the significance level α is set as the concept drift judgment threshold, and the warning threshold β is defined; If p≤α, it is considered that data concept drift has occurred, the system will clear the historical window data, and use the adaptive window data to retrain the integrated Hoeffding tree model; If α<p≤β, it is considered that a drift warning has occurred, and the system automatically enters the elastic buffer stage, dynamically expanding the adaptive window to accumulate potential drift samples; If p>β, it is considered that no drift has occurred and the sliding window continues to update the data; Inputting the original data of the adaptive window when the concept drift occurs into the pre-established learning model based on the integrated Hoeffding tree to obtain the trained learning model based on the integrated Hoeffding tree, and performing real-time classification of the data stream based on the trained learning model based on the integrated Hoeffding tree; The pre-established learning model based on the integrated Hoeffding tree is integrated through a main tree and multiple subtree architecture, where the main tree is trained with a complete indicator set to ensure global feature coverage, and the subtree generates a differentiated classifier by randomly selecting some indicator dimensions. Both the main tree and the multiple subtrees dynamically judge node splitting based on the Gini gain difference and the Hoeffding boundary, and update the tree structure to capture feature distribution changes; The joint decision-making process of the integrated Hoeffding tree in the learning model based on the integrated Hoeffding tree is as follows: Based on the pre-established integrated Hoeffding tree model, when the data stream is input into the model for prediction, the main tree T1 and subtrees T2~T k The prediction results are used to generate the final classification results through a dynamic weighted voting mechanism; Decision window initialization: Use sliding window Q to quantize subtrees T2~T k The historical decision reliability is set to d. impact ×q, the initial state is filled with all 1 values, and subtrees T2~T k The initial decision was correct; Window update rule: Whenever a new sample is input, subtrees T2~T k Perform weighted voting prediction on the sample. If the prediction is correct, fill 1 into the latest bit of the sliding window, otherwise fill 0. k Correct predictions are marked as 1, and incorrect predictions are marked as 0; Credibility calculation: When the data stream arrives, the subtrees T2 to T2 in the statistics window are counted. k The sum of correct decisions Sum(Q) can be expressed as formula (11): Where Q[i] represents the i-th decision value stored in the sliding window Q; Joint decision: Sum(Q) and the preset threshold d impact ×q comparison; If Sum(Q)≥d impact ×q, determine subtree T2~T k The decision is credible, and the majority voting strategy is used to merge the main tree T1 and subtrees T2~T k The prediction results; IfSum(Q) <d impact ×q, only the output of the main tree T1 is used as the final classification result.
2. The ensemble learning classification method with concept drift detection function according to claim 1, characterized in that: Initialization process of the integrated Hoeffding tree: Build a decision tree consisting of k empty trees T1, T2, ..., T k The model consists of T1 and T2, where T1 is the whole tree trained with all indicators, T2~T k The subtree is trained with some indicators, and the complete tree T1 uses the complete indicator set f total =[x1, x2, ..., x 20 ] for training; T2~T k Subtree random selection To ensure that each indicator is within the range of T2 to T k can be selected at least once, set the target indicator coverage probability P observed , the minimum number of subtrees m required can be derived through the probability formula of indicator coverage. The steps are as follows: The indicator coverage of a single subtree, the probability P1 that an indicator is not selected in a single subtree, is shown in formula (1): Coverage of multiple subtrees. Suppose there are m subtrees. The probability P2 that a certain indicator is not selected in the m subtrees is as shown in formula (2): The probability that a certain indicator is selected at least once in m subtrees is expressed as formula (3). Substituting it into formula (2), we get formula (4) as follows: P observed =1-P2 (3) In the target indicator coverage probability P observed When the value is the preset value, calculate the value that satisfies a certain indicator between T2 and T k The condition that a subtree is selected at least once in the tree is met, while minimizing the number of subtrees m required for computing resources.
3. The ensemble learning classification method with concept drift detection function according to claim 2, characterized in that: The pre-established learning model based on the integrated Hoeffding tree, as the data stream samples continue to increase, the leaf nodes of the integrated Hoeffding tree need to determine the candidate area division method, the steps are as follows: Traversing indicator variables: Let the current node sample be S, and the indicator variable set be X = [x1, x2, ..., x n ], traverse each indicator variable x in turn j , where j = 1, 2, ..., n; Generate candidate split points: For each indicator variable x j , the indicator variable x j All the index values contained are sorted from small to large, denoted as {v1, v2, ..., v d }, the candidate split point set is the median of all adjacent index values, and the candidate split point is represented by ∈ i , as shown in formula (5): where v i Represents the indicator variable x j The i-th index value in the sorted set, d represents the index variable x j The number of all indicator values, i is the index of the candidate split point, ranging from 1 to d-1; Calculate the partitioning effect corresponding to the split point: for each candidate split point ∈ i , divide the current node sample S into two sub-areas, as shown in formula (6): Where S1(i, j) and S2(i, j) are the left and right subsets after division, s is the data instance in the sample set S, and x j (s) is the sample instance s in the indicator variable x j The value on .
4. The ensemble learning classification method with concept drift detection function according to claim 3, characterized in that: The optimal selection process of the candidate split points: The purity of a node is quantified by the Gini coefficient. For a node, its Gini coefficient is defined as shown in the formula: Where K is the total number of categories, P k is the sample proportion of category k in the node, and G is the Gini coefficient; Traverse each indicator variable x j All candidate split points are calculated, and the Gini gain after splitting at each split point is calculated. The Gini gain G gain The calculation formula is shown in (8): Among them G parent is the Gini coefficient of the parent node before splitting, indicating the sample purity of the parent node, G left , G right is the Gini coefficient of the left and right child nodes after splitting, representing the sample purity of the left and right child nodes respectively, N left 、N right is the number of samples of left and right child nodes, N total is the total number of samples of the parent node; Select the candidate split point with the largest Gini gain value as the current optimal split point, record the Gini gain value of the sub - optimal split point, and calculate the difference in Gini gain between the two to verify the significance of the split, as shown in formula (9): ΔG=G best -G second (9) Among them G best is the Gini gain of the optimal split point, G second is the Gini gain of the suboptimal split point, ΔG is the difference in Gini gain between the two split points; To ensure the statistical significance of the Gini gain difference, the Hoeffding bound is introduced as a confidence constraint condition, which can be expressed as (10): Where n is the number of samples, δ is the detection threshold, and R is the variable x j The length of the value range of ; If the Gini gain difference ΔG ≥ h, it is considered that the optimal split point is significantly better than the sub - optimal split point, and the optimal split point is accepted for tree splitting; if the Gini gain difference ΔG < h, it means that the current sample is not sufficient to prove the significance of the split point, and re - evaluation is required after data accumulation; For the newly generated child nodes, the process of selecting the optimal split point is repeated until all nodes meet the conditions for stopping splitting.
5. The ensemble learning classification method with concept drift detection function according to claim 1, characterized in that: The process of dimensionality reduction of high - dimensional data of an expandable sliding window using principal component analysis (PCA) is as follows: The proposed principal component analysis (PCA) data dimensionality reduction operation requires the creation of two sliding windows: a fixed-capacity historical window L and a dynamically expandable adaptive window R, where the initial capacity of L is C L , the initial capacity of R is C R And set the maximum expansion limit R MAX , perform dimensionality reduction operations on the historical window L and the adaptive window R respectively; Standardize the data. First, standardize the window data as shown in formula (12): where X is the original data, μ is the mean of the index, σ is the standard deviation of the index, and X′ is the standardized data; Calculate the covariance matrix. Calculate the covariance matrix γ of the standardized data X′. The covariance matrix represents the correlation between each index in the data, as shown in formula (13): Where n is the number of samples in the window, X′ T is the transposed matrix of X′; By solving the eigenvalues and corresponding eigenvectors of the covariance matrix, the direction of the principal component of the data is obtained. The eigenvalue represents the variance of the principal component, and the eigenvector represents the direction of the principal component. The eigenvalues are arranged in descending order, and the eigenvectors corresponding to the first k largest eigenvalues are selected. The principal component with a large eigenvalue represents a large proportion of the original information of the data. The variance contribution rate V of each principal component is calculated. i , as shown in formula (14): Among them, λ i is the eigenvalue of the i-th principal component after sorting, and k is the number of principal components selected; Project the data. Arrange the selected k eigenvectors in columns to form a transformation matrix W, and project the standardized data X′ onto the selected principal components, as shown in formula (15): X reduced =X′W (15) Among them, X reduced It is the data after dimensionality reduction.
6. The ensemble learning classification method with concept drift detection function according to claim 5, characterized in that: The usage process of the adaptive window Kolmogorov - Smirnov early warning detection method is as follows: For the reduced-dimensional data X reduced , extract the i-th dimension data of the historical window and the adaptive window respectively for Kolmogorov-Smirnov test, as shown in formula (16): Among them, X reduced,his is the data matrix after dimension reduction in the historical window, X reduced,cur is the data matrix after adaptive window dimension reduction, X his,i is the data of the historical window on the i-th principal component, X cur,i is the data of the adaptive window on the i-th principal component; The Kolmogorov - Smirnov test process is as follows: Establish the null hypothesis H0: Assume that there is no difference between the two samples and they come from the same distribution; Calculate the cumulative distribution function: For each dimension i, calculate X separately his,i , X cur,i The cumulative distribution function F his,i (x), F cur,i (x), the cumulative distribution function describes the probability that a random variable is less than or equal to a preset specific value, that is, F(x) = P(X≤x). Suppose that n sample data sets X n The samples in are independent and random, then the cumulative distribution function F(x) is expressed as formula (17): Among them, I(x i ≤x) is the indicator function, as shown in formula (18): Calculate the Kolmogorov-Smirnov statistic D i , denoted as F his,i (x) and F cur,i (x) The maximum vertical distance between the cumulative distribution functions is the maximum difference between the two distributions, as shown in formula (19): Define the significance level and calculate p i value, the maximum difference D between the data distributions i Convert to p i Value, formula (20) is the calculation expression: Where n is the effective sample size, that is n1 is X his,i The number of samples in X, n2 is cur,i The number of samples in ; Weighted p-value: According to the variance contribution rate V of each principal component i , for p i The values are weighted to obtain the global drift significance score p, as shown in formula (21):
7. An integrated learning classification system with concept drift detection function, adopting an integrated learning classification method with concept drift detection function as claimed in any one of claims 1 to 6, characterized in that: including: A data dimensionality reduction module, which is used to obtain the high - dimensional data of an expandable sliding window, perform dimensionality reduction on the high - dimensional data of the expandable sliding window using principal component analysis (PCA), and extract key feature data; A data expansion module, which is used to perform real - time monitoring on the key feature data based on the adaptive window Kolmogorov - Smirnov early warning detection method to determine whether concept drift occurs. If concept drift occurs, obtain the original data of the adaptive window when concept drift occurs and automatically expand the adaptive window. If concept drift does not occur, the adaptive window slides normally; The process of real - time concept drift detection for key feature data based on the adaptive window Kolmogorov - Smirnov early warning detection method is as follows: For the k dimensions of the extracted key feature data, calculate the statistics D of the difference in cumulative distribution function respectively i and the corresponding p i value, and then according to the variance contribution rate V of each principal component i The global drift significance score p is obtained by weighting, the significance level α is set as the concept drift judgment threshold, and the warning threshold β is defined; If p ≤ α, it is considered that data concept drift has occurred, and the system will clear the historical window data and retrain the ensemble Hoeffding tree model using the adaptive window data; If α < p ≤ β, it is considered that drift warning has occurred, and the system automatically enters the elastic buffer stage to dynamically expand the adaptive window to accumulate potential drift samples; If p > β, it is considered that no drift has occurred, and the window continues to slide to update the data; A data classification module is used to input the original data of the adaptive window when concept drift occurs into a pre-established learning model based on the integrated Hoeffding tree to obtain a trained learning model based on the integrated Hoeffding tree, and perform real-time classification of data streams based on the trained learning model based on the integrated Hoeffding tree; The pre-established learning model based on the integrated Hoeffding tree is integrated through a main tree and multiple subtree architecture, where the main tree is trained with a complete indicator set to ensure global feature coverage, and the subtree generates a differentiated classifier by randomly selecting some indicator dimensions. Both the main tree and the multiple subtrees dynamically judge node splitting based on the Gini gain difference and the Hoeffding boundary, and update the tree structure to capture feature distribution changes; The joint decision-making process of the integrated Hoeffding tree in the learning model based on the integrated Hoeffding tree is as follows: Based on the pre-established integrated Hoeffding tree model, when the data stream is input into the model for prediction, the main tree T1 and subtrees T2~T k The prediction results are used to generate the final classification results through a dynamic weighted voting mechanism; Decision window initialization: Use sliding window Q to quantize subtrees T2~T k The historical decision reliability is set to d. impact ×q, the initial state is filled with all 1 values, and subtrees T2~T k The initial decision was correct; Window update rule: Whenever a new sample is input, subtrees T2~T k Perform weighted voting prediction on the sample. If the prediction is correct, fill 1 into the latest bit of the sliding window, otherwise fill 0. k Correct predictions are marked as 1, and incorrect predictions are marked as 0; Credibility calculation: When the data stream arrives, the subtrees T2 to T2 in the statistics window are counted. k The sum of correct decisions Sum(Q) can be expressed as formula (11): Where Q[i] represents the i-th decision value stored in the sliding window Q; Joint decision: Sum(Q) and the preset threshold d impact ×q comparison; If Sum(Q)≥d impact ×q, determine subtree T2~T k The decision is credible, and the majority voting strategy is used to merge the main tree T1 and subtrees T2~T k The prediction results; If Sum(Q)<d impact ×q, only the output of the main tree T1 is used as the final classification result.
8. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: The memory stores a computer program that can be run on the processor. When the processor loads and executes the computer program, an integrated learning classification method with a concept drift detection function as described in any one of claims 1 to 6 is adopted.
Citation Information
Patent Citations
Conceptual drift detection and adaptation method based on sub-feature selection
CN118260687A
KR20230099132A