Network data flow feature extraction method based on task guidance and mixed learning
By adopting a feature extraction method based on task-oriented and mixed learning in network data flow analysis, combined with the chi-square test and self-supervised learning of decision tree model, the problem of difficulty in identifying network attack behavior and predicting network load in the existing technology is solved, and efficient network data flow feature extraction and analysis is achieved, improving network security and operational efficiency.
Patent Information
- Application Number
- CN202411898771.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to effectively identify and analyze complex network data flows, especially when facing massive network traffic, it is difficult to accurately identify network attack behaviors and predict network loads, resulting in insecure network security and operational efficiency.
A network data flow feature extraction method based on task-oriented and mixed learning is adopted, and a self-supervised learning method of chi-square test and decision tree model is combined to extract and analyze the network data flow. This method can adapt to different network environments and task requirements through a task goal-oriented feature extraction process, including data preprocessing, feature screening and model construction.
It realizes efficient feature extraction of network data flow, can accurately identify network attack behavior and analyze network traffic patterns, improve network security and operational efficiency, and adapt to complex and changeable network environments.
Smart Images

Figure CN119945720A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network security, and in particular relates to a network data flow feature extraction method based on task orientation and hybrid learning. Background Art
[0002] With the rapid development of network technology, network data flows have shown explosive growth, and their role in various fields of modern society has become increasingly prominent, but at the same time it has also brought a series of severe challenges. The high-dimensional, sparse and heterogeneous characteristics of network data flows make it extremely difficult to directly analyze them effectively. For example, in the field of network security, facing massive and complex network traffic data, traditional analysis methods are difficult to accurately identify various network attack behaviors, such as increasingly rampant DDoS attacks, highly concealed malware traffic, and frequent port scans. These attacks are often hidden in a large amount of normal traffic. If they cannot be detected and prevented in a timely and accurate manner, they will lead to serious consequences such as network system paralysis, user information leakage, and huge economic losses to enterprises.
[0003] In terms of network operation and management, it is difficult to achieve reasonable planning and allocation of network resources due to the lack of in-depth understanding of network traffic patterns. Unreasonable resource allocation may cause network congestion, resulting in a sharp decline in network service quality, affecting user experience and thus reducing corporate competitiveness. For example, during peak hours, if the network load cannot be accurately predicted and resources cannot be prepared in advance, network jams or even service interruptions are likely to occur.
[0004] Existing feature extraction methods often have limitations and cannot adapt well to the complexity and diversity of network data flows. In the field of network data analysis, traditional feature extraction methods focus on simple statistical analysis of data or rule-based feature extraction. For example, early network traffic analysis focused on the basic information of data packets, such as statistical counts of source IP, destination IP, port number, etc., and then judged whether the traffic was abnormal based on some pre-set rules. However, with the development of network technology, the complexity of network data flows has greatly increased, and this simple method has become difficult to cope with.
[0005] The rise of machine learning technology has brought new ideas to network data stream feature extraction. Supervised learning methods can learn the feature representation of data when there is labeled data, but their application is limited for large-scale unlabeled network data streams. Although unsupervised learning methods such as principal component analysis (PCA) and independent component analysis (ICA) can extract features to a certain extent, they are not ideal when dealing with specific tasks of network data streams. For example, PCA mainly focuses on the variance of the data and may ignore some important classification information in the data, resulting in low accuracy in tasks such as network attack behavior identification.
[0006] In recent years, some self-supervised learning methods have gradually attracted attention. As a commonly used statistical test method, the chi-square test can preliminarily screen out features that have a strong correlation with the target variable, but it lacks the ability to build complex models when used alone. The decision tree model has good interpretability and classification capabilities, but it is prone to overfitting and high computational complexity in the face of high-dimensional data. Therefore, the existing technology urgently needs a new technical solution to solve the above problems. Summary of the invention
[0007] The technical problem to be solved by the present invention is to provide a network data flow feature extraction method based on task orientation and hybrid learning, which can efficiently extract valuable features from the network data flow according to different task objectives (such as network attack identification, traffic pattern analysis, network load prediction, etc.), and adapt to the ever-changing network environment through continuous optimization and improvement, so as to ensure network security, improve network operation efficiency, and meet the urgent needs of modern network management.
[0008] A network data stream feature extraction method based on task orientation and hybrid learning includes the following steps:
[0009] Step 1: Construct network data flow feature extraction task objectives for the collected network data flow, including identifying network attack behaviors, analyzing network traffic patterns, and predicting network loads;
[0010] Step 2: Preprocess the original data of the network data stream, perform data cleaning and data normalization;
[0011] Step 3: According to the task objectives established in step 1, supervised learning methods are used to extract features from the preprocessed data;
[0012] Step 4: Apply the feature extraction of step 3 to the actual network traffic analysis. Based on the task objectives constructed in step 1, compare the abnormal traffic characteristics with the normal traffic characteristics, obtain the network attack feature pattern, and perform network protection.
[0013] The data cleaning described in step 2 includes removing duplicate data and removing invalid data.
[0014] The data normalization described in step 2 uses Z-score normalization to normalize each eigenvalue to a mean of 0 and a standard deviation of 1.
[0015] The supervised learning method described in step 3 adopts the chi-square-decision tree model, including the following steps:
[0016] 1. Chi-square feature screening
[0017] Assume that the data set is D = (X, y), where X is the feature matrix and y is the target variable. In the classification problem, the target variable y has C categories.
[0018] For each feature xi∈X (i=1,2,…,n, n is the number of features), a contingency table is constructed; the rows of the contingency table represent the values of the feature xi, and the columns represent the categories of the target variable y;
[0019] Calculate the chi-square statistic The calculation formula is Where R is the number of rows in the contingency table, which is the number of values of feature xi; Orc is the observed frequency, which is the number of samples in the actual data where feature xi takes the rth value and target variable y is the cth category; Erc is the expected frequency, which is the frequency calculated under the assumption that feature xi and y are unrelated;
[0020] According to the pre-set chi-square threshold choose The features of subset ;
[0021] 2. Decision Tree Construction
[0022] In the feature subset χ subset and target variable y to build a decision tree; using the CART decision tree, at each node, calculate each feature χ j ∈χ subset Gini impurity (j = 1, 2, ..., m, m is the number of subset features); Gini impurity Where D is the data subset of the current node, K is the number of categories of the target variable y in D, and p k is the probability of category k in D;
[0023] Select the feature with the largest reduction in Gini impurity as the split feature of the current node; calculate the reduction in Gini impurity by splitting the current node according to the feature: where V is the characteristic χ j The number of values, D v According to the characteristic χ j The sub-node data subset after the v-th value division; select ΔGini j The largest feature is split;
[0024] Repeat the above splitting process until the stopping condition is met; the stopping condition is that the depth of the tree reaches a predetermined value, the number of samples in the leaf node is less than a threshold, or the reduction in Gini impurity is less than a threshold;
[0025] 3. Model Evaluation and Optimization
[0026] Cross-validation was used to evaluate the model performance. A 10-fold cross-validation was used to divide the data set into 10 subsets. In each fold, 9 subsets were used as training sets and 1 subset was used as validation set. A combined model including chi-square test feature screening and decision tree construction was constructed on the training set. The evaluation index of the model was calculated on the validation set.
[0027] Including accuracy
[0028] Among them, TP is the number of true positives, TN is the number of true negatives, FP is the number of false positives, and FN is the number of false negatives;
[0029] Recall
[0030] in,
[0031] Adjusting the Chi-Square Threshold And the decision tree parameters, including the tree depth depth, the minimum number of leaf node samples min_sammples_leaf, repeat the cross-validation process to obtain the parameter combination that makes the model evaluation index optimal.
[0032] Through the above-mentioned design scheme, the present invention can bring the following beneficial effects: a network data flow feature extraction method based on task orientation and hybrid learning, a self-supervised learning method combining the chi-square test with the decision tree model, which is expected to play a unique advantage in network data flow feature extraction, overcome the shortcomings of traditional methods, and provide a more accurate and efficient solution for network traffic analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The present invention is further described below with reference to the accompanying drawings and specific embodiments:
[0034] Figure 1 This is a flow chart of a network data stream feature extraction method based on task orientation and hybrid learning in the present invention. DETAILED DESCRIPTION
[0035] The present invention provides a network data stream feature extraction method based on task orientation and hybrid learning, such as Figure 1 As shown in the figure, firstly, the specific task objectives are clarified for the collected network data flow data, including identifying network attack behaviors, analyzing network traffic patterns, etc. Then, the network data flow data is preprocessed by data cleaning (removing redundant and invalid data). Then, a chi-square-decision tree model is built to extract features, and the extracted features are evaluated and optimized. Finally, the tasks related to traffic analysis are carried out.
[0036] The detailed steps and technical details are as follows:
[0037] 1. First, for the collected network data streams, it is necessary to clarify the specific task objectives of network data stream feature extraction. These objectives may include identifying network attack behaviors, analyzing network traffic patterns, predicting network loads, etc. Clarifying the task objectives helps determine the type and number of features that need to be extracted, as well as the subsequent processing and analysis methods.
[0038] 2. Network data streams are usually high-dimensional, sparse, and heterogeneous. Direct analysis may face problems of high computational complexity and low efficiency. Therefore, before feature extraction, the raw data needs to be preprocessed. The preprocessing steps include data cleaning (removing redundant and invalid data), data normalization (making different features have the same dimension and distribution), etc.
[0039] 2.1 Data cleaning includes two steps: removing duplicate data and removing invalid data
[0040] 2.1.1 Remove duplicate data
[0041] step:
[0042] ① For each data record in the network data stream, use a hash function or comparison algorithm to identify duplicate data. For example, in a data stream containing network connection records (such as source IP, destination IP, port number, protocol, etc.), calculate a hash value for each record. If the hash values of two records are the same, then compare their fields in detail to see if they are completely consistent.
[0043] ② When duplicate data is determined, you can choose to keep one of the records and delete the other duplicate records. Efficiency can be improved by maintaining a collection of processed records (such as using a hash table to store the hash values of records that have been checked).
[0044] Duplicate data determination formula:
[0045] Let the data record be x i =(x i1 ,x i2 ,…,x in ), where i = 1, 2, …, m (m is the total number of data records), and n is the dimension of each record.
[0046] Calculate the hash function h(x i ), if h(x i )=h(x j )(i≠j), then we need to compare x ik =x jk (k=1,2,…,n), if all are equal, then x j This is duplicate data and can be deleted.
[0047] 2.1.2 Remove invalid data
[0048] step:
[0049] ① Determine the rules for invalid data. For example, in the data flow of a network data packet, if the length field of a data packet does not conform to the length range specified by the protocol (such as the length of an IP data packet is less than 20 bytes, which does not conform to the IP protocol specification), it is judged as invalid data.
[0050] Formula for determining invalid data: Assume that the data length range specified by the protocol is [L min ,L max ], for data record x i , if its data length L(x i ) <L min Or L(x i )>L max , then x i For invalid data, delete x i For example, for a TCP data packet, the minimum header length is 20 bytes. Suppose data record x i Represents a TCP data packet, whose header length field is L(x i ), if L(x i )<20, it is invalid data.
[0051] ② Check each data record according to the rules and mark the data records that do not comply with the rules as invalid data.
[0052] ③ Delete invalid data from the data stream. You can quickly locate and delete invalid data through indexes or pointers, or copy valid data to a new data set and discard the original location of the invalid data.
[0053] 2.2 Data normalization is carried out in the following steps
[0054] step:
[0055] Similarly, first traverse the entire original data stream, and for each original data feature j (j = 1, 2, ..., n), calculate the mean of the feature in the entire data stream and standard deviation σ j The mean is calculated by adding the eigenvalues of all data points in the data stream and dividing by the total number of data points. The standard deviation is calculated by first calculating the sum of the squares of the difference between the eigenvalue of each data point and the mean, dividing by the total number of data points, and finally taking the square root.
[0056] Next, for each data point x in the original data stream i =(x i1 ,x i2 ,…,x in), for each feature j, normalization is performed according to the following formula:
[0057] Let the data point in the original data stream be x i =(x i1 ,x i2 ,…,x in ), for each feature j, the mean of the feature in the entire data stream is The standard deviation is σ j After Z-score normalization, the data point x i The normalized value of feature j in for:
[0058] After Z-score normalization, each feature is the normalized value of j. It has the characteristics of a mean of 0 and a standard deviation of 1, which realizes the normalization of the distribution of each feature in the original data stream.
[0059] 3. Task-driven feature extraction method
[0060] 3.1 If there is labeled network traffic data, supervised learning methods can be used to extract features. By training classifiers, feature representations that distinguish different types of traffic can be learned. These methods can automatically select important features and perform subsequent related tasks based on these features.
[0061] 3. Self-supervised learning task design
[0062] Self-supervised model selection Chi-square-decision tree model. The chi-square test can screen out features that are strongly associated with the target variable in the early stage. In high-dimensional data sets, decision trees may become complicated and prone to overfitting due to too many features. By removing those features that are not statistically related to the target variable through the chi-square test, the decision tree construction process can be simplified, so that the decision tree can focus on more valuable features, reduce the amount of calculation and the risk of overfitting. The chi-square test is a method based on statistical tests, and has certain considerations on the distribution of data and other characteristics. The decision tree is easily affected by small changes in the data during the construction process. By first selecting features through the chi-square test, the interference of data noise on the decision tree can be reduced to a certain extent, making the model more stable. This combined model is mainly divided into two stages. The first stage is the feature screening stage, in which the original feature set is selected using the chi-square test. The second stage is the model construction stage, in which the decision tree model is constructed based on the selected features. From the overall structure, it looks like a decision tree model with a preprocessing step (chi-square test screening).
[0063] Chi-square-decision tree model establishment method:
[0064] Feature selection stage: Chi-square test can be used to preliminarily select the original features. Chi-square test can quickly find features that are statistically significantly associated with the target variable. For those features whose chi-square statistic is lower than a certain threshold, they are temporarily excluded from the subsequent decision tree construction. This can reduce the number of features in the decision tree construction process, reduce the computational complexity, and exclude those features that are weakly associated with the target variable to a certain extent.
[0065] Model building phase: Build a decision tree on the feature subset that has been screened by the chi-square test. During the decision tree construction process, the classification contribution of each feature to the target variable is further evaluated. For example, indicators such as information gain, information gain ratio, or Gini impurity are used to select the best split feature.
[0066] Model evaluation and optimization stage: You can use methods such as cross-validation to evaluate the performance of this combination. For example, using k-fold cross-validation, divide the data set into k subsets, use k-1 subsets as training sets each time, and 1 subset as validation set. During the training process, continuously adjust the threshold of the chi-square test or the parameters of the decision tree (such as the depth of the tree, the minimum number of samples of leaf nodes, etc.) to optimize the model performance. Observe the accuracy, recall rate, F1-score and other indicators of the model on the validation set to find the optimal model configuration.
[0067] Specific implementation steps:
[0068] ① Chi-square feature screening
[0069] Assume that the data set is D = (X, y), where X is the feature matrix and y is the target variable. For classification problems, assume that the target variable y has C categories.
[0070] For each feature xi∈X (i=1,2,…,n, n is the number of features), a contingency table is constructed. The rows of the contingency table represent different values of the feature xi, and the columns represent different categories of the target variable y.
[0071] Calculate the chi-square statistic The calculation formula is Where R is the number of rows in the contingency table (the number of values of feature xi), Orc is the observed frequency (the number of samples in the actual data where feature xi takes the rth value and target variable y is the cth category), and Erc is the expected frequency (the frequency calculated under the assumption that feature xi is unrelated to y).
[0072] According to the pre-set chi-square threshold (The threshold value range is set to 10-15), select
[0073] The features of subset .
[0074] ②Decision tree construction
[0075] In the feature subset χ subset And the target variable y is used to build a decision tree. Using the CART decision tree, at each node, calculate each feature χ j ∈χ subset Gini impurity (j=1,2,…,m, m is the number of subset features). Gini impurity Where D is the data subset of the current node, K is the number of categories of the target variable y in D, and p k is the probability of class k in D.
[0076] Select the feature that reduces the Gini impurity the most as the split feature of the current node. For example, calculate the reduction in Gini impurity by splitting the current node according to the feature:
[0077] where V is the characteristic χ j The number of values, D v According to the characteristic χ j The sub-node data subset after the v-th value division. Select ΔGini j The largest feature is split.
[0078] The above splitting process is repeated until the stopping condition is met. The stopping condition may be that the depth of the tree reaches a predetermined value, the number of samples in the leaf node is less than a threshold, or the reduction in Gini impurity is less than a threshold.
[0079] ③Model evaluation and optimization
[0080] Use cross-validation to evaluate model performance. For example, use 10-fold cross-validation to divide the dataset into 10 subsets. For each fold, use 9 subsets as training sets and 1 subset as validation set. Build a combined model on the training set according to the previous steps (chi-square test to select features + build a decision tree), and calculate the model's evaluation indicators on the validation set, including accuracy (TP is the number of true positives, TN is the number of true negatives, FP is the number of false positives, and FN is the number of false negatives), recall rate
[0081] (in )wait.
[0082] Adjusting the Chi-Square Threshold And the parameters of the decision tree (such as the depth of the tree, the minimum number of samples of leaf nodes, min_sammples_leaf, etc.), repeat the cross-validation process to find the parameter combination that makes the model evaluation index optimal.
[0083] 4. Applying the extracted features to actual network traffic analysis can achieve a variety of important mission objectives. In terms of network attack behavior identification, building an accurate detection model based on the extracted features can effectively identify various common network attack forms such as DDoS attacks, malware traffic, port scanning, etc. Through in-depth comparison and learning of normal and abnormal traffic features, when network traffic data shows a match with known attack feature patterns, the system can quickly issue an alarm and take corresponding protective measures, thereby greatly improving network security.
[0084] In the task of analyzing network traffic behavior patterns, these extracted features help to gain in-depth insights into the inherent laws of network traffic. For example, we can clearly understand the distribution characteristics, flow trends, and patterns of traffic peaks in different time periods, different user groups, and different application scenarios. This has key guiding significance for the rational planning and allocation of network resources, ensuring the efficient and stable operation of network services and avoiding network performance degradation caused by traffic congestion or uneven resource allocation.
[0085] At the same time, in order to ensure the continued effectiveness and adaptability of feature extraction methods, it is necessary to actively collect feedback data from practical applications. Through careful analysis of these feedback data, we can accurately discover the shortcomings of current feature extraction methods when facing complex and changing network environments, such as the failure to effectively capture the unique features of certain new network attacks, or the decrease in the accuracy of feature extraction in specific network scenarios. Based on these findings, we can improve and optimize the feature extraction method in a targeted manner, continuously adjust the feature selection strategy, improve the model parameter settings, and even introduce new feature extraction technologies or algorithms, so that the feature extraction method can keep pace with the times, always maintain efficient support capabilities for network traffic analysis tasks, and provide solid and reliable technical support for network operation and management.
[0086] 3. Advantages of the present invention
[0087] ① For the first time, it is proposed to organically combine the chi-square test and the decision tree model for self-supervised feature extraction of network data streams. The chi-square test is used for feature screening, and the decision tree is used for model construction. The two work together to give full play to the advantages of the chi-square test in feature correlation screening and the ability of the decision tree in classification decision-making. It effectively overcomes the problem that the chi-square test alone lacks the ability to build complex models and the decision tree is prone to overfitting and high computational complexity in high-dimensional data, providing a new self-supervised learning model for network data stream feature extraction.
[0088] ② The chi-square-decision tree model uses the chi-square test to screen features, reducing the possibility that the decision tree is affected by irrelevant features, reducing the amount of calculation, and increasing the speed of model training. At the same time, the model comprehensively considers the data distribution characteristics during the construction process, which can effectively reduce the interference of data noise on the decision tree and enhance the stability of the model, so that when facing complex and changeable network data flows, the model can still maintain good performance. The model is evaluated and optimized using cross-validation. By adjusting the chi-square threshold and decision tree parameters (such as the depth of the tree, the minimum number of samples of leaf nodes, etc.), the optimal model configuration can be found, and the model's accuracy, recall rate, score and other evaluation indicators in different network environments and tasks can be further improved to ensure that the model has good generalization ability.
[0089] ③Innovatively integrate the task-driven concept throughout the entire network data flow feature extraction process, and perform comprehensive data preprocessing before feature extraction. Closely linking the task objectives with operations such as data cleaning and normalization makes the feature extraction process more systematic and scientific, which is relatively rare in traditional network data flow analysis methods, laying a solid foundation for subsequent feature extraction and analysis, and improving the effectiveness and adaptability of the overall method.
[0090] ④ A dynamic optimization mechanism for feature extraction methods based on feedback data from actual applications has been established. By continuously collecting and analyzing feedback data, defects in the method can be discovered in a timely manner and targeted improvements can be made, so that the feature extraction method can continuously adapt to changes in the network environment and new task challenges. This dynamic optimization mechanism ensures the effectiveness and reliability of the invention in long-term network operation and management, which is different from the traditional static feature extraction method and reflects the innovation and foresight of the present invention.
Claims
1. A network data stream feature extraction method based on task orientation and hybrid learning, characterized by: The following steps are included: Step 1: Construct network data flow feature extraction task objectives for the collected network data flow, including identifying network attack behaviors, analyzing network traffic patterns, and predicting network loads; Step 2: Preprocess the original data of the network data stream, perform data cleaning and data normalization; Step 3: According to the task objectives established in step 1, supervised learning methods are used to extract features from the preprocessed data; Step 4: Apply the feature extraction of step 3 to the actual network traffic analysis. Based on the task objectives constructed in step 1, compare the abnormal traffic characteristics with the normal traffic characteristics, obtain the network attack feature pattern, and perform network protection.
2. According to claim 1, a method for extracting network data stream features based on task orientation and hybrid learning is characterized by: The data cleaning described in step 2 includes removing duplicate data and removing invalid data.
3. The method for extracting network data stream features based on task orientation and hybrid learning according to claim 1 is characterized in that: The data normalization described in step 2 uses Z-score normalization to normalize each eigenvalue to a mean of 0 and a standard deviation of 1.
4. The method for extracting network data stream features based on task orientation and hybrid learning according to claim 1 is characterized in that: The supervised learning method described in step 3 adopts the chi-square-decision tree model, including the following steps:
1. Chi-square feature screening Assume that the data set is D = (X, y), where X is the feature matrix and y is the target variable. In the classification problem, the target variable y has C categories. For each feature xi∈X (i=1,2,…,n, n is the number of features), a contingency table is constructed; the rows of the contingency table represent the values of the feature xi, and the columns represent the categories of the target variable y; Calculate the chi-square statistic The calculation formula is Where R is the number of rows in the contingency table, which is the number of values of feature xi; Orc is the observed frequency, which is the number of samples in the actual data where feature xi takes the rth value and target variable y is the cth category; Erc is the expected frequency, which is the number of samples where the assumed feature xi and y The frequencies calculated without correlation; According to the pre-set chi-square threshold choose The features of subset ; 2. Decision Tree Construction In the feature subset χ subset and target variable y to build a decision tree; using the CART decision tree, at each node, calculate each feature χ j ∈χ subset Gini impurity (j = 1, 2, ..., m, m is the number of subset features); Gini impurity Where D is the data subset of the current node, K is the number of categories of the target variable y in D, and p k is the probability of category k in D; Select the feature with the largest reduction in Gini impurity as the split feature of the current node; Calculate the reduction of Gini impurity by feature: splitting the current node where V is the characteristic χ j The number of values, D v According to the characteristic χ j The sub-node data subset after the v-th value division; select ΔGini j The largest feature is split; Repeat the above splitting process until the stopping condition is met; the stopping condition is that the depth of the tree reaches a predetermined value, the number of samples in the leaf node is less than a threshold, or the reduction in Gini impurity is less than a threshold; 3. Model Evaluation and Optimization Cross-validation was used to evaluate the model performance. A 10-fold cross-validation was used to divide the data set into 10 subsets. In each fold, 9 subsets were used as training sets and 1 subset was used as validation set. A combined model including chi-square test feature screening and decision tree construction was constructed on the training set. The evaluation index of the model was calculated on the validation set. Including accuracy Among them, TP is the number of true positives, TN is the number of true negatives, FP is the number of false positives, and FN is the number of false negatives; Recall in, Adjusting the Chi-Square Threshold And the decision tree parameters, including the tree depth depth, the minimum number of leaf node samples min_sammples_leaf, repeat the cross-validation process to obtain the parameter combination that makes the model evaluation index optimal.