Internal / private network security baseline construction method based on normal traffic

By building an internal/dedicated network security baseline based on normal traffic, and using the pattern tree leaf node generation mechanism, the problem of high underreport rate of unknown threats in the internal/dedicated network is solved, and efficient abnormal detection effect is achieved.

CN120301710AActive Publication Date: 2025-07-11NO 30 INST OF CHINA ELECTRONIC TECH GRP CORP +1
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510774448.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-11
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect unknown threats in internal/private networks with high concealment and high technical level, resulting in a high misreport rate, and solutions based on abnormal characteristics have limitations in unknown threats.

Method used

Build an internal/dedicated network security baseline based on normal traffic, and generate pattern tree leaf nodes by preprocessing, classifying, clustering and timing feature extraction of normal traffic data, and use dynamic thresholds to judge abnormal traffic, and design a pattern tree leaf node generation mechanism to balance the missed-report rate and false alarm rate.

Benefits of technology

在保证高召回率的同时,降低未知威胁检测的漏报率,实现对内部/专用网络的高效异常检测。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120301710A_ABST
    Figure CN120301710A_ABST
Patent Text Reader

Abstract

The invention discloses an internal / private network security baseline construction method based on normal traffic, and relates to the field of traffic anomaly detection. The method comprises the following steps: firstly, constructing a normal traffic mode tree based on network normal traffic of a certain time span, then modeling internal / private network normal behaviors in combination with normal traffic time sequence characteristics of corresponding ports and protocols in a preset time window, and finally forming a security baseline for a target internal / private network. The missing report rate of unknown threats in an internal / private network is reduced, and the high requirements of the network for safety and management and control are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of traffic anomaly detection, and particularly to a method for constructing an internal / special network security baseline based on normal traffic. Background Art

[0002] The statements in this section only provide background information related to the present disclosure and may not constitute prior art.

[0003] For internal / special networks with high control, strong isolation, and high security levels, anomaly detection is a necessary technical link and deployment step. However, the high concealment and high technical level of network threats increase the difficulty of network behavior modeling, and most existing technical solutions based on anomaly features often have limitations when facing unknown threats due to the inability to obtain anomaly features. Therefore, it is urgent to define the normal behavior of the network driven by the normal data that accounts for the vast majority in the traffic, perform network normal behavior modeling, and solve the problem of false negatives in internal / special network anomaly detection. Summary of the Invention

[0004] The purpose of the present invention is: to make up for the deficiencies of existing work, the present invention mainly aims at the problem of high false negative rate of internal / special networks facing unknown threats, and proposes a method for constructing an internal / special network security baseline based on normal traffic. Specifically, the following problems are solved: (1) How to construct a customized security baseline according to the particularity of the internal / special network to solve the problem of high false negative rate in detecting unknown threats in the internal / special network.

[0005] (2) How to design the generation mechanism of the leaf nodes of the pattern tree to solve the balance problem between false negative rate and false positive rate.

[0006] (3) How to fully exploit the multi-dimensional features of normal network traffic to complete the modeling of normal behavior of the internal / special network.

[0007] Specifically, in view of the above existing problems and requirements, the present invention provides a method for constructing an internal / special network security baseline based on normal traffic, aiming at the actual security requirements of the internal / special network, making full use of the particularity of the internal / special network, fully exploiting the effective knowledge of the normal traffic data that accounts for the vast majority in the traffic data, customizing the normal traffic security baseline of the target network, and solving the problem of high false negative rate of the internal / special network facing unknown threats while ensuring a high level of recall rate.

[0008] The technical solution of the present invention is as follows: A method for constructing an internal / special network security baseline based on normal traffic, comprising: Step S1: Construct a traffic pattern tree based on the collected normal traffic data; Step S2: Deploy the constructed traffic pattern tree to the key nodes. For each network flow, match the leaf nodes of the traffic pattern tree. If a new leaf node needs to be created, it is an anomaly; otherwise, it is normal.

[0009] Further, the step S1 includes: Step S11: Preprocess the collected normal traffic data; Step S12: Classify the preprocessed traffic data; Step S13: Remove the redundant features from the classified traffic data, and cluster the remaining features to obtain a number of feature clusters; Step S14: Extract the time series features of the preprocessed traffic data, and splice them with a number of feature clusters; Step S15: Generate leaf nodes based on statistical features, and judge whether to generate new leaf nodes based on dynamic thresholds, so as to construct a complete traffic pattern tree.

[0010] Further, the preprocessing includes: Perform data cleaning on the collected original data, and remove the data with default values and incorrect data types.

[0011] Further, the data cleaning includes: Except for the destination port, protocol type, timestamp, and label, other attributes are discretized using KBinsDiscretizer, identify and remove redundant constant features, and standardize the remaining attributes to unify the numerical ranges of different features into a standard range.

[0012] Further, the step S12 includes: Classify the preprocessed traffic data according to the port number, and further classify it according to the protocol type on this basis.

[0013] Further, the redundant features include: destination port, protocol type, timestamp, and label.

[0014] Further, the clustering scheme adopted is: a feature grouping scheme based on genetic algorithms, and the evaluation criterion for grouping is the total mutual information of the features within the group.

[0015] Further, use an autoencoder to complete the extraction of the time series features of the traffic data, and use the long short-term memory (LSTM) as the network architecture of the encoder and decoder of the autoencoder.

[0016] Further, the update strategy of the leaf nodes includes: Strategy 1: If the The difference between the group feature mean and the mean of the th leaf node under the th feature cluster node is less than the threshold , and if the number of leaf nodes in the current feature cluster is less than 5, there is no need to create a new leaf node, and only the mean and variance of the th leaf node need to be updated; Strategy 2: If the difference between the mean of the th group of features in the current flow and the mean of the th leaf node under the th feature cluster node is less than the threshold , and if the number of leaf nodes in the current feature cluster is greater than 5, continue with the t-test. Calculate the value based on the mean and variance of the th leaf node and the current flow. If , it indicates that there is a probability of more than 95% that there is no need to create a new leaf node, and only the mean and variance of the th leaf node need to be updated; Strategy 3: If the mean of the current flow is not similar to all the leaf nodes under the feature cluster node, a new leaf node is generated under this feature cluster.

[0017] Furthermore, the threshold is dynamically updated through the entropy weighted value difference method.

[0018] Compared with the existing technologies, the beneficial effects of the present invention are as follows: 1. In view of the particularity of the internal / special network, the present invention designs a heuristic feature clustering algorithm from three dimensions: port, protocol type, and feature, and proposes a traffic security baseline construction scheme, which reduces the false negative rate of unknown threat detection while ensuring the recall rate.

[0019] 2. The present invention designs a generation strategy for the leaf nodes of the pattern tree and a dynamic threshold setting algorithm, forming a constrained pattern tree generation method, which balances the false negative rate and false positive rate of unknown threat detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a flowchart for constructing a traffic pattern tree; Figure 2 is a schematic diagram of the traffic pattern tree structure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] It should be noted that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0022] The features and performance of the present invention will be further described in detail below in conjunction with embodiments.

[0023] Embodiment 1 Considering the situation that normal behavior within an internal / private network is highly standardized and there are limited applications and services within such a network, this embodiment proposes a method for constructing a security baseline for an internal / private network based on normal traffic from three aspects: ports, protocols, and features. Aiming at the actual security requirements of the internal / private network, making full use of the particularity of the internal / private network, fully exploring the effective knowledge of the normal traffic data that accounts for the vast majority in the traffic data, and customizing the normal traffic security baseline of the target network. This method first constructs a normal traffic pattern tree based on the normal network traffic within a certain time span, and then combines the normal traffic time series features of the corresponding ports and protocols within a preset time window to model the normal behavior of the internal / private network, and finally forms a security baseline for the target internal / private network, reducing the false negative rate of unknown threats in the internal / private network and meeting the high requirements for security and control of such networks.

[0024] Specifically, a method for constructing a security baseline for an internal / private network based on normal traffic includes: Step S1: Construct a traffic pattern tree based on the collected normal traffic data; Step S2: Deploy the constructed traffic pattern tree to key nodes. For each network flow, match the leaf nodes of the traffic pattern tree. If a new leaf node needs to be created, it is abnormal; otherwise, it is normal.

[0025] In this embodiment, specifically, please refer to Figure 1 , the said step S1 includes: Step S11: Preprocess the collected normal traffic data; Step S12: Classify the preprocessed traffic data; Step S13: Remove redundant features from the classified traffic data and cluster the remaining features to obtain feature clusters; Step S14: Extract the time series features of the preprocessed traffic data and splice them with feature clusters; that is, use a time series model to extract the time series features of the traffic classified according to the port and protocol type, and splice the time series features with the features of each feature cluster; Step S15: Generate leaf nodes based on statistical features and determine whether to generate new leaf nodes based on dynamic thresholds, thereby constructing a complete traffic pattern tree.

[0026] In this embodiment, specifically, the preprocessing includes: Perform data cleaning on the collected raw data to remove data with default values and incorrect data types; data cleaning includes: Except for the destination port, protocol type, timestamp, and label, other attributes are discretized using KBinsDiscretizer, redundant constant features are identified and removed, and the remaining attributes are standardized to unify the numerical ranges of different features into a standard range; It should be noted that since the pattern tree is an unsupervised model and only learns the patterns of normal traffic during the training process, it is necessary to collect a sufficient number of normal traffic of the target network, and the data during the traffic collection period should cover all normal business behaviors and application service operations within the target network. First, perform data cleaning on the collected raw data to remove data with default values and incorrect data types. Specifically, except for the destination port, protocol type, timestamp, and label, other attributes are discretized using KBinsDiscretizer, redundant constant features are identified and removed, and the remaining attributes are standardized to unify the numerical ranges of different features into a standard range.

[0027] In this embodiment, specifically, the step S12 includes: Classify the preprocessed traffic data according to the port number, and on this basis, classify it according to the protocol type to obtain the first and second layer nodes of the traffic pattern tree.

[0028] In this embodiment, specifically, the redundant features include: destination port, protocol type, timestamp, and label.

[0029] In this embodiment, specifically, the clustering scheme adopted is: a feature grouping scheme based on genetic algorithms, and the evaluation criterion for grouping is the total mutual information of the features within the group.

[0030] It should be noted that there are still a large number of attribute features in the traffic information attribute features after preliminary data cleaning. If leaf nodes are generated based on these features, it will inevitably cause an explosion in the number of leaf nodes. Therefore, it is necessary to group the attribute features after data cleaning, aggregate highly similar features, and perform data dimensionality reduction. In this embodiment, a feature grouping scheme based on the Genetic Algorithm (GA) is adopted, and the evaluation criterion for grouping is the sum of the mutual information (MI) of the features within the group.

[0031] In this embodiment, it should be noted that the genetic algorithm belongs to an existing algorithm. The principle of the genetic algorithm will be described below in combination with the application scenario involved in this embodiment. Those skilled in the art can implement the specific operations based on the following description.

[0032] (1) Feature grouping by genetic algorithm The core idea of the genetic algorithm originates from Darwin's theory of evolution, and it solves complex optimization problems by simulating the processes of natural selection and inheritance. Each "individual" represents a feature grouping scheme. The encoding of the individual adopts integer encoding, and each gene corresponds to a feature, and its value represents the group to which the feature belongs. For example, for the problem of dividing 10 features into 3 groups, a possible individual encoding is , indicating that the 1st, 3rd, 6th, and 9th features belong to the 1st group, the 2nd, 5th, and 8th features belong to the 2nd group, and the rest belong to the 3rd group. Specifically, it includes the following steps: Step 1: Random population initialization, and at the same time introduce heuristic rules to improve the quality of the initial population. For example, using the mutual information matrix, tend to initialize the features with high mutual information into the same group. The "intelligent" initialization strategy helps the algorithm converge to a good solution faster, and the design of the fitness function is the same as that of the ant colony algorithm; Step 2: Tournament selection. Each time individuals are randomly selected from the population, and then the individual with the highest fitness among them is selected . The advantage of tournament selection is that it can maintain the diversity of the population, and at the same time gives better individuals more reproduction opportunities. By adjusting the scale of the tournament to balance the selection pressure and population diversity; Step 3: Crossover operation. Considering the encoding method, this embodiment designs an improved uniform crossover operation for the genetic algorithm to generate new solutions. During the crossover process, not only the genes of the parent individuals are exchanged, but also the validity of the offspring individuals after crossover is ensured (that is, it is ensured that the features will not be repeatedly assigned to multiple groups). Specifically, first perform standard uniform crossover, and then ensure that each feature belongs to only one group through a repair step, which not only maintains the advantages of uniform crossover but also ensures the feasibility of the solution; Step 4: Mutation operation. It is used to maintain the diversity of the population and prevent the algorithm from converging to the local optimal solution prematurely. In this embodiment, two mutation methods are implemented: one is simple random mutation, randomly selecting a feature and changing its group; the other is intelligent mutation based on mutual information, which tends to move the feature to the group where the feature with higher mutual information is located. The intelligent mutation strategy increases the probability of generating better solutions.

[0033] Step 5: To improve the performance of the algorithm, this embodiment also introduces some advanced strategies. For example, the elitist retention strategy is implemented to ensure that the best individuals in each generation are not lost during the evolution process. In addition, adaptive crossover and mutation probabilities are adopted, and these parameters are dynamically adjusted according to the diversity of the population to balance exploration and exploitation at different stages of the algorithm.

[0034] (2)Mutual information matrix calculation In the feature grouping algorithm, the calculation of the mutual information matrix is a core step, providing an important theoretical basis for subsequent feature clustering. As an important concept in information theory, mutual information can effectively measure the degree of mutual dependence between two random variables. Compared with the traditional correlation coefficient, the advantage of MI is that it can capture the non-linear relationship between features, which is particularly important in feature selection and grouping tasks.

[0035] From a theoretical perspective, the mutual information between two random variables and can be expressed as: , where and are marginal entropies, is the joint entropy. However, since the data in the dataset are continuous variables, a histogram-based method is used to estimate the mutual information. First, the equal-width partitioning strategy is used to discretize the continuous variables into several groups, then the joint probability distribution of the two variables is estimated through a two-dimensional histogram, and finally the marginal entropy and joint entropy are calculated based on the discretized data.

[0036] Due to the large number of features, the present invention adopts a parallel computing strategy to accelerate the calculation of the mutual information matrix during implementation. The calculation of the entire mutual information matrix is decomposed into multiple independent subtasks, and a multi-core processor is used to calculate the mutual information between multiple feature pairs simultaneously. Finally, the results of each subtask are combined into a complete mutual information matrix. To process large-scale feature sets, a row-based calculation method is further adopted, and only one row of the mutual information matrix, that is, the mutual information between one feature and all other features, is calculated each time. In terms of data preprocessing, all features are normalized to the same scale to ensure that the calculation of mutual information is not affected by the range of the original data.

[0037] In this embodiment, specifically, an autoencoder is adopted to extract the temporal features of traffic data, and the Long Short-Term Memory (LSTM) is used as the network architecture of the encoder and decoder of the autoencoder.

[0038] That is, in this embodiment, an autoencoder is adopted to complete the extraction of the temporal features of traffic data, and the Long Short-Term Memory (LSTM) is used as the network architecture of the encoder and decoder of the autoencoder. In this embodiment, a time window is set, and the traffic within the window is preprocessed and then sent to the encoder end of the autoencoder for model training. The specific steps are as follows: Step 1: The encoder represents the preprocessed input temporal data as a fixed-length low-dimensional latent representation, representing the temporal features of the input data. The output of the LSTM can be the hidden state at the last time step (as the latent representation) or the hidden states at all time steps; Step 2: The decoder reconstructs the original data from the latent representation. The RepeatVector layer is used to repeat the latent representation multiple times to match the time step of the original data, and the LSTM unit is used to gradually reconstruct the original temporal data. The fully connected layer is used to map the output of the LSTM to the original feature dimension; Step 3: The mean squared error loss function is used to calculate the difference between the reconstructed data and the original data. The loss function is expressed as: , where is the original data, is the reconstructed data, is the number of samples.

[0039] In this embodiment, it should be noted that generally, each port usually corresponds to a relatively fixed service. For example, the web port 80 usually corresponds to the web service. In the internal / private network environment, due to reasons such as high security level and strong management control, the application services within the network are usually determined. Therefore, constructing a tree through the normal network flows of each port is to construct a pattern tree for the service corresponding to each port.

[0040] As Figure 2 shown, the first layer of the pattern tree is the port, the second layer is the protocol, the third layer is the feature cluster, and each feature cluster contains similar features grouped by features. The last layer is the statistical feature. In this embodiment, the mean and variance are used as examples for illustration. Among them, each network flow will become a branch in the pattern tree, and the update strategy of the leaf node is as follows: Strategy 1: If the difference (similarity_threshold) between the mean of the th group of features of the current flow and the mean of the th leaf node under the th feature cluster node (the third layer) is less than the threshold If the number of leaf nodes in the current feature cluster is less than 5 (min_samples), there is no need to create a new leaf node, and only the mean and variance of the th leaf node need to be updated; Strategy 2: If the difference between the mean of the th group of features in the current stream and the mean of the th leaf node under the th feature cluster node (the third layer) is less than the threshold , and the number of leaf nodes in the current feature cluster is greater than 5, then continue with the t-test. Calculate the value based on the mean and variance of the th leaf node and the current stream. If , it means that there is a probability of more than 95% that there is no need to create a new leaf node, and update the mean and variance of the th leaf node; Among them, refers to the statistic ( -statistic), which is the core calculated value of the test ( -test) and is used to measure the degree of difference between the sample data and the hypothesis, defined as:

[0041] Among them, is the sample mean, is the population hypothesized mean (the reference value in the null hypothesis), is the sample standard deviation, is the sample size; Among them, the value ( -value) is the probability of the null hypothesis being true for the observed data (or more extreme data) in a statistical hypothesis test. , usually considered that the result has statistical significance. , then the null hypothesis cannot be rejected.

[0042] Strategy 3: If the mean of the current stream is not similar to all the leaf nodes under the feature cluster node (the third layer), a new leaf node is generated under this feature cluster.

[0043] It should be noted that in this embodiment, the threshold is dynamically updated by the entropy weighted value difference method. The threshold of the feature cluster node is updated, and the distribution entropy of the value difference is used to calculate the threshold, which can capture the complexity and diversity of data changes. The dynamic threshold is defined as:

[0044] Among them: is the scaling factor; is the total number of leaf nodes under the cluster node; ; is the entropy of the weighted difference, used to measure the uncertainty of the weighted difference, and is defined as:

[0045] is the total number of leaf nodes.

[0046] Among them, is the normalized weighted difference. The normalized weighted difference refers to standardizing the weighted difference and is defined as:

[0047] Among them: is the leaf node the mean value stored for the first time; represents the leaf node the dynamic mean value (if similar flow eigenvalues are merged into an existing leaf node, the mean value of the leaf node will change dynamically); represents the leaf node the number of similar flows that have been merged; is the sum of the weighted differences calculated for all leaf nodes under the cluster node.

[0048] The embodiments described above only represent the specific implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the protection scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the technical solution of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application.

[0049] This background technology section is provided to generally present the context of the present invention. The work of the currently named inventors, to the extent described in this background technology section, and aspects described in this section that are not prior art at the time of filing this application are neither expressly nor impliedly admitted to be prior art of the present invention.

Claims

1. A method for constructing an internal / private network security baseline based on normal traffic, characterized in that Including: Step S1: Based on the collected normal traffic data, construct a traffic pattern tree. Step S2: Deploy the constructed traffic pattern tree to key nodes. For each network flow, match the leaf nodes of the traffic pattern tree. If a new leaf node needs to be created, it is abnormal; otherwise, it is normal. The said Step S1 includes: Step S11: Preprocess the collected normal traffic data. Step S12: Classify the preprocessed traffic data. Step S13: Remove the redundant features from the classified traffic data, and cluster the remaining features to obtain feature clusters; Step S14: Extract the time series features of the preprocessed traffic data and splice them with feature clusters; Step S15: Generate leaf nodes based on statistical features, and based on dynamic thresholds, determine whether to generate new leaf nodes, thereby constructing a complete traffic pattern tree.

2. The method for constructing an internal / private network security baseline based on normal traffic according to claim 1, characterized in that The said preprocessing includes: Perform data cleaning on the collected original data to remove data with default values and incorrect data types.

3. The method for constructing an internal / special network security baseline based on normal traffic according to claim 2, wherein Data cleaning includes: Except for the destination port, protocol type, timestamp, and label, other attributes are discretized using KBinsDiscretizer, redundant constant features are identified and removed, and the remaining attributes are standardized so that the numerical ranges of different features are unified into a standard range.

4. A method for constructing an internal / special network security baseline based on normal traffic according to claim 1, characterized in that The said Step S12 includes: Classify the preprocessed traffic data according to the port number, and on this basis, classify it according to the protocol type.

5. A method for constructing an internal / private network security baseline based on normal traffic according to claim 1, characterized in that, The said redundant features include: destination port, protocol type, timestamp, and label.

6. The method for constructing an internal / private network security baseline based on normal traffic according to claim 1, wherein The clustering scheme adopted is: a feature grouping scheme based on the genetic algorithm, and the evaluation criterion for grouping is the sum of the mutual information of the features within the group.

7. A method for constructing an internal / special network security baseline based on normal traffic according to claim 1, characterized in that An autoencoder is used to extract the temporal features of the traffic data, and the long short-term memory (LSTM) is used as the network architecture of the encoder and decoder of the autoencoder.

8. A method for constructing an internal / private network security baseline based on normal traffic according to claim 1, characterized in that The update strategy of the leaf nodes includes: Strategy 1: If the difference between the mean value of the group of feature means of the current stream and the mean value of the th leaf node under the th feature cluster node is less than the threshold , and the number of leaf nodes of the current feature cluster is less than 5, then there is no need to create a new leaf node, and it is only necessary to update the mean value and variance of the th leaf node; Strategy 2: If the difference between the mean value of the group of feature means of the current stream and the mean value of the th leaf node under the th feature cluster node is less than the threshold , and the number of leaf nodes of the current feature cluster is greater than 5, then continue with the t-test. Calculate the value based on the mean and variance of the th leaf node and the current stream. If , it indicates that there is a probability of more than 95% that there is no need to create a new leaf node. Update the mean and variance of the th leaf node; Among them, refers to a statistic, which is the core calculated value of the test, used to measure the degree of difference between the sample data and the hypothesis, and is defined as: Among them, is the sample mean, is the population hypothesized mean, is the sample standard deviation, is the sample size; Strategy three: If the mean of the current flow is not similar to all the leaf nodes under the feature cluster node, a new leaf node is generated under this feature cluster.

9. A method for constructing an internal / private network security baseline based on normal traffic according to claim 8, characterized in that Threshold is dynamically updated by the entropy weighted value difference method ​

Citation Information

Patent Citations

  • Mode tree optimization method and device, and electronic equipment

    CN113821742A

  • Malicious traffic identification method and device, equipment and storage medium

    CN113992349A

  • 5G slice network anomaly detection method based on virtual network flow analysis

    CN114401516A

  • Security situation awareness processing method, system and equipment with built-in data processing unit

    CN117914547A

  • Network behavior baseline-based network anomaly detection method and system

    CN117978538A