Network intrusion detection method and system based on topological data analysis

By generating new minority class samples in low-dimensional feature space, the challenges of network intrusion detection model in category imbalance and topological structure capture are solved, and more accurate network intrusion detection is achieved.

CN120582901APending Publication Date: 2025-09-02GUODIAN DADU RIVER POWER ENG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510924044.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The existing network intrusion detection models are difficult to capture the relationship and topological structure between data categories when processing large-scale high-dimensional traffic data, and face the problem of class imbalance, which leads to the detection model tending to learn the distribution of most class features and ignores the attack behavior of a few classes, resulting in a high false negative rate.

Method used

Through topological data analysis methods, network traffic data is mapped to low-dimensional feature space, new minority class samples are generated, class balance training sets are constructed, and network intrusion detection models are established. New minority class samples are generated using topological graphs and interpolation algorithms to maintain the topological structure and relationship integrity of the data.

Benefits of technology

It effectively improves the feature quality of a few types of attacks, improves the accuracy of network intrusion detection, effectively deals with category imbalance problem, reduces overfitting and out-of-domain samples generation, and improves the performance of the detection model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120582901A_ABST
    Figure CN120582901A_ABST
Patent Text Reader

Abstract

The invention provides a network intrusion detection method and system based on topological data analysis, and relates to the field of data processing, and the method comprises the steps: obtaining network flow data; based on the filtering function, mapping the network traffic data to a feature space, the dimension of which is lower than that of the network traffic data; determining a plurality of coverage of the low-dimensional space; for each coverage, data points included in the coverage are clustered based on the similarity of any two data points through a clustering algorithm, clustering clusters included in the coverage are determined, a topological graph corresponding to network flow data is established, new minority class samples are generated in combination with an interpolation algorithm, and a class balance training set is constructed; establishing a network intrusion detection model; training a network intrusion detection model based on the class balance training set; the real-time network flow data is acquired, the network intrusion type is determined based on the real-time network flow data through the network intrusion detection model, and the method has the advantages of solving the problem of class imbalance in a network intrusion detection scene and improving the accuracy of network intrusion detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular to a network intrusion detection method and system based on topology data analysis. Background Art

[0002] To improve the accuracy and adaptability of NIDS, various machine learning and deep learning techniques have been introduced. These techniques can process large-scale network traffic data and extract meaningful behavioral patterns, thereby enhancing the recognition capabilities of detection models. Although model performance can be improved through continuous improvement and optimization of network structures, the magnitude of this improvement is shrinking. A major challenge is that, faced with large-scale, high-dimensional traffic data, intrusion detection models based on convolutional neural networks often struggle to capture the interrelationships and topological structures between data classes, such as the time series characteristics of traffic and the network topology. This limits their performance in complex network environments. Another major challenge is class imbalance, where data points representing normal behavior far outnumber those representing abnormal behavior. This imbalance causes detection models to tend to learn the feature distribution of the majority class while ignoring potential attack behaviors from the minority class, resulting in a high false negative rate. To address class imbalance, researchers have proposed various strategies. Common approaches include oversampling and undersampling. For example, the synthetic minority oversampling technique (SMOTE) improves model learning by creating new minority class instances in the feature space, increasing the sample size of the minority class. The key advantage of SMOTE is that it maintains data diversity by interpolating between minority class samples to generate new samples, avoiding the overfitting problem caused by simply duplicating minority class samples. However, directly interpolating oversampling in the original input data may generate out-of-domain samples, which may affect model accuracy. Undersampling methods balance the class distribution by reducing the number of majority class samples, but this can result in the loss of important information. Furthermore, ensemble methods such as random forests and boosting have also been used to improve the ability to identify minority classes. For example, existing techniques combine a random forest model for feature selection with a gradient boosted decision tree model for classification, effectively alleviating the class imbalance problem. Another example is the construction of a dynamic ensemble model that incorporates a dynamic undersampling mechanism into the boosting framework to address the overfitting problem of noisy samples caused by undersampling. Deep learning methods have also been applied to address class imbalance. For example, cost-sensitive loss functions are used to optimize the model's learning process for the minority class. These methods have shown effectiveness in addressing class imbalance, but their application in the cybersecurity field remains challenging, particularly in preserving the inherent structure and relationships of the data.

[0003] Therefore, it is necessary to provide a network intrusion detection method and system based on topology data analysis to solve the class imbalance problem in network intrusion detection scenarios and improve the accuracy of network intrusion detection. Summary of the Invention

[0004] The present invention provides a network intrusion detection method based on topological data analysis, comprising: acquiring network traffic data; mapping the network traffic data to a feature space based on a filter function, wherein the dimension of the feature space is lower than the dimension of the network traffic data; determining multiple covers of the low-dimensional space; for each of the covers, clustering the data points included in the cover based on the similarity of any two data points through a clustering algorithm to determine the cluster clusters included in the cover; establishing a topological graph corresponding to the network traffic data based on the cluster clusters included in each of the covers; generating new minority class samples based on the topological graph corresponding to the network traffic data and an interpolation algorithm to construct a class-balanced training set; establishing a network intrusion detection model; training the network intrusion detection model based on the class-balanced training set; acquiring real-time network traffic data, and determining the type of network intrusion based on the real-time network traffic data through the network intrusion detection model.

[0005] Furthermore, the optimization goal of the filtering function is:

[0006] Among them, C(w h (e), w l (e)) is the optimization target of the filtering function, w h (e) is the weight of the e-th dimension in the high-dimensional space corresponding to the network traffic data, w l (e) is the weight of the e-th dimension in the feature space, and E is the total number of dimensions included in the feature space.

[0007] Furthermore, the weight of the e-th dimension in the high-dimensional space corresponding to the network traffic data is calculated based on the following formula: Where k is a specified number, is the data point x i Neighboring data points within a distance k range The minimum distance, ρ i is the data point x i The distance offset, σ i is the parameter used to scale the distance.

[0008] Furthermore, the weight of the e-th dimension in the feature space is calculated based on the following formula: l (e) = (1 + a(z i -z j ) 2b ) -1 , where z iis the coordinate of data point i in the feature space, z j is the coordinate of data point j in the feature space, and a and b are adjustment parameters.

[0009] Furthermore, determining multiple covers of the low-dimensional space includes: determining multiple covers of the low-dimensional space based on the number of intervals, the fixed length of the cover, and the overlap rate.

[0010] Furthermore, the overlap ratio is calculated based on the following formula: Where p is the overlap ratio, r is the fixed length of coverage, l is the number of intervals, and n is the total number of data points.

[0011] Furthermore, based on the topological graph corresponding to the network traffic data and the interpolation algorithm, new minority class samples are generated, including: selecting multiple basic nodes from the minority class nodes in the topological graph corresponding to the network traffic data; for each basic node, determining the nearest neighbor node of the basic node from the same type of nodes of the basic node in the topological graph corresponding to the network traffic data; and generating new minority class samples based on the basic node and the nearest neighbor node of the basic node through the interpolation algorithm.

[0012] Furthermore, the nearest neighbor node of the base node is determined based on the following formula: nn(v)=argmin u∈S ||h v -h u ||, where nn(v) is the base node x v The nearest neighbor node, S is the base node x v The set of similar nodes, h v is the base node x v In the representation of the feature space, h u is the base node x v Similar nodes x u Representation in the feature space.

[0013] Furthermore, the interpolation algorithm is: v′ =(1-δ)·x v +δ·x nn(v) , where x v′ is a new minority class sample, x nn(v) is the nearest neighbor node of the base node, and δ is a random number between 0 and 1.

[0014] The present invention provides a network intrusion detection system based on topological data analysis, comprising: a data acquisition module for acquiring network traffic data; a data filtering module for mapping the network traffic data to a feature space based on a filtering function, wherein the dimension of the feature space is lower than the dimension of the network traffic data; a coverage determination module for determining multiple coverages of the low-dimensional space; a data clustering module for clustering the data points included in each coverage based on the similarity of any two data points using a clustering algorithm to determine the cluster clusters included in the coverage; a topology map establishment module for establishing a topology map corresponding to the network traffic data based on the cluster clusters included in each coverage; a class balancing module for generating new minority class samples based on the topology map corresponding to the network traffic data and an interpolation algorithm to construct a class-balanced training set; a model establishment module for establishing a network intrusion detection model; a model training module for training the network intrusion detection model based on the class-balanced training set; and an intrusion detection module for acquiring real-time network traffic data and determining the type of network intrusion based on the real-time network traffic data using the network intrusion detection model.

[0015] Compared with the existing technology, the network intrusion detection method and system based on topology data analysis provided by the present invention have at least the following beneficial effects:

[0016] The network intrusion detection method based on topological data analysis can effectively improve the feature quality of minority attack types by generating new minority class samples in the topological space, helping to more accurately identify abnormal behavior in network traffic and effectively addressing the class imbalance problem. The network intrusion detection method based on topological data analysis adopts a new sample selection strategy, selecting nearest neighbor samples with topological relationships for synthesis through a distance metric in a low-dimensional space. Not only does it maintain the topological structure of the original data, it also effectively avoids the problems of overfitting and generating out-of-domain samples that may arise in traditional oversampling methods. Experimental results show that the network intrusion detection method based on topological data analysis outperforms other methods overall on the CIC-IDS2017 and CIC-MalMen2022 datasets, verifying the effectiveness of the network intrusion detection method based on topological data analysis in addressing the class imbalance problem. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] This specification will be further described in the form of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting, and in these embodiments, the same numbers represent the same structures, wherein:

[0018] Figure 1 is a flowchart of a network intrusion detection method based on topology data analysis according to some embodiments of this specification;

[0019] Figure 2 is a schematic diagram of a topological diagram according to some embodiments of this specification;

[0020] Figure 3 is a schematic diagram of experimental results under different enhanced samples according to some embodiments of this specification;

[0021] Figure 4 This is a module diagram of a network intrusion detection system based on topology data analysis according to some embodiments of this specification. DETAILED DESCRIPTION

[0022] To more clearly illustrate the technical solutions of the embodiments of this specification, the following briefly describes the drawings required for describing the embodiments. Obviously, the drawings described below are merely examples or embodiments of this specification. Those skilled in the art can apply this specification to other similar scenarios based on these drawings without inventive effort. Unless otherwise apparent from the context or otherwise noted, the same reference numerals in the figures represent the same structure or operation.

[0023] Figure 1 is a flow chart of a network intrusion detection method based on topology data analysis according to some embodiments of this specification, such as Figure 1 As shown, the network intrusion detection method based on topology data analysis may include the following steps.

[0024] Step 110: Obtain network traffic data.

[0025] Network traffic data can include data in multiple dimensions for network intrusion detection, such as basic packet information (e.g., source IP address and destination IP address, port number, protocol type, etc.), packet content characteristics (e.g., data length, actual data content in the packet, etc.), traffic characteristics (e.g., traffic size, session duration, etc.), behavioral model characteristics (e.g., time distribution of traffic, user behavior data, etc.), etc.

[0026] Step 120: Map the network traffic data to a feature space based on the filter function.

[0027] Among them, the dimension of the feature space is lower than that of network traffic data.

[0028] Specifically, the original high-dimensional data X is mapped to a low-dimensional feature space Z through a filter function f, expressed as f:X→Z. The feature space Z is a one-dimensional or low-dimensional Euclidean space.

[0029] Preferably, the optimization goal of the filter function is:

[0030]

[0031] Among them, C(w h (e),w l (e)) is the optimization objective function of the filtering function, w h (e) is the weight of the e-th dimension in the high-dimensional space corresponding to the network traffic data, w l (e) is the weight of the e-th dimension in the feature space, E is the total number of dimensions included in the feature space, w h (e) and w l (e) The optimization objective function used to optimize the filtering function so as to maintain similar structures between the high-dimensional space (i.e., the high-dimensional space corresponding to the network traffic data) and the low-dimensional space (i.e., the feature space).

[0032] Preferably, the weight of the e-th dimension in the high-dimensional space corresponding to the network traffic data is calculated based on the following formula:

[0033]

[0034] Where k is a specified number, is the data point x i Neighboring data points within a distance k range The minimum distance, ρ i is the data point x i The distance offset, σ i is the parameter used to scale the distance.

[0035] Preferably, the weight of the e-th dimension in the feature space is calculated based on the following formula:

[0036] w l (e) = (1 + a(z i -z j ) 2b ) -1

[0037] Among them, z i is the coordinate of data point i in the feature space, z j is the coordinate of data point j in the feature space, a and b are adjustment parameters used to control the distance between data point i and data point j in the low-dimensional space.

[0038] Step 130 : determining multiple covers of the low-dimensional space.

[0039] Specifically include:

[0040] Determine multiple covers of the low-dimensional space based on the number of intervals, fixed length of the covers, and overlap ratio.

[0041] Preferably, the overlap ratio is calculated based on the following formula:

[0042]

[0043] Where p is the overlap ratio, r is the fixed length of coverage, l is the number of intervals, and n is the total number of data points.

[0044] Specifically, a finite set of covers U = {U α} α∈A , where each cover U α are all subintervals in the space Z, and all covers form a finite index set A. Each cover is of equal length and has intervals, and there is p% overlap between adjacent covers. Since the filter function f is continuous, the set f -1 (U α ) forms an open cover of the network traffic data X.

[0045] Step 140 : For each coverage, cluster the data points included in the coverage based on the similarity between any two data points using a clustering algorithm to determine the cluster cluster included in the coverage.

[0046] Specifically, in each coverage f -1 (U α ), the data points are clustered based on the similarity of the original data. In this step, each cluster can be regarded as a set of subsets in the network traffic data X. For each data point set f in each subset α, -1 (U α ) can be further divided into Record Among them, V(α,i) is the cluster center, k α is f -1 (U α ) is the number of clusters.

[0047] Step 150: Establish a topology map corresponding to the network traffic data based on each covered cluster.

[0048] For example only, Figure 2 is a schematic diagram of a topological diagram according to some embodiments of this specification, such as Figure 2 As shown in the figure, the blue circles represented by dotted lines represent nodes in the topology graph, and edges are added between nodes according to specific rules. The black circles and triangles in the nodes represent majority class samples and minority class samples, respectively. The red asterisks represent newly generated samples, and the circles scattered outside the nodes represent noise.

[0049] Each element W(x v ,x u ) describes the topological relationship between sample points:

[0050]

[0051] Among them, V(x v ) and V(x v ) represents the data point x u and data point x v The cluster where the two points are located, 1 is the indicator function, when the two points are in the same cluster or in different but connected clusters, W(x v ,x u ) is 1, otherwise it is 0.

[0052] Step 160 : Based on the topological graph corresponding to the network traffic data and the interpolation algorithm, new minority class samples are generated to construct a class-balanced training set.

[0053] Specifically include:

[0054] Selecting multiple basic nodes from a minority class of nodes in a topology graph corresponding to the network traffic data;

[0055] For each basic node, determine the nearest neighbor node of the basic node from the same type of nodes in the topology graph corresponding to the network traffic data;

[0056] New minority class samples are generated based on the base node and its nearest neighbor nodes through the interpolation algorithm.

[0057] To select minority class samples, randomly select a subset of minority class nodes from the generated topology as base nodes, particularly those near the class boundary. Because these nodes contain more information, the generated samples are more helpful for classifier learning. A class boundary is defined as a node with more than half the size of the majority class. For each sample in a selected base node, find its nearest neighbor samples (similar samples) in the node.

[0058] To better utilize the topology, we not only oversample each node, but also select a small number of samples from two connected nodes. The connections between these nodes reflect the important relationships between data points. By sampling between these nodes, we can more effectively preserve the key relationships in the topology.

[0059] Preferably, the nearest neighbor node of the base node is determined based on the following formula:

[0060] nn(v)=argmin u∈S ||h v -h u ||

[0061] Among them, nn(v) is the basic node x vThe nearest neighbor node, S is the base node x v The set of similar nodes, h v is the base node x v In the feature space representation, h u is the base node x v Similar nodes x u Representation in feature space.

[0062] The interpolation algorithm is:

[0063] x v′ =(1-δ)·x v +δ·x nn(v)

[0064] Among them, x v′ is a new minority class sample, x nn(v) is the nearest neighbor node of the base node, and δ is a random number between 0 and 1.

[0065] Based on the topological graph corresponding to the network traffic data and the newly generated minority class samples, an enhanced topological graph is ultimately generated. New minority class samples are generated on the existing topological structure while preserving the structural characteristics of the graph and the topological relationships between nodes. In this intermediate mapping space, the dimensionality is lower than the original dimensionality, so samples from the same class are more densely distributed. Because intra-class similarities and inter-class differences are fixed by the topological structure, interpolation can more reliably generate in-domain samples.

[0066] Step 170: Establish a network intrusion detection model.

[0067] Among them, the network intrusion detection model is a machine learning model.

[0068] Step 180: training a network intrusion detection model based on the class-balanced training set.

[0069] Step 190: Acquire real-time network traffic data, and determine the type of network intrusion based on the real-time network traffic data using a network intrusion detection model.

[0070] The beneficial effects of this method are described below with reference to experiments.

[0071] Two unbalanced public datasets CIC-IDS2017 (as shown in Table 1) and CIC-MalMen2022 (as shown in Table 2) are used. For data partitioning, 75% of the traffic is randomly selected for training and the remaining 25% is used for testing.

[0072] Table 1

[0073]

[0074]

[0075]

[0076] The CIC-IDS2017 dataset contains benign traffic similar to real-world data and the latest common attack types. The dataset contains 3,119,345 instances and 83 features, including 1 class label and 82 traffic features. The labels include benign traffic and 14 modern network attacks.

[0077] Table 2

[0078]

[0079]

[0080]

[0081] CIC-MalMen2022 consists of spyware, ransomware, and Trojan malware, providing a dataset that closely resembles real-world malware behavior. Malicious records are further categorized into 15 different attack types. Each record includes 55 features, including 26 new memory-based features generated by a feature extractor.

[0082] In the CIC-IDS2017 dataset, normal behavior samples account for over 80%. In contrast, the CIC-MalMen2022 dataset provides a more even distribution of samples. All datasets were preprocessed before model training. This included removing duplicate headers and constant columns, filling blank data with "0," and converting character features into numerical features using one-hot encoding. In network traffic datasets, the numerical values ​​of features vary greatly. To ensure comparability between features, each column was normalized using the Z-Score.

[0083] To comprehensively evaluate model performance in multi-class environments, macro-average and micro-average metrics are generally used. Macro-average treats all classes equally and is therefore significantly affected by attack class detection results. Micro-average treats each sample as a whole, potentially overemphasizing detection performance for the normal class. Given the class imbalance in network intrusion detection, identifying less common attack types is equally important. Therefore, macro-average metrics are used to comprehensively evaluate model performance. These metrics should include at least accuracy, precision, recall, and F1-score.

[0084] In order to verify the effectiveness of the network intrusion detection model of this method, the experiment compared the network intrusion detection model with some representative detection models and imbalanced learning methods. The comparison models include:

[0085] (1) Deep Neural Networks (DNN) model: DNN is the most basic classification model. The purpose of setting up DNN is to observe the detection effect of raw data.

[0086] (2) AB-LightGBM: It proposes to use the ADASYN (Adaptive Synthetic Sampling Approach for Imbalanced Learning) oversampling technology to increase the number of attack samples in the training data to solve the data imbalance problem. At the same time, Bayesian optimization is used to tune the hyperparameters of the LightGBM (Light Gradient Boosting Machine) model to improve detection accuracy and reduce the computational burden.

[0087] (3) SMOTE-TomekLink: This paper proposes a wireless sensor network (WSN) intrusion detection method that combines the SMOTE-TomekLink data balancing technology with a machine learning algorithm. This method uses the SMOTE-TomekLink technology to generate a balanced data set and uses the Random Forest (RF) algorithm to improve detection accuracy.

[0088] (4) TMG-IDS: A data augmentation model based on Generative Adversarial Networks (GAN) to address the data imbalance problem in network intrusion detection. Specifically, the proposed TMG-IDS method generates different types of attack data simultaneously through a multi-generator structure and introduces a classifier structure to optimize the generator and discriminator. The generator loss includes the cosine similarity between the generated samples and the original samples, as well as the cosine similarity between the generated samples and other types of generated samples, thereby improving the quality of the generated samples and reducing the overlap between classes.

[0089] Table 3

[0090]

[0091]

[0092] The proposed network intrusion detection model uses a filter function with parameters a and b set to 1.73 and 0.79, respectively. This aims to find the optimal balance between the Gaussian kernel and the cosine kernel to maximize the preservation of the topological structure of the high-dimensional space within the low-dimensional embedding. The overlap percentage of the coverage intervals is set to 40%, and r is 30. The interpolation parameter for the synthetic samples is a random number between 0 and 1. The detailed framework of the network intrusion detection model is shown in Table 3.

[0093] To verify the effectiveness of the network intrusion detection model, a large number of experimental tests were conducted. The results show that compared with other intrusion detection models, the network intrusion detection model has improved detection accuracy and efficiency, especially in dealing with class imbalance.

[0094] Table 4 compares the F1-score results of the network intrusion detection model and other models for detecting multiple attack types on the CIC-IDS2017 dataset. As can be seen from the table, the network intrusion detection model performs well in detecting most attack types. For example, when detecting Heartbleed, the network intrusion detection model achieved an F1-score of 74.51%, a 63.98% improvement over SMOTE-TomekLink, a 20.97% improvement over TMG-IDS, and a 17.97% improvement over AB-LightGBM. This significant improvement is primarily due to the network intrusion detection model's ability to learn from the distribution of minority attack samples, generating more representative attack samples, thereby improving its detection capabilities for minority attacks.

[0095] The experimental results on the CIC-MalMen2022 dataset are even more significant. The detailed comparative data listed in Table 5 shows that the network intrusion detection model achieved an F1-score of 97.96% when detecting SpywareTIBS, an improvement of 30.28% over TMG-IDS, 27.48% over SMOTE-TomekLink, and 24.20% over AB-Light GBM.

[0096] This performance improvement is not only reflected in certain specific attack types, but the network intrusion detection model also performs well in overall detection accuracy. Tables 6 and 7 show the overall performance comparison of different methods on the CIC-IDS2017 and CIC-MalMen2022 datasets, respectively.

[0097] As can be seen in Table 6, the network intrusion detection model outperforms other methods in all evaluation metrics, except for accuracy, where it falls slightly short of the AB-LightGBM method by 0.16%. This is because the features generated by TopoSMOTE place a greater weight on minority class detection, resulting in a decrease in overall accuracy. For the more complex CIC-MalMen2022 dataset, the performance of each model generally declines, but the network intrusion detection model still outperforms other models. As can be seen in the table, TopoSMOTE achieves an accuracy of 93.41%, a precision of 85.44%, a recall of 86.37%, and a false positive rate of only 1.03%.

[0098] Table 4

[0099]

[0100]

[0101] Table 5

[0102]

[0103] Table 6

[0104]

[0105]

[0106] From the above comparative experimental results, it can be seen that the network intrusion detection model outperforms other methods in all evaluation indicators, especially in dealing with class imbalance. These results verify the effectiveness and robustness of the TopoSMOTE method in network intrusion detection.

[0107] Table 7

[0108]

[0109] Additional ablation experiments are conducted to verify the motivation of this method, namely, the impact of considering nodes with connected edges on performance in the minority class sample selection strategy, and to verify the effectiveness of using distance in low-dimensional space as a metric for selecting nearest neighbor samples. The experimental results under different augmented samples are shown in Figure 2. Figure 3 As shown in the figure, TopoSMOTE-edge excludes edge-connected minority class samples from the selected samples and samples only from nodes. TopoSMOTE-low selects nearest neighbor samples without using a low-dimensional mapping space, but instead uses the original high-dimensional space to calculate distances. The comparison results in the figure show that the performance of the network intrusion detection model decreases under both simplified conditions, validating the effectiveness of considering topological relationships and using low-dimensional space to calculate distances in improving detection performance.

[0110] In summary, the network intrusion detection method based on topological data analysis can effectively improve the feature quality of minority attack types by generating new minority class samples in the topological space, helping to more accurately identify abnormal behavior in network traffic and effectively addressing the class imbalance problem. The network intrusion detection method based on topological data analysis adopts a new sample selection strategy, selecting nearest neighbor samples with topological relationships for synthesis through a distance metric in a low-dimensional space. Not only does it maintain the topological structure of the original data, it also effectively avoids the problems of overfitting and generating out-of-domain samples that may arise in traditional oversampling methods. Experimental results show that the network intrusion detection method based on topological data analysis outperforms other methods overall on the CIC-IDS2017 and CIC-MalMen2022 datasets, verifying the effectiveness of the network intrusion detection method based on topological data analysis in addressing the class imbalance problem.

[0111] Figure 4 is a schematic diagram of a module of a network intrusion detection system based on topology data analysis according to some embodiments of this specification, such as Figure 4 As shown, the network intrusion detection system based on topological data analysis may include a data acquisition module, a data filtering module, a coverage determination module, a data clustering module, a topology map building module, a class balancing module, a model building module, a model training module and an intrusion detection module.

[0112] The data acquisition module can be used to obtain network traffic data;

[0113] The data filtering module may be used to map the network traffic data to a feature space based on a filter function, wherein the dimension of the feature space is lower than the dimension of the network traffic data;

[0114] The cover determination module can be used to determine multiple covers of the low-dimensional space;

[0115] The data clustering module can be used to cluster the data points included in each coverage based on the similarity between any two data points through a clustering algorithm to determine the cluster clusters included in the coverage;

[0116] The topology map building module can be used to build a topology map corresponding to the network traffic data based on each covered cluster;

[0117] The class balancing module can be used to generate new minority class samples based on the topological graph and interpolation algorithm corresponding to the network traffic data and construct a class-balanced training set;

[0118] The model building module can be used to build network intrusion detection models;

[0119] The model training module can be used to train network intrusion detection models based on class-balanced training sets;

[0120] The intrusion detection module can be used to obtain real-time network traffic data and determine the type of network intrusion based on the real-time network traffic data through a network intrusion detection model.

[0121] The network intrusion detection system based on topology data analysis can be used to execute the network intrusion detection method based on topology data analysis. For more description of the network intrusion detection system based on topology data analysis, please refer to the relevant description of the network intrusion detection method based on topology data analysis, which will not be repeated here.

[0122] Finally, it should be understood that the embodiments described in this specification are intended only to illustrate the principles of the embodiments of this specification. Other variations may also fall within the scope of this specification. Therefore, by way of example and not limitation, alternative configurations of the embodiments of this specification may be considered consistent with the teachings of this specification. Accordingly, the embodiments of this specification are not limited to the embodiments explicitly described and illustrated in this specification.

Claims

1. A network intrusion detection method based on topology data analysis, characterized in that: include: Get network traffic data; Mapping the network traffic data to a feature space based on a filter function, wherein the dimension of the feature space is lower than the dimension of the network traffic data; determining a plurality of covers of the low-dimensional space; For each of the coverages, clustering the data points included in the coverage based on the similarity between any two data points using a clustering algorithm to determine a cluster cluster included in the coverage; Based on each of the clusters included in the coverage, establishing a topology map corresponding to the network traffic data; Based on the topological graph corresponding to the network traffic data and the interpolation algorithm, new minority class samples are generated to construct a class-balanced training set; Establish a network intrusion detection model; Training the network intrusion detection model based on the class-balanced training set; Real-time network traffic data is acquired, and the network intrusion type is determined based on the real-time network traffic data using the network intrusion detection model.

2. The network intrusion detection method based on topology data analysis according to claim 1, characterized in that: The optimization goal of the filtering function is: Among them, C(w h (e), w l (e)) is the optimization target of the filtering function, w h (e) is the weight of the e-th dimension in the high-dimensional space corresponding to the network traffic data, w l (e) is the weight of the e-th dimension in the feature space, and E is the total number of dimensions included in the feature space.

3. The network intrusion detection method based on topology data analysis according to claim 2, characterized in that: The weight of the e-th dimension in the high-dimensional space corresponding to the network traffic data is calculated based on the following formula: Where k is a specified number, is the data point x i Neighboring data points within a distance k range The minimum distance, ρ i is the data point x i The distance offset, σ i is the parameter used to scale the distance.

4. The network intrusion detection method based on topology data analysis according to claim 2, characterized in that: The weight of the e-th dimension in the feature space is calculated based on the following formula: In l (e)=(1+a(z i -With j ) 2b ) -1 Among them, z i is the coordinate of data point i in the feature space, z j is the coordinate of data point j in the feature space, and a and b are adjustment parameters.

5. The network intrusion detection method based on topology data analysis according to any one of claims 1 to 4, characterized in that: Determining a plurality of covers of the low-dimensional space includes: Based on the number of intervals, the fixed length of the cover, and the overlap ratio, multiple covers of the low-dimensional space are determined.

6. The network intrusion detection method based on topology data analysis according to claim 6, characterized in that: The overlap ratio is calculated based on the following formula: Where p is the overlap ratio, r is the fixed length of coverage, l is the number of intervals, and n is the total number of data points.

7. The network intrusion detection method based on topology data analysis according to any one of claims 1 to 4, characterized in that: Based on the topological graph corresponding to the network traffic data and the interpolation algorithm, new minority class samples are generated, including: Selecting a plurality of basic nodes from a minority of nodes in a topological graph corresponding to the network traffic data; For each of the basic nodes, determining the nearest neighbor node of the basic node from the same type of nodes as the basic node in the topology graph corresponding to the network traffic data; Generate new minority class samples based on the base node and the nearest neighbor nodes of the base node through an interpolation algorithm.

8. The network intrusion detection method based on topology data analysis according to any one of claims 1 to 6, characterized in that: The nearest neighbor node of the base node is determined based on the following formula: nn(v)=arg min u∈S ||h v -h u || Among them, nn(v) is the basic node x v The nearest neighbor node, S is the base node x v The set of similar nodes, h v is the base node x v In the representation of the feature space, h u is the base node x v Similar nodes x u Representation in the feature space.

9. The network intrusion detection method based on topology data analysis according to any one of claims 1 to 6, characterized in that: The interpolation algorithm is: x v′ =(1-δ)·x v +δ·x nn(v) Among them, x v′ is a new minority class sample, x nn(v) is the nearest neighbor node of the base node, and δ is a random number between 0 and 1.

10. A network intrusion detection system based on topological data analysis, characterized in that: A network intrusion detection method based on topology data analysis for executing any one of claims 1 to 9, comprising: Data acquisition module, used to obtain network traffic data; a data filtering module, configured to map the network traffic data to a feature space based on a filtering function, wherein the dimension of the feature space is lower than the dimension of the network traffic data; a cover determination module, configured to determine a plurality of covers of the low-dimensional space; a data clustering module, configured to cluster the data points included in each of the coverages based on the similarity between any two data points using a clustering algorithm, and determine a cluster cluster included in the coverage; A topology map establishing module, configured to establish a topology map corresponding to the network traffic data based on each cluster included in the coverage; A class balancing module, configured to generate new minority class samples based on a topological graph corresponding to the network traffic data and an interpolation algorithm, and construct a class-balanced training set; Model building module, used to build network intrusion detection model; A model training module, configured to train the network intrusion detection model based on the class-balanced training set; The intrusion detection module is used to obtain real-time network traffic data and determine the type of network intrusion based on the real-time network traffic data through the network intrusion detection model.