Network intrusion detection method and system based on deep learning

Through the deep learning-based network intrusion detection method, the use of spatiotemporal matrix and topological complex features, combined with DTW and evolutionary fitness analysis, and using K-means clustering and long-term short-term memory network model, the problems of poor adaptability to new attacks and insufficient utilization of nonlinear relationships in the existing technology are solved, achieving more efficient and accurate network intrusion detection.

CN120017400AActive Publication Date: 2025-05-16NANJING FORESTRY UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510283732.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-05-16
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

The prior art relies on simple network traffic feature extraction in network intrusion detection, and fails to fully mine spatio-temporal information and topological features, resulting in poor adaptability to new attacks in dynamic environments. In the recognition of mutation points and abnormal detection, traditional distance measurements or static threshold settings are used, and non-linear relationships and local change characteristics in the time series are not fully utilized.

Method used

The network intrusion detection method based on deep learning is adopted, network traffic data is collected through network packet capture tool, preprocessed and constructed a spatiotemporal matrix, converted into grayscale images, dynamic adjustment is used to dynamically adjust, topological complexes are constructed and important topological features are screened, the spatiotemporal mapping generation method is used for segmentation, DTW distance is calculated and mutation point fragments are marked, intrusion behavior is identified using evolutionary fitness formula, and K-means clustering and long-term short-term memory network model are used for prediction.

Benefits of technology

It improves the accuracy and real-time response capabilities of network intrusion detection, enhances the detection accuracy and adaptability of intrusion behavior, and can more effectively identify and predict new attack behaviors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017400A_ABST
    Figure CN120017400A_ABST
Patent Text Reader

Abstract

The invention discloses a network intrusion detection method and system based on deep learning, and relates to the technical field of network security, and the method comprises the steps: collecting network flow data through employing a network packet capturing tool, constructing a space-time matrix through employing a mapping matrix construction method, converting the space-time matrix into a gray level image, and carrying out the dynamic adjustment through employing a dynamic weight adjustment formula. Constructing a topological complex and screening important topological features, respectively performing enhancement processing on the important topological features, and performing fusion by using a weighted average method to generate a fused grayscale image; segmenting by using a space-time mapping generation method, calculating a DTW distance by using dynamic time warping and marking a mutation point fragment, and calculating the evolution fitness of the mutation point fragment by using an evolution fitness formula and identifying an intrusion behavior. The neural biological signal processing technology is combined with topological data analysis, the data processability and the feature identifiability are improved, mutation point fragments are analyzed by using an evolution fitness formula, and the detection precision and adaptability of intrusion behaviors are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technology, and in particular to a network intrusion detection method and system based on deep learning. Background Art

[0002] With the rapid development of information technology, network security issues are becoming increasingly severe, especially the prevention and detection of network intrusion behaviors has become one of the core issues in the field of network security. Traditional network intrusion detection methods mostly rely on rule matching, packet analysis and other technologies, mainly based on matching and analysis of known attack features. These methods have shown obvious shortcomings when facing unknown attacks or rapidly changing attack patterns. The rule matching method can only identify known attack patterns and has limited detection capabilities for new and variant attacks. The packet analysis-based method is prone to efficiency bottlenecks when processing large-scale network traffic, and it is difficult to effectively identify potential intrusion behaviors in real time. Researchers have gradually begun to introduce machine learning, deep learning and other technologies into the field of network intrusion detection to improve its automation and intelligence, and enhance the adaptability and accuracy of the system.

[0003] Existing deep learning applications still have some challenges to be solved in network intrusion detection. Existing technologies rely on simple network traffic feature extraction and fail to fully exploit spatiotemporal information and topological features, resulting in poor adaptability to new attacks in dynamic environments. When dealing with mutation point identification and anomaly detection, existing technologies mostly use traditional distance metrics or static threshold settings, failing to fully utilize nonlinear relationships and local change characteristics in time series. Summary of the invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides a network intrusion detection method and system based on deep learning, which solves the problem that the existing technology relies on simple network traffic feature extraction and fails to fully explore spatiotemporal information and topological features, resulting in poor adaptability to new attacks in dynamic environments. When dealing with mutation point identification and anomaly detection, the existing technology mostly uses traditional distance metrics or static threshold settings, which fails to fully utilize the nonlinear relationships and local change characteristics in time series.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0007] In a first aspect, the present invention provides a network intrusion detection method based on deep learning, which comprises:

[0008] Use network packet capture tools to collect network traffic data, pre-process the network traffic data, use the mapping matrix construction method to build a space-time matrix, and convert it into a grayscale image. Use the dynamic weight adjustment formula to dynamically adjust it, build a topological complex and screen important topological features, enhance the important topological features respectively, and use the weighted average method to fuse them to generate a fused grayscale image.

[0009] Use the space-time mapping generation method for segmentation, use dynamic time warping to calculate the DTW distance and mark the mutation point fragments, and use the evolutionary fitness formula to calculate the evolutionary fitness of the mutation point fragments and identify invasion behavior;

[0010] The K-means clustering method is used to cluster the mutation point fragments corresponding to the invasion behavior, and a long short-term memory network model is constructed to predict the invasion behavior at time t;

[0011] Stores, collects and analyzes the resulting network traffic data.

[0012] As a preferred solution of the network intrusion detection method based on deep learning described in the present invention, wherein: the network traffic data is collected by using a network packet capture tool, a space-time matrix is ​​constructed by using a mapping matrix construction method, and converted into a grayscale image, a dynamic weight adjustment formula is used for dynamic adjustment, a topological complex is constructed and important topological features are screened, important topological features are enhanced respectively, and a weighted average method is used for fusion to generate a fused grayscale image, including:

[0013] The network traffic data includes packet size, transmission delay, source port, destination port, transmission time, traffic frequency and timestamp;

[0014] The preprocessed network traffic data is sorted in chronological order, and the timestamps corresponding to the preprocessed network traffic data are mapped to the rows of the space-time matrix using a mapping matrix construction method, and the preprocessed network traffic data are mapped to the columns of the space-time matrix to construct the space-time matrix;

[0015] The spatiotemporal matrix is ​​converted into a grayscale image using a grayscale value mapping formula;

[0016] Use the sliding window method to set the neighborhood window, calculate the mean and standard deviation of the neighborhood window, and use the local relative contrast formula to calculate the local relative contrast of the (x, y) pixel. , where x and y are the position indices of the row and column in the grayscale image respectively;

[0017] Use the time window selection method to set the time window size, use the experimental experience method to set the time decay rate, and use the Gaussian time-weighted activation function to calculate the neural signal strength of the (x, y) pixel. ;

[0018] Using statistical analysis to set constant factors for enhanced strength , using random search to set the threshold of neural signals , use the dynamic weight adjustment formula to dynamically adjust the gray value of the (x, y)th pixel in the gray image;

[0019] Calculate the grayscale difference of the pixels of the adjusted grayscale image and set it as the connection strength. Use statistical analysis to set the connection threshold. Mark the connection strength greater than the connection threshold as 1, and mark the connection strength less than or equal to the connection threshold as 0 to construct an adjacency matrix.

[0020] The topological complex is constructed using the Rips complex construction method, and the topological features are extracted using the topological feature extraction method, including connected components and ring structures. The persistence calculation method is used to calculate the persistence of connected components and ring structures respectively, and the persistence screening method is used to screen the important topological features of connected components and ring structures respectively.

[0021] The grayscale image of the important topological features of the screening connected components is enhanced using the brightness enhancement method, and the grayscale image of the important topological features of the screening ring structure is enhanced using the standardization enhancement method;

[0022] The weighting coefficient is set by using the back propagation algorithm, the enhanced grayscale values ​​are fused by using the weighted average method to obtain the fused grayscale values, the fused grayscale values ​​are used to reconstruct the image, and a fused grayscale image is generated.

[0023] As a preferred solution of the network intrusion detection method based on deep learning described in the present invention, the preprocessing of network traffic data includes:

[0024] Gaussian filtering is used for denoising, the interquartile range method is used to identify and delete abnormal data, the majority filling method is used to fill missing data, and the network traffic data is normalized.

[0025] As a preferred solution of the network intrusion detection method based on deep learning described in the present invention, wherein: the segmentation is performed using the spatiotemporal mapping generation method, the DTW distance is calculated using dynamic time warping and the mutation point segments are marked, and the evolutionary fitness formula is used to calculate the evolutionary fitness of the mutation point segments and identify the intrusion behavior, including:

[0026] Use the spatiotemporal mapping generation method to segment the fused grayscale image according to the time dimension and extract spatiotemporal data fragments;

[0027] Collect network traffic data in the MongoDB database, segment the collected network traffic data in the database according to timestamps using the spatiotemporal data segmentation method, and extract database data segments;

[0028] Use the BLAST algorithm to locally compare the spatiotemporal data segments with the database data segments, calculate the similarity scores of the time segments, and sort the similarity scores from large to small;

[0029] Use dynamic time warping to calculate the DTW distance between the data segment corresponding to the maximum similarity score and the remaining data segments, use the percentile method to set the anomaly threshold, and mark the segment with a DTW distance greater than the anomaly threshold as a mutation point segment;

[0030] The Jaccard similarity coefficient formula is used to calculate the similarity score between the mutation point fragment and the database data fragment, and the evolutionary fitness formula is used to calculate the evolutionary fitness of the mutation point fragment;

[0031] The detection threshold is set using empirical rules, and the mutation point fragments whose evolutionary fitness is greater than the detection threshold are marked as invasive behaviors.

[0032] As a preferred solution of the network intrusion detection method based on deep learning described in the present invention, wherein: the use of the K-means clustering method to cluster the mutation point fragments corresponding to the intrusion behavior includes:

[0033] Normalize the mutation point fragments corresponding to the invasion behavior;

[0034] The ROC curve method is used to set the classification threshold, the elbow rule is used to set the number of clusters K, K initial cluster centers of load data are randomly selected from the normalized mutation point fragments, and the Euclidean distance formula is used to calculate the Euclidean distance from the normalized mutation point fragments to the K distance centers. The normalized mutation point fragments are assigned to the cluster center with the nearest distance. After each assignment, the K cluster centers are recalculated. When the calculated cluster center is less than the classification threshold, the iteration is stopped to obtain the clustered K intrusion types.

[0035] As a preferred solution of the network intrusion detection method based on deep learning of the present invention, the method of constructing a long short-term memory network model to predict the intrusion behavior at time t includes:

[0036] Collect and preprocess historical network traffic data with labels to generate a training set;

[0037] Construct a long short-term memory network model, including input layer, LSTM layer, Dense layer and output layer;

[0038] Set the input layer format to network traffic data;

[0039] Use the training set to train the LSTM network model, and use the loss function and Adam optimizer to iteratively optimize the model parameters;

[0040] The preprocessed network traffic data is fed into the trained long short-term memory network model to obtain the intrusion behavior at time t;

[0041] The intrusion behaviors include SQL injection, phishing attacks, brute force cracking, and lateral movement.

[0042] As a preferred solution of the network intrusion detection method based on deep learning of the present invention, the network traffic data generated by the storage, collection and analysis includes:

[0043] The collected network traffic data and all data generated by the analysis are stored in the central database, and security access measures are set up. The central database will back up the stored data to the cloud and regularly perform integrity checks on the stored data and backup data. After the test is completed, the integrity test record is generated and stored synchronously in the central database.

[0044] In a second aspect, the present invention provides a network intrusion detection system based on deep learning, comprising:

[0045] The collection enhancement module is used to collect network traffic data using a network packet capture tool, pre-process the network traffic data, construct a spatiotemporal matrix using a mapping matrix construction method, and convert it into a grayscale image. Dynamic adjustment is performed using a dynamic weight adjustment formula, a topological complex is constructed, and important topological features are screened. The important topological features are enhanced respectively, and the weighted average method is used for fusion to generate a fused grayscale image.

[0046] The recognition module is used to segment using the spatiotemporal mapping generation method, calculate the DTW distance using dynamic time warping and mark the mutation point segments, calculate the evolutionary fitness of the mutation point segments using the evolutionary fitness formula and identify the invasion behavior;

[0047] The clustering prediction module is used to cluster the mutation point fragments corresponding to the invasion behavior using the K-means clustering method, and to construct a long short-term memory network model to predict the invasion behavior at time t;

[0048] The storage module is used to store the network traffic data generated by collection and analysis.

[0049] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the deep learning-based network intrusion detection method as described in the first aspect of the present invention is implemented.

[0050] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the deep learning-based network intrusion detection method as described in the first aspect of the present invention.

[0051] The beneficial effects of the present invention are as follows: the present invention collects network traffic data by using a network packet capture tool, pre-processes the network traffic data, constructs a space-time matrix by using a mapping matrix construction method, converts it into a grayscale image, uses a dynamic weight adjustment formula for dynamic adjustment, constructs a topological complex and screens important topological features, respectively enhances the important topological features, fuses them by using a weighted average method, and generates a fused grayscale image; uses a space-time mapping generation method for segmentation, uses dynamic time warping to calculate the DTW distance and mark mutation point fragments, uses an evolutionary fitness formula to calculate the evolutionary fitness of the mutation point fragments and identifies intrusion behavior; uses a K-means clustering method to cluster the mutation point fragments corresponding to the intrusion behavior, and constructs a long short-term memory network model to predict the intrusion behavior at time t; improves the detection accuracy and real-time response capability, and enhances the detection accuracy and adaptability of intrusion behavior. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0053] Figure 1 This is a flow chart of the network intrusion detection method based on deep learning in Example 1.

[0054] Figure 2 This is a structural diagram of the deep learning-based network intrusion detection system in Example 1. DETAILED DESCRIPTION

[0055] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.

[0056] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0057] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.

[0058] Example 1, reference Figure 1 and Figure 2 , which is the first embodiment of the present invention, and provides a network intrusion detection method based on deep learning, comprising the following steps:

[0059] S1. Use network packet capture tools to collect network traffic data, pre-process the network traffic data, use the mapping matrix construction method to construct a space-time matrix, and convert it into a grayscale image. Use the dynamic weight adjustment formula to dynamically adjust it, build a topological complex and screen important topological features, enhance the important topological features respectively, use the weighted average method to fuse them, and generate a fused grayscale image.

[0060] Specifically, network traffic data is collected using a network packet capture tool, a space-time matrix is ​​constructed using a mapping matrix construction method, and converted into a grayscale image. Dynamic adjustment is performed using a dynamic weight adjustment formula, a topological complex is constructed, and important topological features are screened. Important topological features are enhanced respectively, and a weighted average method is used for fusion to generate a fused grayscale image, including:

[0061] The network traffic data includes packet size, transmission delay, source port, destination port, transmission time, traffic frequency and timestamp;

[0062] The preprocessed network traffic data is sorted in chronological order, and the timestamps corresponding to the preprocessed network traffic data are mapped to the rows of the space-time matrix using a mapping matrix construction method, and the preprocessed network traffic data are mapped to the columns of the space-time matrix to construct the space-time matrix;

[0063] The spatiotemporal matrix is ​​converted into a grayscale image using a grayscale value mapping formula;

[0064] Use the sliding window method to set the neighborhood window, calculate the mean and standard deviation of the neighborhood window, and use the local relative contrast formula to calculate the local relative contrast of the (x, y) pixel. , where x and y are the position indexes of the row and column in the grayscale image respectively, and the formula is:

[0065] ,

[0066] in is the gray value of the (x,y)th pixel in the gray image, is the mean value of the (x,y)th pixel in the neighborhood window, is the standard deviation of the (x,y)th pixel in the neighborhood window, is a very small constant to avoid division by zero errors;

[0067] Use the time window selection method to set the time window size, use the experimental experience method to set the time decay rate, and use the Gaussian time-weighted activation function to calculate the neural signal strength of the (x, y) pixel. , the formula is:

[0068] ,

[0069] Where T is the time window size, is the time decay rate at time point t;

[0070] Using statistical analysis to set constant factors for enhanced strength , using random search to set the threshold of neural signals , use the dynamic weight adjustment formula to dynamically adjust the gray value of the (x, y)th pixel in the gray image. The formula is:

[0071] ,

[0072] in is the gray value of the (x, y)th pixel in the adjusted gray image;

[0073] Calculate the grayscale difference of the pixels of the adjusted grayscale image and set it as the connection strength. Use statistical analysis to set the connection threshold. Mark the connection strength greater than the connection threshold as 1, and mark the connection strength less than or equal to the connection threshold as 0. Construct an adjacency matrix to represent the topological structure of the pixels in the image.

[0074] The topological complex is constructed using the Rips complex construction method, and the topological features are extracted using the topological feature extraction method, including connected components and ring structures. The persistence calculation method is used to calculate the persistence of connected components and ring structures respectively, and the persistence screening method is used to screen the important topological features of connected components and ring structures respectively.

[0075] The brightness enhancement method is used to enhance the grayscale image of the important topological features of the screening connected components, and the standardized enhancement method is used to enhance the grayscale image of the important topological features of the screening ring structure. The formula is:

[0076] ,

[0077] ,

[0078] in is the gray value of the (x, y)th pixel after connected component enhancement, is the gray value of the (x,y)th pixel after ring structure enhancement, is the persistence of the connected component, is the maximum value of the persistence of connected components, is the mean of the grayscale values ​​of the ring structure, indicating the average value of the grayscale values ​​of all ring structures, is the standard deviation of the grayscale values ​​of the ring structure, which represents the standard deviation of the grayscale values ​​of all ring structures;

[0079] The weight coefficient is set by using the back propagation algorithm, and the enhanced grayscale values ​​are fused by using the weighted average method to obtain the fused grayscale values. The fused grayscale values ​​are used to reconstruct the image to generate the fused grayscale image. The formula is:

[0080] ,

[0081] in is the gray value of the (x,y)th pixel of the fused image, is the weighting coefficient, is the grayscale value of the (x, y)th pixel after all enhancements, including the grayscale value of the (x, y)th pixel after enhancement by the connected component and the ring structure.

[0082] The present invention uses a mapping matrix construction method to construct a spatiotemporal matrix, and improves the processability of data and the identifiability of features through grayscale image conversion technology. The construction of a dynamic weight adjustment formula and a topological complex effectively enhances the ability to extract important features, thereby improving the accuracy and real-time performance of intrusion detection. The topological complex construction method is used to extract the topological features of network traffic, and the persistence calculation method is used to screen out important topological features. The quality of the image is further improved through enhanced processing, so that the deep learning model can more accurately identify potential network intrusion behaviors.

[0083] By using network packet capture tools to collect information such as the size, transmission delay, source port, and destination port of data packets, we can fully understand the communication behavior in the network, sort the data in chronological order, and ensure the temporal nature of the data. By converting the space-time matrix into a grayscale image, we can simplify the data representation, making subsequent image processing and deep learning algorithms more efficient, reducing the complexity of the data, and highlighting the key features in the network traffic. By combining the Gaussian time-weighted activation function to calculate the neural signal strength, we can simulate the neural response to changes in network traffic and effectively identify sudden abnormal points that may be intrusion behaviors. The construction of topological complexes and the extraction of topological features help to extract more representative features from complex data. The persistence calculation method is used to screen out features that are meaningful to intrusion detection, effectively reducing the impact of noise on the detection results and improving the detection accuracy. The enhanced grayscale images are fused by the weighted average method to generate a fused image, which improves the overall quality of the image features, enabling the intrusion detection algorithm to more accurately identify potential security threats and improve the processing efficiency of network traffic data and the accuracy of intrusion detection.

[0084] Furthermore, the network traffic data is preprocessed, including:

[0085] Gaussian filtering is used for denoising, the interquartile range method is used to identify and delete abnormal data, the majority filling method is used to fill missing data, and the network traffic data is normalized.

[0086] Gaussian filtering reduces these interferences by smoothing the data, providing a more stable data foundation for model training and analysis. By using the interquartile range method, these outliers can be automatically detected and eliminated during the data analysis process to avoid abnormal data interfering with model training and prediction, improve data stability and model robustness. After filling in the missing data, the integrity of the data set is guaranteed, allowing the model to effectively learn the characteristics of all data and avoid the imbalance problem caused by missing data. Normalization processing prevents features with a large numerical range from dominating the model learning process, thereby improving the training speed and accuracy of the model. The data preprocessing method provides a reliable data foundation for subsequent intrusion detection, network traffic analysis and other tasks, and improves the prediction accuracy and computational efficiency of the model.

[0087] S2, use the space-time mapping generation method to segment, use dynamic time warping to calculate the DTW distance and mark the mutation point fragments, use the evolutionary fitness formula to calculate the evolutionary fitness of the mutation point fragments and identify the invasion behavior;

[0088] Specifically, the space-time mapping generation method is used for segmentation, the DTW distance is calculated using dynamic time warping and the mutation point segments are marked, and the evolutionary fitness formula is used to calculate the evolutionary fitness of the mutation point segments and identify the invasion behavior, including:

[0089] The fused grayscale image is segmented according to the time dimension using the spatiotemporal mapping generation method, and spatiotemporal data segments are extracted from each row;

[0090] Collect network traffic data from the MongoDB database, segment the collected network traffic data in the database according to timestamps using the spatiotemporal data segmentation method, and extract database data segments from each row;

[0091] Use the BLAST algorithm to locally compare the spatiotemporal data segments with the database data segments, calculate the similarity scores of the time segments, and sort the similarity scores from large to small;

[0092] Use dynamic time warping to calculate the DTW distance between the data segment corresponding to the maximum similarity score and the remaining data segments, use the percentile method to set the anomaly threshold, and mark the segment with a DTW distance greater than the anomaly threshold as a mutation point segment;

[0093] The Jaccard similarity coefficient formula is used to calculate the similarity score between the mutation point fragment and the database data fragment, and the evolutionary fitness formula is used to calculate the evolutionary fitness of the mutation point fragment. The formula is:

[0094] ,

[0095] in is the evolutionary fitness of the mutation point fragment w, is the similarity score between the mutation point segment w and the database data segment z;

[0096] The detection threshold is set using empirical rules, and the mutation point fragments whose evolutionary fitness is greater than the detection threshold are marked as invasive behaviors.

[0097] The present invention solves the deficiencies of the prior art in mutation point identification and dynamic intrusion detection by introducing the division of spatiotemporal data segments and dynamic time warping calculation, combined with evolutionary fitness analysis. The network traffic data is segmented by introducing the spatiotemporal mapping generation method, and the similarity of the data segments is calculated by combining the dynamic time warping (DTW) algorithm, so as to effectively identify the mutation point segments, thereby improving the detection accuracy of intrusion behavior. The mutation point segments are analyzed by using the evolutionary fitness formula, and the detection threshold can be automatically evaluated and adjusted, further improving the adaptability and accuracy of the model. It not only has obvious advantages in capturing spatiotemporal dynamic characteristics, but also can maintain efficient computing performance and low false alarm rate when facing large-scale network traffic data, significantly improving the effect of network intrusion detection.

[0098] The grayscale image is segmented according to the time dimension using the spatiotemporal mapping generation method, which can effectively extract the time features of network traffic data and structure it into easy-to-process spatiotemporal data segments. The BLAST algorithm is used to calculate the similarity score, which can effectively identify the similarity between historical data and current network traffic, thereby providing a basis for intrusion detection. By calculating the DTW distance between the maximum similarity segment and other segments, the mutation point segments that are significantly different from the normal behavior pattern can be accurately marked. The introduction of DTW can make up for the shortcomings of traditional time series analysis methods, especially when dealing with network traffic with inconsistent signal rate changes, it can achieve more accurate similarity calculation. Jaccard similarity can quickly evaluate the similarity between data segments, especially when facing massive network traffic data, it can effectively identify which data segments have abnormal patterns. The evolutionary fitness formula can adaptively adjust the detection strategy by introducing a dynamic fitness evaluation mechanism, gradually identify the most potentially threatening behavior segments, and identify new intrusion behaviors that are difficult to detect with traditional methods. It has strong universality and adaptability.

[0099] S3, cluster the mutation point fragments corresponding to the invasion behavior using the K-means clustering method, and construct a long short-term memory network model to predict the invasion behavior at time t;

[0100] Specifically, the K-means clustering method is used to cluster the mutation point fragments corresponding to the invasion behavior, including:

[0101] Normalize the mutation point fragments corresponding to the invasion behavior;

[0102] The ROC curve method is used to set the classification threshold, the elbow rule is used to set the number of clusters K, K initial cluster centers of load data are randomly selected from the normalized mutation point fragments, and the Euclidean distance formula is used to calculate the Euclidean distance from the normalized mutation point fragments to the K distance centers. The normalized mutation point fragments are assigned to the cluster center with the nearest distance. After each assignment, the K cluster centers are recalculated. When the calculated cluster center is less than the classification threshold, the iteration is stopped to obtain the clustered K intrusion types.

[0103] Normalization can eliminate the influence caused by the difference in feature scales, so that each feature has the same weight in the clustering process, and can avoid some features being over- or ignored in the clustering process due to large or small values, thereby improving the accuracy and stability of the K-means clustering algorithm. The ROC curve method helps to select an optimal threshold by comparing the true positive rate and false positive rate under different classification thresholds. The appropriate K value ensures the stability and accuracy of the clustering process and avoids the influence of too many or too few cluster centers. Selecting a suitable K value makes the clustering result more consistent with the distribution of actual data, improving the classification accuracy of intrusion behavior. Clustering allocation makes similar intrusion behaviors clustered in the same cluster, enhancing the system's ability to identify different intrusion types. By setting the classification threshold, it can ensure that the algorithm converges after an appropriate number of iterations, avoiding the waste of calculation caused by excessive iterations and improving the clustering efficiency. The clustering method can identify potential attack patterns in large-scale network traffic and classify them into different types of intrusion behaviors. It can accurately identify multiple types of intrusion behaviors, which helps to take targeted defense measures and improve the practicality and effectiveness of intrusion detection.

[0104] Furthermore, a long short-term memory network model is constructed to predict the intrusion behavior at time t, including:

[0105] Collect and preprocess historical network traffic data with labels to generate a training set;

[0106] Construct a long short-term memory network model, including input layer, LSTM layer, Dense layer and output layer;

[0107] Set the input layer format to network traffic data;

[0108] Use the training set to train the LSTM network model, and use the loss function and Adam optimizer to iteratively optimize the model parameters;

[0109] The preprocessed network traffic data is fed into the trained long short-term memory network model to obtain the intrusion behavior at time t;

[0110] The intrusion behaviors include SQL injection, phishing attacks, brute force cracking, and lateral movement.

[0111] By collecting and organizing historical data, the accuracy and reliability of training data can be ensured. Data preprocessing can improve the quality of network traffic data and ensure the learning effect of the model during the training process. The use of the LSTM layer can help the model learn long-term dependencies, that is, capture historical patterns in network traffic and use them to predict future intrusion behaviors. By predicting intrusion behaviors, the system can identify potential attack patterns in advance and take corresponding defense measures, thereby effectively improving network security protection capabilities. The output layer of the LSTM network converts the prediction results of network traffic data into intrusion behavior categories. The accurate identification of these categories helps the network security system respond in a timely manner. Through continuous training and optimization, the LSTM network can improve its generalization ability in different network environments and adapt to various complex intrusion detection scenarios.

[0112] S4, store, collect and analyze the generated network traffic data;

[0113] Specifically, the network traffic data generated by storage, collection and analysis includes:

[0114] The collected network traffic data and all data generated by the analysis are stored in the central database, and security access measures are set. The central database will back up the stored data to the cloud, and regularly perform integrity checks on the stored data and backup data. After the test is completed, the integrity test record is generated and stored synchronously in the central database;

[0115] All data generated by the analysis include clustering results and predicted intrusion behavior at time t.

[0116] During the network security monitoring process, network traffic and prediction results can be obtained in real time, helping security analysts respond to potential threats and attacks in a timely manner. Security access measures help protect sensitive information stored in the central database, avoid data leakage or abuse, and enhance the security of the system, especially when processing sensitive data such as network traffic and intrusion prediction information. Cloud backup is highly reliable and can restore data in the event of system failure, natural disasters or other unforeseen situations. Data integrity detection can be achieved by calculating hash values, checksums, etc. to ensure the correctness and consistency of data. Cloud backup and integrity detection provide double protection for data. In the face of data loss, damage or tampering, it can be quickly restored and ensure data integrity. Data security and integrity protection also improve the credibility of intrusion detection results and avoid misjudgments due to data problems.

[0117] This embodiment also provides a network intrusion detection system based on deep learning, including:

[0118] The collection enhancement module is used to collect network traffic data using a network packet capture tool, pre-process the network traffic data, construct a spatiotemporal matrix using a mapping matrix construction method, and convert it into a grayscale image. Dynamic adjustment is performed using a dynamic weight adjustment formula, a topological complex is constructed, and important topological features are screened. The important topological features are enhanced respectively, and the weighted average method is used for fusion to generate a fused grayscale image.

[0119] The recognition module is used to segment using the spatiotemporal mapping generation method, calculate the DTW distance using dynamic time warping and mark the mutation point segments, calculate the evolutionary fitness of the mutation point segments using the evolutionary fitness formula and identify the invasion behavior;

[0120] The clustering prediction module is used to cluster the mutation point fragments corresponding to the invasion behavior using the K-means clustering method, and to construct a long short-term memory network model to predict the invasion behavior at time t;

[0121] The storage module is used to store the network traffic data generated by collection and analysis.

[0122] This embodiment also provides a computer device, which is suitable for the network intrusion detection method based on deep learning, including: a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute computer executable instructions to implement the network intrusion detection method based on deep learning proposed in the above embodiment.

[0123] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered on the display screen, or a key, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.

[0124] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, the network intrusion detection method based on deep learning as proposed in the above embodiment is implemented; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, referred to as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, referred to as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, referred to as EPROM), programmable read-only memory (Programmable Red-Only Memory, referred to as PROM), read-only memory (Read-Only Memory, referred to as ROM), magnetic storage, flash memory, disk or optical disk.

[0125] In summary, the present invention collects network traffic data by using a network packet capture tool, pre-processes the network traffic data, constructs a space-time matrix by using a mapping matrix construction method, and converts it into a grayscale image, uses a dynamic weight adjustment formula for dynamic adjustment, constructs a topological complex and screens important topological features, enhances the important topological features respectively, and fuses them using a weighted average method to generate a fused grayscale image; uses a space-time mapping generation method for segmentation, uses dynamic time warping to calculate the DTW distance and mark mutation point fragments, uses an evolutionary fitness formula to calculate the evolutionary fitness of the mutation point fragments and identify intrusion behaviors; uses a K-means clustering method to cluster the mutation point fragments corresponding to the intrusion behavior, and constructs a long short-term memory network model to predict the intrusion behavior at time t; improves the detection accuracy and real-time response capability, and enhances the detection accuracy and adaptability of intrusion behaviors.

[0126] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A network intrusion detection method based on deep learning, characterized in that: include, Use network packet capture tools to collect network traffic data, pre-process the network traffic data, use the mapping matrix construction method to build a space-time matrix, and convert it into a grayscale image. Use the dynamic weight adjustment formula to dynamically adjust it, build a topological complex and screen important topological features, enhance the important topological features respectively, and use the weighted average method to fuse them to generate a fused grayscale image. Use the space-time mapping generation method for segmentation, use dynamic time warping to calculate the DTW distance and mark the mutation point fragments, and use the evolutionary fitness formula to calculate the evolutionary fitness of the mutation point fragments and identify invasion behavior; The K-means clustering method is used to cluster the mutation point fragments corresponding to the invasion behavior, and a long short-term memory network model is constructed to predict the invasion behavior at time t; Stores, collects and analyzes the resulting network traffic data.

2. The network intrusion detection method based on deep learning according to claim 1, characterized in that: The network packet capture tool is used to collect network traffic data, a mapping matrix construction method is used to construct a space-time matrix, and the matrix is ​​converted into a grayscale image, a dynamic weight adjustment formula is used for dynamic adjustment, a topological complex is constructed and important topological features are screened, important topological features are enhanced respectively, and a weighted average method is used for fusion to generate a fused grayscale image, including: The network traffic data includes packet size, transmission delay, source port, destination port, transmission time, traffic frequency and timestamp; The preprocessed network traffic data is sorted in chronological order, and the timestamps corresponding to the preprocessed network traffic data are mapped to the rows of the space-time matrix using a mapping matrix construction method, and the preprocessed network traffic data are mapped to the columns of the space-time matrix to construct the space-time matrix; The spatiotemporal matrix is ​​converted into a grayscale image using a grayscale value mapping formula; Use the sliding window method to set the neighborhood window, calculate the mean and standard deviation of the neighborhood window, and use the local relative contrast formula to calculate the local relative contrast of the (x, y) pixel. , where x and y are the position indices of the row and column in the grayscale image respectively; Use the time window selection method to set the time window size, use the experimental experience method to set the time decay rate, and use the Gaussian time-weighted activation function to calculate the neural signal strength of the (x, y) pixel. ; Using statistical analysis to set constant factors for enhanced strength , using random search to set the threshold of neural signals , use the dynamic weight adjustment formula to dynamically adjust the gray value of the (x, y)th pixel in the gray image; Calculate the grayscale difference of the pixels of the adjusted grayscale image and set it as the connection strength. Use statistical analysis to set the connection threshold. Mark the connection strength greater than the connection threshold as 1, and mark the connection strength less than or equal to the connection threshold as 0 to construct an adjacency matrix. The topological complex is constructed using the Rips complex construction method, and the topological features are extracted using the topological feature extraction method, including connected components and ring structures. The persistence calculation method is used to calculate the persistence of connected components and ring structures respectively, and the persistence screening method is used to screen the important topological features of connected components and ring structures respectively. The grayscale image of the important topological features of the screening connected components is enhanced using the brightness enhancement method, and the grayscale image of the important topological features of the screening ring structure is enhanced using the standardization enhancement method; The weighting coefficient is set by using the back propagation algorithm, the enhanced grayscale values ​​are fused by using the weighted average method to obtain the fused grayscale values, the fused grayscale values ​​are used to reconstruct the image, and a fused grayscale image is generated.

3. The network intrusion detection method based on deep learning as claimed in claim 2, characterized in that: The method of using the space-time mapping generation method for segmentation, using dynamic time warping to calculate the DTW distance and mark the mutation point segments, and using the evolutionary fitness formula to calculate the evolutionary fitness of the mutation point segments and identify the invasion behavior includes: Use the spatiotemporal mapping generation method to segment the fused grayscale image according to the time dimension and extract spatiotemporal data fragments; Collect network traffic data in the MongoDB database, segment the collected network traffic data in the database according to timestamps using the spatiotemporal data segmentation method, and extract database data segments; Use the BLAST algorithm to locally compare the spatiotemporal data segments with the database data segments, calculate the similarity scores of the time segments, and sort the similarity scores from large to small; Use dynamic time warping to calculate the DTW distance between the data segment corresponding to the maximum similarity score and the remaining data segments, use the percentile method to set the anomaly threshold, and mark the segment with a DTW distance greater than the anomaly threshold as a mutation point segment; The Jaccard similarity coefficient formula is used to calculate the similarity score between the mutation point fragment and the database data fragment, and the evolutionary fitness formula is used to calculate the evolutionary fitness of the mutation point fragment; The detection threshold is set using empirical rules, and the mutation point fragments whose evolutionary fitness is greater than the detection threshold are marked as invasive behaviors.

4. The network intrusion detection method based on deep learning as claimed in claim 3, characterized in that: The method of clustering the mutation point fragments corresponding to the invasion behavior using the K-means clustering method includes: Normalize the mutation point fragments corresponding to the invasion behavior; The ROC curve method is used to set the classification threshold, the elbow rule is used to set the number of clusters K, K initial cluster centers of load data are randomly selected from the normalized mutation point fragments, and the Euclidean distance formula is used to calculate the Euclidean distance from the normalized mutation point fragments to the K distance centers. The normalized mutation point fragments are assigned to the cluster center with the nearest distance. After each assignment, the K cluster centers are recalculated. When the calculated cluster center is less than the classification threshold, the iteration is stopped to obtain the clustered K intrusion types.

5. The network intrusion detection method based on deep learning as claimed in claim 4, characterized in that: The constructing of the long short-term memory network model to predict the intrusion behavior at time t includes: Collect and preprocess historical network traffic data with labels to generate a training set; Construct a long short-term memory network model, including input layer, LSTM layer, Dense layer and output layer; Set the input layer format to network traffic data; Use the training set to train the LSTM network model, and use the loss function and Adam optimizer to iteratively optimize the model parameters; The preprocessed network traffic data is fed into the trained long short-term memory network model to obtain the intrusion behavior at time t; The intrusion behaviors include SQL injection, phishing attacks, brute force cracking, and lateral movement.

6. The network intrusion detection method based on deep learning as claimed in claim 5, characterized in that: The preprocessing of network traffic data includes: Gaussian filtering is used for denoising, the interquartile range method is used to identify and delete abnormal data, the majority filling method is used to fill missing data, and the network traffic data is normalized.

7. The network intrusion detection method based on deep learning according to claim 6, characterized in that: The network traffic data generated by the storage, collection and analysis includes: The collected network traffic data and all data generated by the analysis are stored in the central database, and security access measures are set up. The central database will back up the stored data to the cloud and regularly perform integrity checks on the stored data and backup data. After the test is completed, the integrity test record is generated and stored synchronously in the central database.

8. A network intrusion detection system based on deep learning, based on the network intrusion detection method based on deep learning according to any one of claims 1 to 7, characterized in that: include, The collection enhancement module is used to collect network traffic data using a network packet capture tool, pre-process the network traffic data, construct a spatiotemporal matrix using a mapping matrix construction method, and convert it into a grayscale image. Dynamic adjustment is performed using a dynamic weight adjustment formula, a topological complex is constructed, and important topological features are screened. The important topological features are enhanced respectively, and the weighted average method is used for fusion to generate a fused grayscale image. The recognition module is used to segment using the spatiotemporal mapping generation method, calculate the DTW distance using dynamic time warping and mark the mutation point segments, calculate the evolutionary fitness of the mutation point segments using the evolutionary fitness formula and identify the invasion behavior; The clustering prediction module is used to cluster the mutation point fragments corresponding to the invasion behavior using the K-means clustering method, and to construct a long short-term memory network model to predict the invasion behavior at time t; The storage module is used to store the network traffic data generated by collection and analysis.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the deep learning-based network intrusion detection method described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the deep learning-based network intrusion detection method described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Intrusion detection system and method based on intelligent network

    CN118413406A

  • Perimeter security intrusion signal identification method in migration scene based on distributed optical fiber sensing

    CN118781723A

  • Synchronizing security systems in a geographic network

    US12063458B1