Multi-model fusion efficient intrusion detection method and system
Through the feature screening and attention mechanism combined with multi-model fusion method, the spatiotemporal characteristics of network traffic are extracted and fused, and feature weights are allocated through the multi-head attention mechanism, the problems of data distribution imbalance and computational overhead in the existing network intrusion detection methods are solved, and efficient and accurate intrusion detection is achieved.
Patent Information
- Application Number
- CN202510136418.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-07
AI Technical Summary
Existing network intrusion detection methods are difficult to achieve efficient and accurate intrusion detection when facing problems such as unbalanced data distribution, large overhead for testing model training and unstable accuracy.
Through feature screening and attention mechanisms, the model focuses on the key features of the attack traffic, combines multi-model fusion methods, including data sample cleaning, sample generation, spatiotemporal feature extraction and feature weight allocation, use TCN and BiLSTM to extract spatiotemporal features, and use multi-head attention mechanism to distribute feature weights, and finally output classification results at the full connection layer.
It effectively alleviates the problem of data distribution imbalance, reduces computing overhead, improves the stability and generalization capabilities of the model, and ensures high efficiency and accuracy in complex traffic environments.
Smart Images

Figure CN119995978A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a multi-model fusion efficient intrusion detection method and system. Background Art
[0002] With the rapid development of information technology and the widespread application of Internet of Things technology, network security issues are still unavoidable; the forms of network attacks are constantly evolving, attackers use more complex and covert means to launch attacks, the complexity of network behavior information is increasing, and traditional detection technology faces unprecedented unstable challenges. For the detection of network attacks, analysis based on network traffic has become one of the mainstream methods. Network traffic detection can help security systems identify potential threats and attack behaviors in real time. However, this method always faces problems such as unbalanced data distribution, high detection model training overhead, and unstable accuracy when the actual network environment continues to change.
[0003] In order to solve the detection obstacles that still exist in network-based intrusion detection, feature screening and attention mechanisms can be used to focus the model's attention on the key features of attack traffic, enhance the recognition ability of minority classes (attacks), and thus alleviate the impact of data imbalance. Secondly, by improving the efficiency of traffic data feature extraction, the reliance on long time series and a large number of redundant features can be reduced, while parallel training and feature simplification techniques can be used to reduce computational overhead. On the basis of solving the problem of imbalanced distribution, the processing and extraction efficiency of stable features, the stability and generalization ability of the model will also be enhanced simultaneously.
[0004] In addition, considering that the multi-dimensional features and their inherent correlation in network traffic data are the key to continuously improving the accuracy and robustness of detection, a correlation analysis mechanism between features is constructed to dynamically capture the interactive relationship between spatiotemporal features under different traffic patterns, and combined with an adaptive weight allocation strategy, the importance of feature expression can be adjusted in real time during the classification process, so that the model can focus more accurately on key feature areas and abnormal patterns. At the same time, the dynamic feature adjustment mechanism effectively alleviates the misjudgment problem caused by feature imbalance or feature redundancy in traditional methods, ensuring the efficiency and generalization ability of the detection system in complex traffic environments.
[0005] Publication number CN109462583A, the invention name is "Network intrusion detection method and network intrusion detection device based on deep learning". This patent integrates CNN+GRU and attention mechanism technology as an intrusion detection method. Taking into account the rapid extraction effect of the GRU model on time features, it does not combine the actual time granularity and time step to reasonably explore the needs of spatiotemporal feature extraction, and does not consider the impact of the distribution of the data set on the detection accuracy. On the basis of enhanced data cleaning, this patent clarifies seconds and minutes as the basic time granularity, designs a multi-head attention mechanism around the time granularity of seconds and minutes, and can fully explore the weight distribution in the mixed data of seconds and minutes. And add a sample generation stage to solve the problem of distribution imbalance in the future trend of distributed networks and stabilize the detection accuracy.
[0006] To sum up, based on the data distribution problems and overhead problems faced by network intrusion detection and the importance of in-depth feature analysis, compared with the insufficient exploration of the granularity of seconds in existing solutions and the lack of attention to the problem of imbalanced distribution, this application proposes an efficient intrusion detection method and system with multi-model fusion to further adapt to the stable detection performance of network attacks in various environments. Summary of the invention
[0007] In view of the deficiencies in the prior art, the present invention discloses an efficient intrusion detection method and system with multi-model fusion to solve the problems raised in the above background technology.
[0008] To achieve the above object, the present invention provides the following technical solution: an efficient intrusion detection method based on multi-model fusion, comprising the following steps:
[0009] S1. Data sample cleaning, including removing all inf and nan outliers in the sample data, using TSODE to screen for time features, using custom regular specifications to screen for fixed format features, and using z-score to screen for statistical features;
[0010] S2, sample generation, including using the sliding window method to confirm the minority class that meets the properties of the decision boundary, using BorderLine-SMOTE to generate samples, and using ENN to verify the generated samples;
[0011] S3, spatiotemporal feature extraction, including using TCN to extract spatial features of sample data and using BiLSTM to extract flow features of sample data;
[0012] S4, spatiotemporal feature fusion, including fusing and normalizing temporal and spatial features, including feature dimension alignment, feature splicing, and normalization processing;
[0013] S5, feature weight assignment, including generating an initial weight matrix by using statistical sample means of time windows with different time step lengths, and calculating a QKV matrix based on the initial weight matrix;
[0014] S6. Based on the attention mechanism results, the classification results are output in the fully connected layer.
[0015] Preferably, in step S1, based on the statistical analysis of the traffic data, the traffic data features are classified into four categories: time features, fixed format features, statistical features, and other features; outlier removal, TSODE method, and Z-score method are used to process each type of feature to complete data sample cleaning, which specifically includes the following steps:
[0016] Step 1.1, scan the samples for outlier samples with characteristic values of inf and nan, and remove them directly;
[0017] Step 1.2: Based on the outlier cleaning in step 1.1, the TSODE method is used to perform two stages of single feature deviation evaluation and multi-feature joint evaluation for the time feature column, and the abnormal samples whose time feature values deviate from the normal model are eliminated;
[0018] Step 1.3: For fixed format features, they usually have a fixed format or specification. Therefore, based on regular grammar, combined with relevant specifications and protocol formats of network traffic, custom feature rules are set; for samples that fail to match, the custom rules will locate the location of the sample, and a manual review method is used to determine whether the fixed format feature value is "the format is incorrect, but the content is feasible". The samples with real erroneous content are eliminated, otherwise the correct format is restored;
[0019] Step 1.4: The measure derived from the original data or primary features is the statistical feature. Since it usually conforms to the approximate normal distribution, the z-score method is proposed to remove samples with feature value evaluation scores exceeding 3. The z-score converts the original feature value into a dimensionless standard score, which no longer depends on a specific numerical range, making the anomaly detection of different features comparable. For the statistical feature x j :
[0020]
[0021] Among them, μ j ,σ j is the mean and standard deviation of the current feature at all time steps; samples with scores out of range will be directly eliminated;
[0022] Through four outlier processing steps, the quality of network traffic data can be effectively improved, and the interference of abnormal noise on model training can be reduced, thereby improving the robustness and generalization ability of the model; these processing measures can ensure that time features reflect real dynamic behaviors, fixed format features such as IP addresses and port numbers conform to the standard format, and statistical features such as traffic rate and packet size are more evenly distributed.
[0023] Preferably, step 1.2 specifically includes the following steps:
[0024] Step 1.2.1, identify whether there are outliers in the time features that deviate from the normal pattern one by one; use the sliding window method to j ={x 1j ,x 2j ,…,x kj}, calculate the rolling mean μ of the feature in k time steps t and standard deviation α t :
[0025]
[0026] Where k is the window size; after the results of the single variable learning curve method and grid search, it is determined that the value range of k is [5,11]. In the present invention, setting k to 7 can effectively take into account the time consumption of outlier discovery and calculation; and if the current sample meets:
[0027] |x tj -μ t |≤λ·α t
[0028] If the current sample time feature conforms to the normal mode of the data set, then the mean μ of the current time window is t Replace x tj , so as to ensure that other information of the sample is not lost; where λ = 2.5;
[0029] Step 1.2.2, establish joint feature screening of multiple time features; take the time step as the basic unit, and calculate the joint outlier score of d time features on the sample corresponding to the current time step:
[0030]
[0031] Where α = 2, and when the feature score is greater than the threshold T, it is considered an outlier and the sample is removed; considering the multi-scale traffic data at the second and minute levels that may appear in the data set, and the joint score is approximately a normal distribution or a right-skewed distribution, the threshold T is set:
[0032] T=μ+β·σ
[0033] Where μ, σ are the mean and standard deviation of all time features of the current sample; β takes the value of 2.5 when the proportion of minute-level data exceeds two-thirds, otherwise it takes the value of 1.5; and for Score t >T, it is considered that the joint distribution of the current sample deviates and the sample is directly eliminated.
[0034] Preferably, in step S2, after determining that there is a problem with the data set to be detected, the BorderLine-SMOTE combined with the ENN method is used to solve the problem, which specifically includes the following steps:
[0035] The first step is to confirm the boundary samples, for the minority class samples x i , use KNN to calculate the minority class ratio, where the minority class ratio refers to the proportion of samples that belong to the same category as the sample in the first k samples of feature similarity:
[0036]
[0037] The k neighbors are selected by calculating the formula:
[0038]
[0039] Where D is the feature dimension, Represents sample x i The d eigenvalues are set, and k is 5; the boundary thresholds T1 = 0.5 and T2 = 0.2 are set, and the minority class samples can be classified as follows:
[0040]
[0041] Next, from all boundary samples, new boundary samples are synthesized based on the nearest minority class samples to supplement the key categories in the data:
[0042] x new =x i +λ·(x neighbor -x i ),λ~Uniform(0,1)
[0043] The essence of SMOTE is to expand the minority class through the idea of resampling, but there is a problem of unstable sample quality. Therefore, the ENN method is further introduced to delete the synthesized bad samples by undersampling; according to the category distribution judgment formula Set 0.5 as the ENN bad sample evaluation score threshold; when the Noise score is less than 0.5, it means that the synthesized sample is inconsistent with more than half of the sample categories in its neighbors, that is, a new noise sample is generated, and the new sample x' i :
[0044] Noisei =Count(Neighbors(x′ i ≠y i ))
[0045] After completing the above synthesis and evaluation steps, the new and old samples are combined as the data set to be tested.
[0046] Preferably, in step S3, under the goal of refining the time step, BiLSTM is introduced to complete the extraction of time features, and TCN is used to complete the extraction of spatial features. BiLSTM simultaneously captures information from the forward and backward directions of the time series, so it can model time features more comprehensively. Especially in network traffic data, bidirectional dependence has significant advantages in identifying contextual patterns and capturing complex time series features (such as periodic changes or burst traffic). TCN effectively captures features of different scales through causal convolution and dilated convolution. While having privacy protection capabilities, it has less computational complexity than the self-attention mechanism, and can effectively improve the real-time requirements of the detection system.
[0047] Step 3.1, time feature extraction, BiLSTM four layers and its application in data The output results are:
[0048] BiLSTM layer, output
[0049] Dropout layer, to prevent overfitting, H'1 = Dropout (H (1) ,rate=0.2);
[0050] BiLSTM layer, output
[0051] Dropout layer, to prevent overfitting, H'1 = Dropout (H (1) ,rate=0.2);
[0052] The final output time feature matrix is
[0053] Step 3.2, spatial feature extraction, using the TCN model to output the spatial feature matrix. In addition to the input layer and the output layer, the TCN model in the present invention designs an on-demand adjusted TCN network layer for mixed minute-level and second-level multi-scale traffic collection data;
[0054] First, the network layer of TCN is composed of the basic TCN model layer unit consisting of the dilated causal convolution layer, the regularization layer, the activation function layer and the normalization layer;
[0055] Secondly, according to the data time step information counted in step S2, if the proportion of second-level time step samples exceeds Then four serial model layer units are set in TCN:
[0056] H1=ReLU(Norm(DildatedConv1D(X,k,d1,C)
[0057] H2=ReLU(Norm(DildatedConv1d(H1,k,d2,C)
[0058] H3=ReLU(Norm(DildatedConv1D(H2,k,d3,C)
[0059] H4=ReLU(Norm(DildatedConv1D(H3,k,d4,C)
[0060] Otherwise, the TCN is set to have three serial model layer units; the convolution kernel size in the model layer units of each layer is kept consistent, the number of input and output channels between each layer is kept consistent, and the expansion rate d is taken as {1, 2, 4, 8} and {1, 2, 4} respectively; The proposed spatial feature matrix is
[0061] By evaluating the time steps of different granularities in the traffic data and setting the number of different TCN model layer units, we can perform batch data detection tasks such as second-level traffic, improve the receptive field of model feature correlation, obtain more global feature dependencies, and output high-quality spatial features.
[0062] Preferably, in step S4, the fusion of temporal features and spatial features specifically includes the following steps:
[0063] Step 4.1, feature dimension alignment, determine whether the time step dimension of the temporal feature matrix is consistent with that of the spatial feature matrix. Otherwise, adjust the two model structures in step S3 to output feature matrices with the same time step dimension;
[0064] Step 4.2, feature concatenation, concatenates the feature dimension of the temporal feature matrix with the feature dimension of the spatial feature matrix to form a new joint feature matrix F, whose feature dimension is F Ttime +F relete ;
[0065] Step 4.3, normalization processing, normalizes the concatenated joint feature matrix to eliminate the differences between the distribution ranges of different features; specifically, calculates the mean and standard deviation of the joint feature matrix, and uses the formula:
[0066]
[0067] Normalize the joint feature matrix to a standard distribution to improve the feature weight calculation performance of the subsequent attention mechanism; the final spatiotemporal feature matrix This is the input feature of the multi-head TPA attention mechanism.
[0068] Preferably, in step S5, the fusion features calculated in step 4 are used to propose a multi-head attention mechanism with sub-second as input to achieve accurate detection of sub-second scale granularity flow data, which specifically includes the following steps:
[0069] Step 5.1, determine the QKV matrix of each head; the multi-head attention opportunity decomposes the input features into multiple heads, and each head independently calculates the QKV matrix; in order to better fit the characteristics of mixed traffic data, the present invention has a higher feature identification ability at the mixed granularity of seconds and minutes; the time window analysis strategy is used to determine the initialized linear matrix W Q ,W K ,W V :
[0070] 1) For attention head a, calculate the rolling mean and rate of change As a matrix value:
[0071]
[0072] Then the KQV matrices are:
[0073]
[0074] The time step unit is seconds, k = 60 or the sequence length value is the time window size;
[0075] 2) For attention head b, it focuses on exploring minute-level traffic, with seconds as the time step unit, k = 120 or the sequence length value as the time window size, and the statistical rolling mean and rate of change As the value of the matrix, the calculation method is consistent with 1), then the KQV matrix is:
[0076]
[0077] 3) For the attention head c, it focuses on the exploration of the mixed direction of minutes and seconds, taking the mean of all sample data and rate of change As the value of the matrix, the calculation method is consistent with 1), then the KQV matrix is:
[0078] Q (c) =diag(μ t )·F
[0079] K (c) =diag(Δ t )·F
[0080] V (c) =[diag(μ t )+diag(Δ t )]·F
[0081] Step 5.2, attention matrix calculation, each head is calculated according to the following formula:
[0082]
[0083] Step 5.3, result matrix output:
[0084] Z (m) =A (m) V (m) ,
[0085] Step 5.4, merge the multi-head results:
[0086] Z=Concat(Z (a) ,Z (b) ,Z (c) ),
[0087] Preferably, in step S6, the last layer of the multi-head attention mechanism is a fully connected layer, which is used to map the concatenation result Z calculated by the attention mechanism to the final classification label
[0088]
[0089] E is a matrix of all 1s,
[0090]
[0091] The weight matrix W of the fully connected layer O and bias b O It is calculated through the back-propagation process; during the back-propagation process, the model evaluates the difference between the predicted result and the true label through the Softmax and cross entropy loss functions, and updates W according to the gradient information O and b O Gradually optimize the classification ability of the model;
[0092] Through the designed multi-head attention mechanism with second, minute and mixed granularity, it is possible to capture rapid traffic changes in a short time range (seconds), traffic trends in a longer time span (minutes), and comprehensive patterns that integrate short and long time features (mixed granularity). This mechanism can fully adapt to the complex characteristics of multi-time scale traffic data, promote the dynamic allocation and balance of time granularity, and effectively meet the needs of high-quality and timely intrusion detection.
[0093] The present invention also provides a multi-model fusion efficient intrusion detection system, including a data sample cleaning module, a sample generation module, a spatiotemporal feature extraction module, and a feature weight adjustment and classification module.
[0094] The data sample cleaning module is used for sample outlier removal, TSODE time feature screening, rule matching independent feature screening, and z-score statistical feature screening;
[0095] The sample generation module is used for BorderLine-SMOTE sample generation and ENN elimination to generate noise samples;
[0096] The spatiotemporal feature extraction module is used for extracting spatial features using the TCN model, extracting temporal features using the BiLSTM model, and feature fusion;
[0097] The feature weight adjustment and classification module is used to generate initial weights, split-second multi-head attention mechanism and fully connected layer classification based on the time post method.
[0098] Compared with the prior art, the present invention has the following beneficial effects: the present invention provides an efficient intrusion detection method and system with multi-model fusion, combines technologies such as TSODE to enhance sample cleaning, uses the SMOTE method to target the minority classes on the decision boundary for supplementation, proposes that the minute and second are the basic granularity, designs TCN+BiLSTM to comprehensively extract spatiotemporal features, combines the TPA attention mechanism and the multi-head attention mechanism to design the minute and second multi-head attention mechanism to deepen the dynamic allocation of feature weights, and effectively provides a stable and high-quality intrusion detection model. BRIEF DESCRIPTION OF THE DRAWINGS
[0099] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.
[0100] In the attached picture:
[0101] Figure 1 It is a schematic diagram of a module of a multi-model fusion efficient intrusion detection system of the present invention;
[0102] Figure 2It is a schematic diagram of a process flow of a multi-model fusion efficient intrusion detection method of the present invention;
[0103] Figure 3 It is a TCN model diagram of the present invention;
[0104] Figure 4 It is the BiLSTM model diagram of the present invention;
[0105] Figure 5 It is a model diagram of the multi-head attention mechanism of the present invention. DETAILED DESCRIPTION
[0106] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0107] Embodiment: The present invention combines multiple sequence models of deep learning to improve data cleaning and sample classification capabilities. To this end, in the data cleaning stage, outlier removal and TSODE and Z-score technologies are integrated to deepen the removal of noise samples; in the feature extraction stage, on the basis of basic extraction of spatiotemporal features, the attention mechanism is introduced to deepen the dynamic allocation of feature weights; in the classification detection stage, high-quality detection results are output in the fully connected layer based on the attention weights.
[0108] The present invention provides an efficient intrusion detection method of multi-model fusion. The overall execution process of the Bi3T method is as follows: Figure 2 As shown, the specific steps of the method include:
[0109] Step 1: Data sample cleaning.
[0110] In order to more generally handle features of different types and functions, based on the statistical analysis of traffic data, the traffic data features are classified into four categories: time features, fixed format features, statistical features, and other features. Outlier removal, TSODE method, and z-score method are used to process each type of feature to complete data sample cleaning.
[0111] Step 1.1, scan the samples for outlier samples with characteristic values of inf and nan, and remove them directly;
[0112] Step 1.2: Based on the outlier cleaning in step 1.1, the TSODE method is used to perform two stages of single feature deviation evaluation and multi-feature joint evaluation for the time feature column, and the abnormal samples whose time feature values deviate from the normal model are eliminated;
[0113] Step 1.2.1, identify whether there are outliers in the time features that deviate from the normal pattern. Use the sliding window method toj ={x 1j ,x 2j ,…,x kj}, calculate the rolling mean μ of the feature in k time steps t and standard deviation α t :
[0114]
[0115]
[0116] Where k is the window size. After the results of the single variable learning curve method and grid search, it is determined that the value range of k is [5,11]. In the present invention, setting k to 7 can effectively take into account the time consumption of outlier discovery and calculation; and if the current sample meets:
[0117] |x tj -μ t |≤λ·α t
[0118] If the current sample time feature conforms to the normal mode of the data set, then the mean μ of the current time window is t Replace x tj , so as to ensure that other information of the sample is not lost. Where λ = 2.5.
[0119] Step 1.2.2, establish joint feature screening of multiple time features. Taking the time step as the basic unit, calculate the joint outlier score of the d time features on the sample corresponding to the current time step:
[0120]
[0121] Where α = 2, and when the feature score is greater than the threshold T, it is considered an outlier and the sample is removed. This application takes into account the multi-scale traffic data at the second and minute levels that may appear in the data set, and the joint score is approximately a normal distribution or a right-skewed distribution, so the threshold T is set:
[0122] T=μ+β·σ
[0123] Where μ and σ are the mean and standard deviation of all time features of the current sample. When the proportion of minute-level data exceeds two-thirds, β is 2.5, otherwise it is 1.5. t >T, it is considered that the joint distribution of the current sample deviates and the sample is directly eliminated.
[0124] Step 1.3, for fixed format features, they usually have fixed formats or specifications. Therefore, based on regular grammar, combined with relevant specifications and protocol formats of network traffic, custom special rules are set as shown in the following table:
[0125] Table 1 Fixed format feature customization rules (partial)
[0126]
[0127] For samples that fail to match, custom rules can effectively locate the location of the sample. By manually reviewing the sample, it can be determined whether the fixed format feature value is "the format is incorrect, but the content is feasible". Samples with real erroneous content will be eliminated, otherwise the correct format will be restored.
[0128] In step 1.4, the measure derived from the original data or primary features is the statistical feature. Since it usually conforms to the approximate normal distribution, the z-score method is proposed to remove samples with feature value evaluation scores exceeding 3; the z-score can convert the original feature value into a dimensionless standard score, which no longer depends on a specific numerical range, making the anomaly detection of different features comparable. j :
[0129]
[0130] Among them, μ j ,σ j is the mean and standard deviation of the current feature at all time steps. Samples with scores out of range will be directly removed.
[0131] Through four outlier processing steps, the quality of network traffic data can be effectively improved, the interference of abnormal noise on model training can be reduced, and the robustness and generalization ability of the model can be improved. These processing measures can ensure that time features reflect real dynamic behaviors, fixed format features such as IP addresses and port numbers conform to the standard format, and statistical features such as traffic rate and packet size are more evenly distributed.
[0132] Step 2: Sample generation.
[0133] Abnormal traffic is a key analysis target in intrusion detection tasks. The importance of abnormal samples is much higher than that of normal samples. And abnormal traffic is usually rarely mixed in normal traffic. Therefore, the present invention adopts the BorderLine-SMOTE-ENN method to balance the problem of uneven sample distribution.
[0134] The first step is to confirm the boundary samples, for the minority class samples x i , use KNN to calculate the minority class ratio. The minority class ratio refers to the proportion of samples that belong to the same category as the sample in the first k samples of feature similarity:
[0135]
[0136] The k neighbors are selected by calculating the formula:
[0137]
[0138] Where D is the feature dimension, Represents sample x i The d eigenvalues are set, and k is 5; the boundary thresholds T1 = 0.5 and T2 = 0.2 are set, and the minority class samples can be classified as follows:
[0139]
[0140] Next, from all boundary samples, new boundary samples are synthesized based on the nearest minority class samples to supplement the key categories in the data:
[0141] x new =x i +λ·(x neighbor -x i ),λ~Uniform(0,1)
[0142] The essence of SMOTE is to expand the minority class through the idea of resampling, but there is a problem of unstable sample quality. Therefore, the present invention further introduces the ENN method to delete the synthesized bad samples by undersampling. By setting 0.5 as the ENN bad sample evaluation score threshold. When the Noise score is less than 0.5, it means that the synthesized sample is inconsistent with more than half of the sample categories in its neighbors, that is, a new noise sample is generated. i :
[0143] Noise i =Count(Neighbors(x′ i ≠y i ))
[0144] After completing the above synthesis and evaluation steps, the new and old samples are combined as the data set to be tested.
[0145] Step 3: Extract spatiotemporal features.
[0146] In view of the rich time series dependencies contained in network traffic data and the wide distribution of multi-dimensional features, in-depth exploration of the time series dependencies of traffic data is a key goal of intrusion detection. Under the established goal of refining the time step, the present invention introduces BiLSTM to complete the extraction of time features and uses TCN to complete the extraction of spatial features. BiLSTM captures information from both the forward and backward directions of the time series, so it can model time features more comprehensively. Especially in network traffic data, bidirectional dependencies have significant advantages in identifying contextual patterns and capturing complex time series features (such as periodic changes or burst traffic). TCN effectively captures features of different scales through causal convolution and dilated convolution. While having privacy protection capabilities, it has less computational complexity than the self-attention mechanism, and can effectively improve the real-time requirements of the detection system. Specifically, the following steps are included:
[0147] Step 3.1, time feature extraction, BiLSTM four layers and its application in data The output results are:
[0148] BiLSTM layer, output
[0149] Dropout layer, to prevent overfitting, H'1 = Dropout (H (1) ,rate=0.2);
[0150] BiLSTM layer, output
[0151] Dropout layer, to prevent overfitting, H'1 = Dropout (H (1) ,rate=0.2);
[0152] The final output time feature matrix is The specific Bi3T spatial feature extraction model TCN structure is shown as follows Figure 3 shown.
[0153] Step 3.2, spatial feature extraction, using the TCN model to output the spatial feature matrix. In addition to the input layer and the output layer, the TCN model in the present invention designs an on-demand adjusted TCN network layer for mixed minute-level and second-level multi-scale traffic collection data.
[0154] First, the network layer of TCN is composed of {expanded causal convolution layer, regularization layer, activation function layer and normalization layer}, which are the basic TCN model layer units;
[0155] Secondly, according to the data time step information counted in step 2, if the proportion of second-level time step samples exceeds Then four serial model layer units are set in TCN; otherwise, three serial model layer units are set in TCN. The convolution kernel size in the model layer units of each layer is kept consistent, the number of input and output channels between each layer is kept consistent, and the expansion rate d is taken as {1,2,4,8} and {1,2,4} respectively.
[0156] H1=ReLU(Norm(DildatedConv1D(X,k,d1,C)
[0157] H2=ReLU(Norm(DildatedConv1D(H1,k,d2,C)
[0158] H3=ReLU(Norm(DildatedConv1D(H2,k,d3,C)
[0159] H4=ReLU(Norm(DildatedConv1D(H3,k,d4,C)
[0160] The proposed spatial feature matrix is The specific BiLSTM structure of the Bi3T temporal feature extraction model is shown as follows Figure 4 shown.
[0161] By evaluating the time steps of different granularities in the traffic data and setting the number of different TCN model layer units, we can perform batch data detection tasks such as second-level traffic, improve the receptive field of model feature correlation, obtain more global feature dependencies, and output high-quality spatial features.
[0162] Step 4: fusion of spatiotemporal features.
[0163] The temporal and spatial features are integrated to provide a more comprehensive feature representation and improve the model's ability to capture complex behavior patterns. The specific steps include:
[0164] Step 4.1: Align feature dimensions. Determine whether the time step dimensions of the temporal feature matrix and the spatial feature matrix are consistent. Otherwise, adjust the two model structures in step 3 to output feature matrices with consistent time step dimensions.
[0165] Step 4.2, feature concatenation. The feature dimension of the temporal feature matrix is concatenated with the feature dimension of the spatial feature matrix to form a new joint feature matrix F, whose feature dimension is F Ttime +F relete
[0166] Step 4.3, normalization. The concatenated joint feature matrix is normalized to eliminate the differences between the distribution ranges of different features. Specifically, the mean and standard deviation of the joint feature matrix are calculated and the formula is:
[0167]
[0168] Normalize the joint feature matrix to a standard distribution to improve the feature weight calculation performance of the subsequent attention mechanism. The final spatiotemporal feature matrix This is the input feature of the multi-head TPA attention mechanism.
[0169] Step 5: Feature weight assignment
[0170] In view of the fact that the TPA attention mechanism adds a temporal pattern extraction module to display the local information characteristics of the modeled time series, and uses this feature as the input of the attention mechanism, it can achieve an effect that focuses more on processing time series data than the traditional attention mechanism. It introduces position awareness to extract key information from the feature matrix, while strengthening the importance of time features and position to the final output. Although it tends to be local in terms of dependency capture range, it has a faster calculation speed. Therefore, based on this idea, and the fusion feature F calculated in step 4, a split-second multi-head attention mechanism with F as input is proposed to achieve accurate detection of split-second scale granularity traffic data; the specific Bi3T split-second multi-head attention mechanism model structure is shown as follows Figure 5 As shown, the specific steps include:
[0171] Step 5.1, determine the QKV matrix of each head. The multi-head attention machine decomposes the input features into multiple heads, and each head independently calculates the QKV matrix. In order to better fit the characteristics of mixed traffic data, the present invention has a higher feature identification ability at the mixed granularity of seconds and minutes. Determine the initialized linear matrix W using the time window analysis strategy Q ,W K ,W V :
[0172] For attention head a, the time step unit is seconds, k = 60 or the sequence length value is the time window size, and the statistical rolling mean and rate of change As a matrix value:
[0173]
[0174] Then the KQV matrices are:
[0175]
[0176]
[0177] For attention head b, it focuses on exploring minute-level traffic, with seconds as the time step unit, k = 120 or the sequence length value as the time window size, and the statistical rolling mean and rate of change As the value of the matrix, the calculation method is the same as If the internal structure is consistent, the KQV matrices are:
[0178]
[0179] For the attention head c, it focuses on the exploration of the mixed direction of minutes and seconds, taking the mean of all sample data and rate of change As the value of the matrix, the calculation method is the same as If the internal structure is consistent, the KQV matrices are:
[0180] Q (c) =diag(μ t )·F
[0181] K (c) =diag(Δ t )·F
[0182] V (c) =[diag(μ t )+diag(Δ t )]·F
[0183] Step 5.2, attention matrix calculation, each head is calculated according to the following formula:
[0184]
[0185] Step 5.3, result matrix output:
[0186] Z (m) =A (m) V (m) ,
[0187] Step 5.4, merge the multi-head results:
[0188] Z=Convat(Z (a) ,Z (b) ,Z (c) ),
[0189] Step 6: Classification and result output.
[0190] The last layer of the multi-head attention mechanism is a fully connected layer, which is used to map the concatenation result Z calculated by the attention mechanism to the final classification label.
[0191]
[0192] E is a matrix of all 1s,
[0193]
[0194] The weight matrix W of the fully connected layer O and bias b O It is calculated through the back propagation process. During the back propagation process, the model evaluates the difference between the predicted result and the true label through the Softmax and cross entropy loss functions, and updates W based on the gradient information. O and b O Progressively optimize the classification ability of the model.
[0195] Through the designed multi-head attention mechanism with second, minute and mixed granularity, it is possible to capture rapid traffic changes in a short time range (seconds), traffic trends in a longer time span (minutes), and a comprehensive pattern that integrates short and long time features (mixed granularity). This mechanism can fully adapt to the complex characteristics of multi-time scale traffic data, promote the dynamic allocation and balance of time granularity, and effectively meet the needs of high-quality and timely intrusion detection.
[0196] like Figure 1 As shown, the present invention also provides an efficient intrusion detection system with multi-model fusion, including a data sample cleaning module, a sample generation module, a spatiotemporal feature extraction module, and a feature weight adjustment and classification module.
[0197] The data sample cleaning module is used for sample outlier removal, TSODE time feature screening, rule matching independent feature screening, and z-score statistical feature screening;
[0198] The sample generation module is used for BorderLine-SMOTE sample generation and ENN elimination to generate noise samples;
[0199] The spatiotemporal feature extraction module is used for extracting spatial features using the TCN model, extracting temporal features using the BiLSTM model, and feature fusion;
[0200] The feature weight adjustment and classification module is used to generate initial weights, split-second multi-head attention mechanism and fully connected layer classification based on the time post method.
[0201] The present invention proposes a multi-model fusion efficient intrusion detection method, which combines the deep extraction and fusion of temporal and spatial features of BiLSTM and TCN networks through cleaning operations for different categories of features, and further proposes a multi-head attention mechanism at the second, minute and mixed granularity, which effectively solves key problems such as uneven sample distribution and unbalanced feature importance in traffic data. In the feature modeling of multiple time scales, the scheme can adapt to the mixed granularity traffic features, ensure classification accuracy while providing stable high-quality detection capabilities, which is suitable for distributed networks and real-time detection needs, and provides a forward-looking and practical solution for intrusion detection tasks in current network security scenarios.
[0202] Finally, it should be noted that the above description is only a preferred example of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An efficient intrusion detection method based on multi-model fusion, characterized in that: The following steps are involved: S1. Data sample cleaning, including removing all inf and nan outliers in the sample data, using TSODE to screen for time features, using custom regular specifications to screen for fixed format features, and using z-score to screen for statistical features; S2, sample generation, including using the sliding window method to confirm the minority class that meets the properties of the decision boundary, using BorderLine-SMOTE to generate samples, and using ENN to verify the generated samples; S3, spatiotemporal feature extraction, including using TCN to extract spatial features of sample data and using BiLSTM to extract flow features of sample data; S4, spatiotemporal feature fusion, including fusing and normalizing temporal and spatial features, including feature dimension alignment, feature splicing, and normalization processing; S5, feature weight assignment, including generating an initial weight matrix by using statistical sample means of time windows with different time step lengths, and calculating a QKV matrix based on the initial weight matrix; S6. Based on the attention mechanism results, the classification results are output in the fully connected layer.
2. According to claim 1, the efficient intrusion detection method of multi-model fusion is characterized by: In step S1, based on the statistical analysis of traffic data, traffic data features are classified into four categories: time features, fixed format features, statistical features, and other features; outlier removal, TSODE method, and Z-score method are used to process each type of feature to complete data sample cleaning, which specifically includes the following steps: Step 1.1, scan the samples for outlier samples with characteristic values of inf and nan, and remove them directly; Step 1.2: Based on the outlier cleaning in step 1.1, the TSODE method is used to perform two stages of single feature deviation evaluation and multi-feature joint evaluation for the time feature column, and the abnormal samples whose time feature values deviate from the normal model are eliminated; Step 1.3: For fixed format features, they usually have a fixed format or specification. Therefore, based on regular grammar, combined with relevant specifications and protocol formats of network traffic, custom feature rules are set; for samples that fail to match, custom rules will locate the location of the sample, and a manual review method is used to determine whether the fixed format feature value is "the format is incorrect, but the content is feasible". Samples with real erroneous content are eliminated, otherwise the correct format is restored; Step 1.4: The measure derived from the original data or primary features is the statistical feature. Since it usually conforms to the approximate normal distribution, the z-score method is proposed to remove samples with feature value evaluation scores exceeding 3. The z-score converts the original feature value into a dimensionless standard score, which no longer depends on a specific numerical range, making the anomaly detection of different features comparable. For the statistical feature x j : Among them, μ j ,σ j is the mean and standard deviation of the current feature at all time steps; samples with scores out of range will be directly removed.
3. The efficient intrusion detection method of multi-model fusion according to claim 2 is characterized by: Step 1.2 specifically includes the following steps: Step 1.2.1, identify whether there are outliers in the time features that deviate from the normal pattern one by one; use the sliding window method to j ={x 1j ,x 2j ,…,x kj }, calculate the rolling mean μ of the feature in k time steps t and standard deviation α t : Where k is the window size; after the results of the single variable learning curve method and grid search, it is determined that the value range of k is [5,11]. In the present invention, setting k to 7 can effectively take into account the time consumption of outlier discovery and calculation; and if the current sample meets: |x tj -m t |≤λ·a t If the current sample time feature conforms to the normal mode of the data set, otherwise the mean μ of the current time window is t Replace x tj , so as to ensure that other information of the sample is not lost; where λ = 2.5; Step 1.2.2, establish joint feature screening of multiple time features; take the time step as the basic unit, and calculate the joint outlier score of d time features on the sample corresponding to the current time step: Where α = 2, and when the feature score is greater than the threshold T, it is considered an outlier and the sample is removed; considering the multi-scale traffic data at the second and minute levels that may appear in the data set, and the joint score is approximately a normal distribution or a right-skewed distribution, the threshold T is set: T=μ+β·σ Where μ, σ are the mean and standard deviation of all time features of the current sample; β takes the value of 2.5 when the proportion of minute-level data exceeds two-thirds, otherwise it takes the value of 1.5; and for Score t >T, it is considered that the joint distribution of the current sample deviates and the sample is directly eliminated.
4. The efficient intrusion detection method of multi-model fusion according to claim 1 is characterized by: In step S2, after determining that there is a problem with the dataset to be detected, the BorderLine-SMOTE combined with the ENN method is used to solve the problem, which specifically includes the following steps: The first step is to confirm the boundary samples, for the minority class samples x i , use KNN to calculate the minority class ratio, where the minority class ratio refers to the proportion of samples that belong to the same category as the sample in the first k samples of feature similarity: The k neighbors are selected by calculating the formula: Where D is the feature dimension, Represents sample x i The d eigenvalues are set, and k is 5; the boundary thresholds T1 = 0.5 and T2 = 0.2 are set, and the minority class samples can be classified as follows: Next, from all boundary samples, new boundary samples are synthesized based on the nearest minority class samples to supplement the key categories in the data: x new =x i +λ·(x neighbor -x i ),λ~Uniform(0,1) The ENN method is further introduced to delete the synthesized bad samples by undersampling; according to the category distribution judgment formula Set 0.5 as the ENN bad sample evaluation score threshold; when the Noise score is less than 0.5, it means that the synthesized sample is inconsistent with more than half of the sample categories in its neighbors, that is, a new noise sample is generated, and the new sample x' i : Noise i =Count(Neighbors(x′ i ≠y i )) After completing the above synthesis and evaluation steps, the new and old samples are combined as the data set to be tested.
5. The efficient intrusion detection method of multi-model fusion according to claim 1 is characterized by: In step S3, under the established goal of refining the time step, BiLSTM is introduced to complete the extraction of temporal features, and TCN is used to complete the extraction of spatial features. BiLSTM captures information from both the forward and backward directions of the time series. Step 3.1, time feature extraction, BiLSTM four layers and its application in data The output results are: BiLSTM layer, output Dropout layer, to prevent overfitting, H'1 = Dropout (H (1) ,rate=0.2); BiLSTM layer, output Dropout layer, to prevent overfitting, H'1 = Dropout (H (1) ,rate=0.2); The final output time feature matrix is Step 3.2, spatial feature extraction, specifically includes the following steps: First, the network layer of TCN is composed of the basic TCN model layer unit consisting of the dilated causal convolution layer, the regularization layer, the activation function layer and the normalization layer; Secondly, according to the data time step information counted in step S2, if the proportion of second-level time step samples exceeds Then four serial model layer units are set in TCN: H1=ReLU(Norm(DildatedConv1D(X,k,d1,C) H2=ReLU(Norm(DildatedConv1D(H1,k,d2,C) H3=ReLU(Norm(DildatedConv1D(H2,k,d3,C) H4=ReLU(Norm(DildatedConv1D(H3,k,d4,C) Otherwise, the TCN is set to have three serial model layer units; the convolution kernel size in the model layer units of each layer is kept consistent, the number of input and output channels between each layer is kept consistent, and the expansion rate d is taken as {1, 2, 4, 8} and {1, 2, 4} respectively; The proposed spatial feature matrix is 6. The multi-model fusion efficient intrusion detection method according to claim 1 is characterized by: In step S4, the temporal features and spatial features are fused. The specific steps include: Step 4.1, feature dimension alignment, determine whether the time step dimension of the temporal feature matrix is consistent with that of the spatial feature matrix. Otherwise, adjust the two model structures in step S3 to output feature matrices with the same time step dimension; Step 4.2, feature concatenation, concatenates the feature dimension of the temporal feature matrix with the feature dimension of the spatial feature matrix to form a new joint feature matrix F, whose feature dimension is F Ttime +F relete ; Step 4.3, normalization processing, normalizes the concatenated joint feature matrix to eliminate the differences between the distribution ranges of different features; specifically, calculates the mean and standard deviation of the joint feature matrix, and uses the formula: Normalize the joint feature matrix to the standard distribution, and finally the spatiotemporal feature matrix This is the input feature of the multi-head TPA attention mechanism.
7. The efficient intrusion detection method of multi-model fusion according to claim 1 is characterized by: In step S5, the fusion features calculated in step 4 are used to propose a multi-head attention mechanism with sub-second as input to achieve accurate detection of sub-second scale granularity flow data, which specifically includes the following steps: Step 5.1, determine the QKV matrix of each head; the multi-head attention opportunity decomposes the input features into multiple heads, and each head independently calculates the QKV matrix; determine the initialized linear matrix W using the time window analysis strategy Q ,W K ,W V : 1) For attention head a, calculate the rolling mean and rate of change As a matrix value: Then the KQV matrices are: The time step unit is seconds, k = 60 or the sequence length value is the time window size; 2) For attention head b, it focuses on exploring minute-level traffic, with seconds as the time step unit, k = 120 or the sequence length value as the time window size, and the statistical rolling mean and rate of change As the value of the matrix, the calculation method is consistent with 1), then the KQV matrix is: 3) For the attention head c, it focuses on the exploration of the mixed direction of minutes and seconds, taking the mean of all sample data and rate of change As the value of the matrix, the calculation method is consistent with 1), then the KQV matrix is: Q (c) =diag(μ t )·F K (c) =diag(Δ t )·F V (c) =[diag(μ t )+diag(Δ t )]·F Step 5.2, attention matrix calculation, each head is calculated according to the following formula: Step 5.3, result matrix output: Step 5.4, merge the multi-head results:
8. The efficient intrusion detection method of multi-model fusion according to claim 1 is characterized by: In step S6, the last layer of the multi-head attention mechanism is a fully connected layer, which is used to map the concatenation result Z calculated by the attention mechanism to the final classification label E is a matrix of all 1s, The weight matrix W of the fully connected layer O and bias b O It is calculated through the back-propagation process; during the back-propagation process, the model evaluates the difference between the predicted result and the true label through the Softmax and cross entropy loss functions, and updates W according to the gradient information O and b O Progressively optimize the classification ability of the model.
9. An efficient intrusion detection system with multi-model fusion, characterized by: It includes data sample cleaning module, sample generation module, spatiotemporal feature extraction module, and feature weight adjustment and classification module. The data sample cleaning module is used for sample outlier removal, TSODE time feature screening, rule matching independent feature screening, and z-score statistical feature screening; The sample generation module is used for BorderLine-SMOTE sample generation and ENN elimination to generate noise samples; The spatiotemporal feature extraction module is used for extracting spatial features using the TCN model, extracting temporal features using the BiLSTM model, and feature fusion; The feature weight adjustment and classification module is used to generate initial weights, split-second multi-head attention mechanism and fully connected layer classification based on the time post method.
Citation Information
Patent Citations
Reflective vulnerability detection method based on static and dynamic combination
CN109462583A
Abnormal network flow detection method based on bidirectional time convolutional neural network
CN115037543A
Transform and neural network intrusion detection method and system
CN118199969A
Intrusion detection method and system combining attention mechanism and MSCNN + BiLSTM
CN118839242A
Cited By
RC component impact failure mode intelligent prediction method, device and equipment and storage medium
CN121413345A
Intelligent prediction method, device and equipment for impact failure mode of rc component and storage medium
CN121413345B
Test bed pneumatic valve state intelligent detection method and system
CN121521484A
Intelligent detection method and system for the status of pneumatic valves on test bench
CN121521484B