Secondary shunt detection method of intrusion detection system by using deep learning

By adopting the secondary shunt detection methods of Swin Transformer and MLP-Mixer in the intrusion detection system, the problem that traditional deep learning models are difficult to classify boundary data in intrusion detection is solved, and higher detection accuracy and real-time performance are achieved.

CN119966654AActive Publication Date: 2025-05-09INST OF INT RELATIONS
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202411832314.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-05-09
Estimated Expiration
2044-12-12

AI Technical Summary

Technical Problem

Traditional deep learning models are difficult to effectively classify boundary data in intrusion detection, resulting in underreporting problems and unable to identify potential threats in a timely manner.

Method used

Swin Transformer is used as the primary shunt detection platform, and network traffic is divided into abnormal, normal and suspicious traffic through a phased attention mechanism. Then, suspicious traffic is delivered to MLP-Mixer as a secondary shunt detection platform, and each block is extracted and combined through a multi-layer perceptron to finally distinguish suspicious traffic.

Benefits of technology

It effectively solves the problem of fuzzy classification of boundary data, improves the accuracy of intrusion detection and real-time identification, and reduces the risk of underreport.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119966654A_ABST
    Figure CN119966654A_ABST
Patent Text Reader

Abstract

In order to solve the problem that boundary data classification is fuzzy when a machine learning technology is applied to the field of intrusion detection, the invention provides a deep learning-based intrusion detection system two-stage shunting detection method, which can effectively solve the problem that boundary data classification is fuzzy. Comprising the following steps: step 1, network traffic is divided into abnormal traffic, normal traffic and suspicious traffic through a primary shunting detection platform based on a Swin Transform classifier, and the suspicious traffic is transmitted to a secondary detection platform; and 2, dividing the suspicious traffic data into a plurality of blocks by adopting an MLP-Mixer as a secondary shunting detection platform, then performing feature extraction and combination on each block through a multi-layer sensor, and finally performing final distinguishing on the suspicious traffic to obtain a normal traffic result and an abnormal traffic result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intrusion detection, and in particular relates to a secondary shunt detection method for an intrusion detection system using deep learning. Background Art

[0002] Network traffic security is extremely important. In today's digital age, a large amount of personal, commercial and government information is transmitted and stored over the network. This information includes personal privacy, financial data, business secrets and national security information. Therefore, protecting network traffic security is crucial to maintaining personal privacy, preventing economic losses, protecting corporate interests and maintaining national security.

[0003] Unprotected network traffic is vulnerable to various threats, such as hacker attacks, phishing, ransomware, spyware, and data leaks. These threats may lead to serious consequences such as theft of personal information, leakage of business secrets, financial losses, and social instability. Therefore, ensuring the security of network traffic is an important measure to protect personal privacy, commercial interests, and national security. Through encrypted communications, access control, firewalls, intrusion detection systems, security updates, network traffic analysis, security training and other measures, the risk of network traffic being attacked and leaked can be effectively reduced, and the security and stability of the network ecology can be maintained.

[0004] Intrusion detection system is one of the important strategies for detecting traffic security. Intrusion detection technology based on deep learning is an emerging research direction in the field of network security in recent years. Traditional intrusion detection systems usually rely on manually defined rules or features to identify malicious behaviors, but this method is difficult to capture complex and unknown attack patterns. Deep learning technology can more effectively identify potential intrusion behaviors by learning a large amount of data and automatically extracting features. Intrusion detection systems based on deep learning usually use convolutional neural networks (CNN), recurrent neural networks (RNN), long short-term memory networks (LSTM) or variant deep neural networks to process network traffic or host log data.

[0005] In the boundary data division part, patent CN201610427861.7 discloses a boundary data division method and device, which obtains the associated high-density interval of the boundary data through the undisputed data of the associated cluster group in the clustering result, and intercepts the concentrated data in the associated high-density interval from the undisputed data of the associated cluster group, and then analyzes the similarity between the boundary data and the concentrated data in the associated high-density interval, and finally divides the boundary data based on the similarity, which can accurately classify the boundary data, so that the data classification is accurate and lossless. Patent CN201911075244.5 discloses a method and system for boundary data analysis using clustering method, which presets key variables and thresholds in various types of log data generated by boundary data exchange behavior, and then uses clustering algorithm to classify the data to obtain cluster analysis results, and then performs cluster analysis on the data generated by the new boundary data exchange behavior, and compares the results with the constructed form to find outliers and mark and count them, and issue an alarm after exceeding the threshold.

[0006] In the intrusion detection system part, patent CN202111044147.7 discloses an intrusion detection system that can classify events through a pre-built event analyzer. The pre-built event analyzer is trained based on historical event records and can automatically identify events as normal events or intrusion events based on event characteristics, thereby improving the real-time recognition and accuracy of the intrusion detection system. Patent CN202211485456.2 discloses a network security intrusion detection system and detection method, which are used to solve the problem that the existing network security intrusion detection system can only intercept specific or continuous intrusion behaviors, has a high false alarm rate, and the network security intrusion detection system cannot restrict the intrusion terminal, resulting in the intrusion terminal being able to easily invade again, causing serious losses.

[0007] Traditional intrusion detection systems usually rely on manually defined rules or features to identify malicious behavior, but this method has difficulty capturing complex and unknown attack patterns. Deep learning technology can more effectively identify potential intrusion behaviors by learning large amounts of data and automatically extracting features. These neural network models can automatically learn and extract features from data, thereby reducing the reliance on manual feature engineering to a certain extent. However, intrusion detection solutions based on deep learning generally have a disadvantage, that is, they cannot fully classify boundary data. Boundary data usually refers to those data points that are ambiguous or unclear between normal and abnormal. These data points may have potential threats, but due to their high similarity to normal behavior, traditional deep learning models often have difficulty effectively classifying them as abnormal or normal. In this case, deep learning models tend to classify boundary data as normal, resulting in the problem of underreporting, that is, failure to identify potential threats in a timely manner. Summary of the invention

[0008] Aiming at the problem of fuzzy classification of boundary data that occurs when machine learning technology is applied in the field of intrusion detection, the present invention proposes a two-level diversion detection method of an intrusion detection system using deep learning, which can effectively solve the problem of fuzzy classification of boundary data.

[0009] The present invention is achieved through the following technical solutions.

[0010] A secondary diversion detection method for an intrusion detection system using deep learning comprises the following steps:

[0011] Step 1: The network traffic is divided into abnormal traffic, normal traffic and suspicious traffic through the primary diversion detection platform based on the Swin Transformer classifier, and the suspicious traffic is transmitted to the secondary detection platform;

[0012] Step 2: Use MLP-Mixer as a secondary diversion detection platform to divide the suspicious traffic data into several blocks, then use a multi-layer perceptron to extract and combine features of each block, and finally distinguish the suspicious traffic to obtain the results of normal traffic and abnormal traffic.

[0013] Beneficial effects of the present invention:

[0014] 1. The present invention uses the primary traffic diversion detection platform to divide network traffic into abnormal, normal and suspicious traffic, where suspicious traffic refers to data that cannot be accurately classified by the primary detection. Then, the suspicious traffic is finally judged by the secondary detection platform;

[0015] 2. The present invention adopts Swin Transformer and introduces a staged attention mechanism to divide the input data into several blocks and apply self-attention mechanism at different levels to capture information of different scales and resolutions. This staged attention mechanism helps to effectively model long-distance dependencies while reducing the number of parameters and computational complexity. In addition, Swin Transformer also adopts the strategy of exchanging attention windows, that is, exchanging the positions of attention windows at different levels to enhance the flexibility and generalization ability of the model;

[0016] 3. The present invention adopts Swin Transformer and MLP-Mixer to undertake the tasks of advanced feature extraction and deep analysis respectively, so that the intrusion detection system can effectively cope with complex traffic patterns;

[0017] 4. The intrusion detection technology based on Swin Transformer and MLP-Mixer deep learning realizes automatic learning and extraction of data features, which can more effectively identify potential intrusion behaviors. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is a block diagram of a two-level shunt detection method for an intrusion detection system using deep learning in the present invention; DETAILED DESCRIPTION

[0019] The exemplary embodiments of the present invention are described in detail below with reference to the accompanying drawings. It should be understood that the embodiments shown and described in the accompanying drawings are only exemplary and are intended to illustrate the principles and spirit of the present invention, rather than to limit the scope of the present invention.

[0020] like Figure 1 As shown, a secondary diversion detection method of an intrusion detection system using deep learning of the present invention specifically includes the following steps:

[0021] Step 1: The network traffic is divided into abnormal traffic, normal traffic and suspicious traffic through the primary diversion detection platform based on the Swin Transformer classifier, and the suspicious traffic is transmitted to the secondary detection platform; the specific steps are as follows:

[0022] The suspicious traffic mentioned in the present invention refers to the data flow that the primary diversion detection platform cannot accurately determine whether it is abnormal traffic.

[0023] 1.1 The Linear layer (input layer) is the initial layer of the model, which is used to reshape the dimension of the input feature (traffic data) from 87 to 32. The Swin Transformer divides the input network traffic data into several blocks, each of which has its own attention mechanism, expressed in the following formula:

[0024] X traffic =Linear(X traffic )

[0025] Attention(X traffic )=BlockAttention(StageNorm(X traffic ))+X traffic

[0026] Among them, X traffic represents the network traffic characteristics, StageNorm represents the standardization operation performed on each stage, BlockAttention represents the attention mechanism within each block; Attention(X traffic) indicates that the features of each stage are weighted by the attention mechanism and then added to the original features to obtain the final attention features;

[0027] From this step, we can see that this staged attention mechanism helps the model effectively capture information of features of different scales.

[0028] 1.2Swin Transformer adopts a cross-stage interaction method to share information between different stages, which is expressed by the following formula:

[0029] X traffic (l+1) =X traffic (l) +Block(StagePooling(X traffic (l) ))

[0030] Among them, X traffic (l) represents the traffic features of the first stage, StagePooling represents the pooling operation on the features of the current stage, and Block represents the processing of the pooled features; X traffic (l+1) It means that the features of the current stage are pooled and processed, and then added to the features of the previous stage to obtain the feature representation of the next stage;

[0031] 1.3Swin Transformer uses a local attention mechanism to reduce computational complexity and memory consumption. Local attention allows the model to focus only on information in a local area rather than information in a global range. Therefore, it is expressed in the following formula:

[0032] Attention(X traffic ) = LocalAttention(X traffic )+X traffic

[0033] Among them, X traffic represents the input traffic feature, LocalAttention is the local attention mechanism; Attention(X traffic ) indicates that the local attention mechanism processes the input feature and then adds it to the original feature to obtain the final attention feature;

[0034] 1.4Swin Transformer introduces cross-channel interaction to enhance the representation ability of features, which is expressed by the following formula:

[0035] X traffic (l+1) =X traffic(l) +LayerNorm(MlpBlock(X traffic (l) ))

[0036] Among them, X traffic (l) represents the traffic characteristics of the lth stage, MlpBlock is a module containing a multi-layer perceptron (MLP), LayerNorm represents the layer normalization of the MLP module output, and X traffic (l+1) It means that the features of the current stage are processed by the MLP module and layer normalized, and then added to the features of the previous stage to obtain the feature representation of the next stage;

[0037] 1.5 The Linear layer (output layer) converts the output dimension of the Transformer to the final output dimension, that is, converts the dimension from 32 to 1. The converted output will be used for the final binary classification prediction; the Sigmoid function in the Sigmoid layer is used to perform the activation operation in the binary classification task and map the output value to the range of [0,1];

[0038] (x suspect ,(x abnormal ,x normal )1)=Sigmoid(MLP(x))

[0039] Among them, x is the data of the layer, x suspevt is the suspicious outgoing traffic, (x abnormal ,x normal )1 is the abnormal flow and normal flow determined by output;

[0040] When implementing: set x normal >0.9 is normal flow, x abnormal <0.1 indicates abnormal traffic, and the output value range of suspicious traffic is 0.1 <x suspect <0.9.

[0041] Step 2: Use MLP-Mixer as a secondary diversion detection platform to divide the suspicious traffic data into several blocks, then use a multi-layer perceptron to extract and combine features of each block, and finally distinguish the suspicious traffic to obtain the results of normal traffic and abnormal traffic; the specific steps are as follows:

[0042] 2.1 Use PatchEmbed to represent the characteristics of suspicious traffic data; the specific formula is as follows:

[0043] T = PatchEmbed(X suspect )

[0044] Among them, Xsuspect represents the input suspicious traffic data, T represents the feature representation obtained by PatchEmbed, and PatchEmbed is the process of dividing the input traffic into uniform data blocks (or patches) and converting each patch into a low-dimensional vector;

[0045] 2.2 The feature vector of each patch is processed by a multi-layer perceptron to capture local information and perform feature mixing. The specific formula is as follows:

[0046] T′=TokenMix(T)

[0047] T″=ChannelMix(T′)

[0048] Among them, Token Mix means processing the feature vector of each patch through a multi-layer perceptron (MLP), and ChannelMix means processing the feature vector after Token Mix through global average pooling and projection across channels. It can be seen that this step helps to capture global information and mix features between different channels;

[0049] 2.3 The features are finally projected and represented through the fully connected layer. The specific formula is as follows:

[0050] Y=FC(T″)

[0051] Among them, FC means that the features are finally projected and represented through a fully connected layer, and the fully connected layer has multiple hidden layers for mapping the mixed features to the final output space.

[0052] 2.4 Using the MLP-Mixer secondary detection platform to suspect Classification as (x abnormal ,x normal )2. This completes x traffic →{(x abnormal ,x normal )1,(x abnormal ,x normal )2}, classify network traffic into normal traffic and abnormal traffic;

[0053] (x abnormal ,x normal )2=MLP-Mixer(x suspect ).

[0054] In summary, the above are only preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

[0055] It is obvious to those skilled in the art that the embodiments of the present invention are not limited to the details of the above exemplary embodiments, and that the embodiments of the present invention can be implemented in other specific forms without departing from the spirit or basic features of the embodiments of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive, and the scope of the embodiments of the present invention is limited by the attached claims rather than the above description, so it is intended to include all changes that fall within the meaning and scope of the equivalent elements of the claims in the embodiments of the present invention. Any figure mark in the claims should not be regarded as limiting the claims involved. In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units, modules or devices stated in the system, device or terminal claims can also be implemented by the same unit, module or device through software or hardware. The words first, second, etc. are used to indicate names, and do not indicate any particular order.

[0056] Finally, it should be noted that the above implementation modes are only used to illustrate the technical solutions of the embodiments of the present invention and are not intended to limit them. Although the embodiments of the present invention have been described in detail with reference to the above preferred implementation modes, those skilled in the art should understand that the technical solutions of the embodiments of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A secondary diversion detection method for an intrusion detection system using deep learning, characterized in that: The following steps are involved: Step 1: The network traffic is divided into abnormal traffic, normal traffic and suspicious traffic through the primary diversion detection platform based on the Swin Transformer classifier, and the suspicious traffic is transmitted to the secondary detection platform; Step 2: Use MLP-Mixer as a secondary diversion detection platform to divide the suspicious traffic data into several blocks, then use a multi-layer perceptron to extract and combine features of each block, and finally distinguish the suspicious traffic to obtain the results of normal traffic and abnormal traffic.

2. The method for secondary flow diversion detection of an intrusion detection system using deep learning as claimed in claim 1, characterized in that: The suspicious traffic refers to the data flow that the primary traffic diversion detection platform cannot accurately determine whether it is abnormal traffic. The specific judgment steps are as follows: 1.1 The Linear layer (input layer) is the initial layer of the model, which is used to reshape the dimension of the input feature from 87 to 32. SwinTransformer divides the input network traffic data into several blocks, each of which has its own attention mechanism, expressed by the following formula: X traffic =Linear(X traffic ) Attention(X traffic )=BlockAttention(StageNorm(X traffic ))+X traffic Among them, X traffic represents the network traffic characteristics, StageNorm represents the standardization operation performed on each stage, BlockAttention represents the attention mechanism within each block; Attention(X traffic ) indicates that the features of each stage are weighted by the attention mechanism and then added to the original features to obtain the final attention features; 1.2Swin Transformer adopts a cross-stage interaction method to share information between different stages, which is expressed by the following formula: X traffic (l+1) =X traffic (l) +Block(StagePooling(X traffic (l) )) Among them, X traffic (l) represents the traffic features of the first stage, StagePooling represents the pooling operation on the features of the current stage, and Block represents the processing of the pooled features; X traffic (l+1) It means that the features of the current stage are pooled and processed, and then added to the features of the previous stage to obtain the feature representation of the next stage; 1.3Swin Transformer uses a local attention mechanism to reduce computational complexity and memory consumption. Local attention allows the model to focus only on information in a local area rather than information in a global range. Therefore, it is expressed in the following formula: Attention(X traffic )=LocalAttention(X traffic )+X traffic Among them, X traffic represents the input traffic feature, LocalAttention is the local attention mechanism; Attention(X traffic ) indicates that the local attention mechanism processes the input feature and then adds it to the original feature to obtain the final attention feature; 1.4Swin Transformer introduces cross-channel interaction to enhance the representation ability of features, which is expressed by the following formula: X traffic (l+1) =X traffic (l) +LayerNorm(MlpBlock(X traffic (l) )) Among them, X traffic (l) represents the traffic characteristics of the lth stage, MlpBlock is a module containing a multi-layer perceptron (MLP), LayerNorm represents the layer normalization of the MLP module output, and X traffic (l+1) It means that the features of the current stage are processed by the MLP module and layer normalized, and then added to the features of the previous stage to obtain the feature representation of the next stage; 1.5 The Linear layer (output layer) converts the output dimension of the Transformer to the final output dimension, that is, converts the dimension from 32 to 1. The converted output will be used for the final binary classification prediction; the Sigmoid function in the Sigmoid layer is used to perform the activation operation in the binary classification task and map the output value to the range of [0,1]; (x suspect ,(x abnormal ,x normal )1)=Sigmoid(MLP(x)) Among them, x is the data of the layer, x suspect is the suspicious outgoing traffic, (x abnormal ,x normal )1 is the abnormal flow and normal flow determined by the output.

3. The intrusion detection system secondary diversion detection method using deep learning as claimed in claim 2, characterized in that: Set x normal >0.9 is normal flow, x abnornal <0.1 indicates abnormal traffic, and the output value range of suspicious traffic is 0.1 <x suspect <0.

9.

4. A two-level diversion detection method for an intrusion detection system using deep learning as described in claim 1, 2 or 3, characterized in that: Step 2 is as follows: 2.1 Use PatchEmbed to represent the characteristics of suspicious traffic data; the specific formula is as follows: T=PatchEmbed(X suspect ) Among them, X suspect represents the input suspicious traffic data, T represents the feature representation obtained by PatchEmbed, and PatchEmbed is the process of dividing the input traffic into uniform data blocks (or patches) and converting each patch into a low-dimensional vector; 2.2 The feature vector of each patch is processed by a multi-layer perceptron to capture local information and perform feature mixing. The specific formula is as follows: T ′ =TokenMix(T) T ″ =ChannelMix(T ′ ) Among them, Token Mix means processing the feature vector of each patch through a multi-layer perceptron (MLP), and ChannelMix means processing the feature vector after Token Mix through global average pooling and projection across channels; 2.3 The features are finally projected and represented through the fully connected layer. The specific formula is as follows: Y=FC(T ″ ) Among them, FC means that the features are finally projected and represented through a fully connected layer, and the fully connected layer has multiple hidden layers for mapping the mixed features to the final output space; 2.4 Using the MLP-Mixer secondary detection platform to suspect Classification as (x abnormal ,x normal )2, thus completing x traffic →{(x abnormal ,x normal )1,(x abnormal ,x normal )2}, classify network traffic into normal traffic and abnormal traffic; (x abnormal ,x normal )2=MLP-Mixer(x suspect )。 5. The method for secondary flow diversion detection of an intrusion detection system using deep learning as claimed in claim 4, characterized in that: Set x normal >0.9 is normal flow, x abnormal <0.1 indicates abnormal traffic, and the output value range of suspicious traffic is 0.1 <x suspect <0.9.

Citation Information

Patent Citations

  • Boundary data partitioning method and device

    CN107516101A

  • Method and system for analyzing boundary data by clustering method

    CN110851414A

  • Intrusion detection system

    CN113726810A

  • Network security intrusion detection system and detection method

    CN115865451A

  • Encrypted traffic classification method based on GNST

    CN117354012A