Intrusion detection method for interpretable fine-grained industrial control network
By adopting a deep learning model in industrial control network intrusion detection, combining byte-level, packet-level and stream-level feature extraction and fusion, the problem of low accuracy in detection of complex attacks in the prior art is solved, and a more efficient and interpretable intrusion detection effect is achieved.
Patent Information
- Application Number
- CN202510137699.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-07
AI Technical Summary
The prior art is difficult to effectively detect complex attacks in industrial control network intrusion detection, such as fragment noise filling attacks, and the detection accuracy based on rule matching and byte feature matching is low, making it difficult to update security rules in a timely manner.
A deep learning-based intrusion detection system model is adopted to extract and fusion features through three levels of processing at byte level, packet level and stream level. The byte-level module uses a self-attention mechanism to analyze the importance of each byte. The packet-level module uses a convolutional neural network to extract the spatial characteristics of the data packet. The stream-level module captures the timing characteristics of the data flow through a multi-head self-attention mechanism.
Improves the accuracy, interpretability and efficiency of intrusion detection, especially in complex attack scenarios, and can detect and respond to potential threats faster and more efficiently.
Smart Images

Figure CN119995979A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of control network intrusion detection, and in particular to an interpretable fine-grained industrial control network intrusion detection method. Background Art
[0002] Industrial Control Systems (ICS) are computerized systems used to monitor and control industrial processes. They consist of a series of hardware and software components used to manage the communication between industrial equipment, sensors and actuators, and to control the production process. ICS systems are commonly used in factories, manufacturing and production facilities to ensure the safe, efficient and reliable operation of the production process. A typical industrial control system is as follows, consisting of a control center, an engineer station, and a programmable logic controller.
[0003] The engineer station is an important part of the industrial control system, providing a graphical user interface for monitoring and managing the operation of the industrial control system. The engineer station is usually used by engineers to write, debug and modify control logic programs, and configure and maintain various parameters of the industrial control network.
[0004] A programmable logic controller (PLC) is a specialized computer control system used to control and automate physical processes in real time. Due to its widespread use in industrial control, PLCs are vulnerable to cyberattacks, so ensuring the security of PLC systems is critical to industrial control networks. PLCs have multiple input / output (I / O) ports connected to sensors and actuators to monitor and control the operation of industrial equipment.
[0005] Two typical network attacks against PLC: (used to help understand the advantages of this method over other intrusion detection solutions)
[0006] 1) Traditional control logic injection attack, this attack is performed by transmitting malicious control logic to PLC through industrial control protocol. There are two main ways to implement this attack. The first is to directly write the control logic into PLC by writing control logic command. This method can be easily discovered by detecting the command to write control logic command. The second is to write the control logic into PLC data block by writing data block command. The attacker further modifies the control flow of PLC to make PLC execute the malicious control logic in the data block. Since data block often needs to interact with control service center such as SCADA to exchange PLC status, it is difficult to detect by detecting the command to write data block. However, due to the detection strategy, it is still possible to detect the injection of malicious control logic by capturing the byte characteristics of malicious control logic.
[0007] 2) Fragmented noise filling attack is also a novel network attack that injects malicious control logic into PLC. The main idea is to divide the malicious control logic into X different data packets (fragmented), each of which contains only N bytes of the malicious control logic, and the rest are similar to 0000 or random noise bytes (noise). Through the write data block command, all bytes of the malicious control logic are injected into the PLC one by one. On the one hand, this makes the injection process of the control logic look like the exchange of normal data blocks. On the other hand, because the N bytes transmitted to the data block in each data packet are too few, it is difficult to detect by capturing the byte characteristics of the malicious control logic.
[0008] In the field of industrial control networks, for industrial control network attack detection, the rule-matching-based method requires a lot of time and manpower to analyze private protocols, which makes it difficult to update security rules in a timely manner and increases the risk of network attacks.
[0009] Byte feature matching methods have difficulty detecting attack packets containing smaller attack payloads because these packets tend to be mixed with normal packets. This type of method has low detection accuracy.
[0010] The deep packet inspection method based on CNN / CNN-RNN can only extract traffic features at the packet level, and it is difficult to effectively detect certain attack types such as fragment noise attacks;
[0011] The RNN-based packet flow detection method does not effectively utilize the byte features of the data packets, and it is difficult to detect attacks where the attack payload is concentrated in a small number of data packets (for example, some attacks are completed in one data packet). In addition, the current flow-based attack detection granularity is too large to accurately locate abnormal data packets. Summary of the invention
[0012] The technical problem to be solved by the present invention is to provide an explainable fine-grained industrial control network intrusion detection method in view of the deficiencies of the prior art.
[0013] To solve the above technical problems, the technical solution provided by the present invention is: an interpretable fine-grained industrial control network intrusion detection method, which includes a model structure part and a data processing part;
[0014] The model structure part includes a byte-level interpretable module, a packet-level feature extraction module, a flow-level timing analysis module and a result output module;
[0015] The data processing part includes byte-level, packet-level, and flow-level network models and data processing processes, and the data processing process includes data preprocessing, feature extraction, and final intrusion detection decision;
[0016] The byte level, packet level, and flow level all include network model processing and result output. Ultimately, through processing at the byte level, packet level, and flow level, the model will output its own attack judgment results respectively, and obtain the final intrusion detection conclusion by fusing these results.
[0017] Furthermore, the byte-level interpretable module includes the following contents:
[0018] A: The byte sequence of the input data packet is converted into an embedding vector, including raw byte embedding, position embedding, and segment embedding, and the importance weight of each byte is calculated through the self-attention mechanism;
[0019] B: Self-attention mechanism of shared query vector: This mechanism can calculate the contribution of each byte to the judgment of the entire data packet, improving the byte-level interpretability of the model.
[0020] Furthermore, the packet-level feature extraction module includes the following contents:
[0021] A: The byte features of the input data packet are processed through a convolutional neural network (CNN) to extract spatial features, capture the local dependencies between bytes, and then determine whether the data packet is malicious;
[0022] B: This module extracts useful features from data packets through convolution operations, combines the pooling layer to reduce redundant information, and finally outputs packet-level intrusion detection results through the fully connected layer.
[0023] Furthermore, the flow-level timing analysis module includes the following contents:
[0024] A: During stream-level processing, the model analyzes the temporal relationship between data packets through a multi-head self-attention mechanism to capture malicious disturbances in the stream.
[0025] B: This module predicts whether there is abnormal behavior in the entire data flow through the timing dependency of the traffic, and is particularly suitable for detecting advanced attack patterns.
[0026] Furthermore, the data preprocessing is specifically as follows:
[0027] A: The data flow is divided through the five-tuple to ensure the independence of each connection flow.
[0028] B: Perform flow filtering and packet filtering to remove traffic not related to the ICS system and ensure that only relevant traffic is analyzed.
[0029] C: Data packet content cleaning, zero padding and truncation operations are performed to unify the input format and adapt it to the requirements of the deep learning model.
[0030] D: A sliding window mechanism is introduced to enhance the robustness of the model to changes in attack positions by dynamically adjusting the window size and step size.
[0031] Furthermore, the feature extraction is specifically as follows:
[0032] A: Byte-level features are processed through embedding and self-attention mechanisms;
[0033] B: Packet-level features are extracted through convolutional neural networks;
[0034] C: Stream-level features capture temporal features through a multi-head self-attention mechanism.
[0035] Furthermore, flow-level and packet-level fusion: the temporal features of flow-level analysis are combined with the spatial features of packet-level analysis to ensure a balance between efficiency and accuracy.
[0036] Furthermore, the byte level complements the packet level and the flow level: the byte level provides detailed explainability, the packet level strengthens feature extraction, and the flow level provides overall timing perception. The combination of these three can ensure that the intrusion detection model has stronger attack identification capabilities.
[0037] After adopting the above structure, the present invention has the following advantages: The present invention proposes an intrusion detection system model based on deep learning, which is used for the detection of advanced attacks in ICS networks. Through byte, packet and flow-level feature extraction and fusion, the accuracy, interpretability and efficiency of intrusion detection are improved, especially in complex attack scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a partial schematic diagram of a model of an explainable fine-grained industrial control network intrusion detection method. DETAILED DESCRIPTION
[0039] The technical problems to be solved by this application proposal are:
[0040] 1) Byte-level interpretability: Through the attention mechanism, this method can analyze the contribution of each byte in the data packet to the malicious attack, thereby helping reliable intrusion detection.
[0041] ICS uses many private protocols, and the protocols between different manufacturers and even different models of equipment vary significantly. Manual analysis of protocol fields is not only time-consuming and laborious, but existing intrusion detection systems (IDS) are deficient in data packet byte-level analysis and are unable to effectively label and explain the importance of each byte in the data packet. This byte-level interpretability is crucial for security personnel to deeply analyze protocol fields, accurately locate the source of anomalies, and achieve more reliable intrusion detection.
[0042] Since ICS widely uses private protocols to ensure data confidentiality, its traffic is usually transmitted in plain text, giving it grammatical and semantic features similar to natural language. The attention mechanism in deep learning has been widely used in natural language processing, which can assign weights to each word in a sentence to highlight important vocabulary. In intrusion detection, the attention mechanism can also be applied to the analysis of ICS traffic, highlighting key bytes that may carry abnormal information by dynamically assigning weights to each byte. This mechanism not only enables the detection model to focus on important bytes more accurately, but also provides security personnel with more fine-grained explanatory information to help them better understand and respond to complex network attacks.
[0043] 2) Packet-granular detection with flow feature fusion: This method integrates flow-level features into each data packet, thus achieving packet-granular intrusion detection while ensuring detection accuracy.
[0044] In ICS, some advanced attack methods make single-packet detection difficult to work. Attackers can not only spread malicious payloads into multiple data packets, making the attack payload of each packet very small or even the same as a normal packet, but also insert one or more seemingly normal data packets at key positions in certain operation sequences by analyzing the vulnerability of the process flow. The content of these data packets may fall within the threshold range allowed by the register, making it difficult for traditional detection methods to identify them.
[0045] The data streams of these advanced attacks are usually highly sequential, which means that abnormal behavior can be perceived by analyzing the temporal relationship between packets. For example, in a fragment noise attack, the attacker breaks down the continuous binary control logic into multiple packets of write memory operations and writes them to the PLC one by one, thereby implementing the delivery of malicious control logic. Since each packet contains only a small amount of malicious code fragments, sometimes these packets are almost indistinguishable from the normal physical process packets exchanged between the HMI (human-machine interface) and the control center, making them difficult to detect in packet-level detection.
[0046] At the same time, the normal traffic of ICS networks also has significant time series characteristics. Since the business processes of ICS are usually fixed, network traffic shows a high degree of periodicity and timing. For example, sensors send data regularly and controllers issue instructions regularly, forming a stable traffic pattern. When attackers insert data packets that conform to normal operations but have abnormal timing, even if there is no problem with these operations themselves, the insertion behavior disrupts the time series characteristics of the normal data flow. This anomaly can be perceived and identified by the intrusion detection system (IDS).
[0047] Current intrusion detection systems usually rely on sequential neural networks such as RNN (recurrent neural network) to process time series data. RNN can capture the temporal dependency between packets by gradually processing the packet sequence. However, a major limitation of RNN is that it needs to accumulate sufficiently long sequence data to make effective inferences, which affects the real-time performance of detection. We hope to be able to perform real-time inference after accumulating as few packets as possible, so as to respond to potential threats more quickly.
[0048] To this end, this method selects the Multi-Head Self-Attention Mechanism. Multi-Head Self-Attention can focus on multiple different temporal relationships in ICS network traffic in parallel when processing shorter data packet sequences, thereby improving reasoning speed and accuracy. It captures key temporal features in a shorter time by assigning attention weights to different parts of the data packet sequence, achieving faster and more effective intrusion detection than RNN. Multi-Head Attention performs well in the case of short data packet sequences, significantly improving the detection accuracy, verifying the effectiveness of this method.
[0049] The present invention is further described in detail below in conjunction with the accompanying drawings.
[0050] Combined with Figure 1 ,An interpretable fine-grained industrial control network intrusion detection method includes a model structure part, a data processing part;
[0051] The model structure part includes a byte-level interpretable module, a packet-level feature extraction module, a flow-level timing analysis module and a result output module;
[0052] The byte-level interpretable module includes the following contents:
[0053] A: The byte sequence of the input data packet is converted into an embedding vector, including raw byte embedding, position embedding, and segment embedding, and the importance weight of each byte is calculated through the self-attention mechanism;
[0054] B: Self-attention mechanism of shared query vector: This mechanism can calculate the contribution of each byte to the judgment of the entire data packet, improving the byte-level interpretability of the model.
[0055] This module deeply analyzes the byte-level features of data packets by introducing the self-attention mechanism of triple embedding and shared query vectors. First, the bytes in the data packet are converted into three dimensions: raw byte embedding, position embedding, and segment embedding. Then, the self-attention mechanism is applied to calculate the importance weight of each byte to enhance the detection capability of byte-level attacks and provide highly interpretable analysis results. The key technical feature of this module is to generate the weight distribution of bytes through shared query vectors, thereby helping security personnel quickly locate attack bytes and improve their understanding and response to abnormal data packets.
[0056]
[0057] Segment Embedding:
[0058] The original bytes are distinguished from the zero-padded bytes by segment embedding, which has a dimension of 2. The segment embedding matrix is:
[0059]
[0060] Finally, after triple embedding, each byte in the data packet becomes a 3×1 vector, and the total data packet is represented as:
[0061]
[0062] Self-attention mechanism for shared query vectors: The self-attention mechanism has achieved remarkable results in the field of natural language processing. Usually, the self-attention mechanism generates weights by calculating the relationship between each byte and other bytes. However, in intrusion detection, it is more meaningful to focus on the contribution of each byte to the overall detection result of the data packet. Therefore, Avocado proposes a self-attention mechanism for shared query vectors, which generates the importance weight of each byte through the shared query vector.
[0063] Calculate the K, Q, and V matrices:
[0064] First, the embedding matrix E is transposed and input into the linear transformation matrix W K ,W Q ,W V , generate the query, key, and value matrices:
[0065] K=E T W K ,Q=E T W Q ,V=E T W V
[0066] in d is the embedding dimension, which represents the hidden layer dimension of the attention matrix.
[0067] Shared query vector generation:
[0068] To obtain a shared query vector, the model performs a mean operation on the query matrix Q along the byte dimension to generate a packet-level query vector:
[0069]
[0070] in is the mean query vector.
[0071] Shared query vector weight calculation:
[0072] The shared query vector Q packet Multiply it with the key-value matrix K and normalize it through the Softmax function to generate the weight of each byte:
[0073]
[0074] in Indicates the weight of each byte, d k is the K matrix embedding dimension.
[0075] Output calculation:
[0076] The byte attention weight W B Apply element by element to the value matrix V and calculate the weighted byte representation:
[0077]
[0078] Then, after a linear change W O The output matrix is the same as the original input E T Perform residual connection to get the final byte-level output:
[0079] B=ZW O +E T
[0080] in Keep the same shape as the original input
[0081] The packet-level feature extraction module includes the following contents:
[0082] A: The byte features of the input data packet are processed through a convolutional neural network (CNN) to extract spatial features, capture the local dependencies between bytes, and then determine whether the data packet is malicious;
[0083] B: This module extracts useful features from data packets through convolution operations, combines the pooling layer to reduce redundant information, and finally outputs packet-level intrusion detection results through the fully connected layer.
[0084] The flow-level timing analysis module includes the following contents:
[0085] A: During stream-level processing, the model analyzes the temporal relationship between data packets through a multi-head self-attention mechanism to capture malicious disturbances in the stream.
[0086] B: This module predicts whether there is abnormal behavior in the entire data flow through the timing dependency of the traffic, and is particularly suitable for detecting advanced attack patterns.
[0087] The data processing part includes byte-level, packet-level, and flow-level network models and data processing processes, and the data processing process includes data preprocessing, feature extraction, and final intrusion detection decision;
[0088] The byte level, packet level, and flow level all include network model processing and result output. Ultimately, through processing at the byte level, packet level, and flow level, the model will output its own attack judgment results respectively, and obtain the final intrusion detection conclusion by fusing these results.
[0089] The byte-level processing described:
[0090] In the byte-level processing stage, we focus on each byte in the data packet. The goal of this stage is to analyze the contribution of each byte in attack detection by introducing the attention mechanism in deep learning, so as to ensure that the specific bytes of malicious attacks can be accurately located. At this level, the model weights the features of each byte to dynamically highlight those bytes related to the attack. In this way, the model can provide byte-level interpretability to help security personnel accurately understand the potential attack behaviors in abnormal data packets.
[0091] A: Network model processing: For each byte, the model first embeds it into a vector representation, and then uses a self-attention mechanism to calculate the correlation between bytes to evaluate the impact of each byte on malicious activities.
[0092] B: Output: The output of this process is the weight of each byte and the corresponding classification result, which helps to analyze which bytes play an important role in the entire data packet.
[0093] Packet level processing:
[0094] In the packet-level processing stage, we extend our focus to the level of individual packets. This level of detection includes not only byte-level features, but also certain timing features. By using convolutional neural networks (CNNs), the model is able to extract spatial features and temporal dependencies within packets to identify more complex attack patterns, especially in some advanced attacks (such as fragmented noise attacks) where malicious payloads are dispersed across multiple packets.
[0095] A: Network model processing: The convolution layer is used to extract the spatial features between bytes in the data packet, and the maximum pooling operation is used to reduce redundant information. Then, the fully connected layer is used for classification to determine whether the packet is a malicious packet.
[0096] B: Output results: Packet-level output is the attack judgment result for each data packet, marking whether each data packet has abnormal behavior.
[0097] This module is designed to perform deep feature extraction on each data packet. Since each data packet in the ICS network has complex correlations in space and time, CNN can effectively capture these local features, especially the dependencies between bytes in the data packet. Using CNN to extract spatial features can not only improve the representation of data packets, but also reduce the number of parameters through its weight sharing mechanism, thereby reducing computational complexity and meeting the requirements of ICS real-time detection.
[0098] Input transformation: Before performing the convolution operation, the byte embedding matrix processed by the Byte-Level Interpretable Module needs to be Perform dimension conversion to adapt it to the CNN input format. First, perform the Permute operation to change the input dimension to The 3 is derived from the original byte embedding dimension, which represents the original number of channels of the input convolution, and 1 is the virtual height dimension (used to maintain 1D convolution).
[0099] One-dimensional convolution layer: Next, the data is subjected to two one-dimensional convolution operations to extract spatial local features. Assuming that C1 and C2 are the number of output channels of the convolution kernel, k1 and k2 are the sizes of the convolution kernel, s1 and s2 are the convolution steps, the results X1 and X2 after the two convolutions are expressed as:
[0100] X1=Conv1D(B′,C1,k1,s1)
[0101] X2=Conv1D(X1,C2,k2,s2)
[0102] Max pooling layer: The output after convolution is further extracted through the max pooling layer (Max Pool). The max pooling operation reduces redundant information by reducing the size of the output feature map and improves the computational efficiency of the model. The result F after pooling is expressed as:
[0103] X=MaxPool(X2)
[0104] This operation helps to reduce the dimensionality of features while retaining the most salient features.
[0105] Feature flattening layer: The pooled features are converted into one-dimensional vectors through a flattening operation to generate packet-level feature representations:
[0106] P = Flatten(X)
[0107] Through this module, the model uses a convolutional neural network to extract local spatial features in each data packet, then effectively reduces the data dimension through maximum pooling, and finally flattens it into a one-dimensional vector for stream feature fusion. The convolution layer captures the spatial dependencies in the byte sequence, and the combination of the pooling layer and the flattening operation provides a compact and high-dimensional feature representation for the subsequent stream-level feature fusion module. This extraction method not only improves the expressiveness of the features, but also takes into account the computational efficiency.
[0108] Stream-level processing:
[0109] Stream-level processing focuses on the timing characteristics of the entire traffic and the dependencies across packets. At this stage, the multi-head self-attention mechanism is used to process the time series characteristics in the data stream. By analyzing the timing relationship between packets, the model can capture the attacker's malicious manipulation of the timing, such as disrupting normal industrial control traffic by inserting forged packets.
[0110] A: Network model processing: The multi-head self-attention mechanism focuses on multiple temporal relationships in parallel, so as to analyze important flow features in a shorter time. This stage is particularly suitable for advanced attacks with significant temporal features, such as fragment noise filling attacks.
[0111] B: Output results: The flow-level output is a prediction of the attack possibility of the entire flow, which is summarized based on the judgment results at the packet level to ultimately determine whether the entire flow is abnormal.
[0112] The goal of this module is to enhance the model's ability to detect advanced industrial control network attacks by integrating the timing features between data packets. Single-packet detection is often ineffective when facing advanced attacks where malicious payloads are dispersed or embedded in normal traffic, and the timing of data flows is an important basis for capturing such attacks. Attackers can not only disperse malicious code to make the payload of each data packet as small as a normal packet, but also insert seemingly normal packets at key locations in certain process flows, making it difficult for traditional methods to accurately identify attack behaviors.
[0113] In order to effectively capture the temporal characteristics of these advanced attacks, this paper adopts a multi-head self-attention mechanism in the Flow-Level Feature Fusion Module. The input of this part is a set of data packet feature sequences extracted by packet-level features. Assume that a group has a total of n consecutive data packets, and the hidden layer dimension of each data packet is dp , then input
[0114] The self-attention mechanism calculates the correlation through the query, key, and value matrices. Unlike 4.4.1, each data packet has an independent query vector, which is combined into the query matrix Q. Self-attention calculates the correlation between each data packet, so that each data packet can integrate the previous and next context information and capture the potential timing pattern. This mechanism helps the model integrate the flow features between data packets and integrate the timing features in the data stream into the features of each data packet, thereby effectively identifying complex attack behaviors.
[0115] The query, key, and value matrices are calculated as follows:
[0116] K=FW K ,Q=FW Q ,V=FW V
[0117] Where W K , W Q , W V is a learnable weight matrix.
[0118] Self-attention is calculated by calculating the weighted feature matrix Z between data packets, which is calculated as follows:
[0119]
[0120] Here k It is the hidden layer dimension of the key vector, which is used for normalization to ensure that the inner product result is not too large.
[0121] Each Z represents an attention head. The multi-head attention mechanism further improves the model's ability to cope with complex attack scenarios by computing multiple independent attention heads in parallel, with each head observing the timing relationship between data packets from a different perspective. Each attention head can independently focus on different timing levels and features, thereby providing the model with more diverse information. Finally, the outputs of the multi-head attention are combined through a concatenation operation to form a new feature representation, which is then residually connected to the original input F. The calculation formula is as follows:
[0122] F′=MultiHead(F)+F=Concat(Z1,…,Z h )W O +F
[0123] in, The shape is consistent with the original input F, h is the number of attention heads, W O is a linearly varying weight matrix.
[0124] The result output module is specifically:
[0125] The goal of this module is to generate a feature matrix based on the flow-level feature fusion module. Count each packet And finally output the intrusion detection classification probability label of each data packet.
[0126] In order to classify the feature vector of each data packet, it is first necessary to further perform nonlinear mapping and feature extraction on the data packet features through a set of fully connected neural networks (Feedforward Neural Network, FFNN). The FFNN module can further enhance the expressiveness of the model and help the model capture deeper feature relationships.
[0127] Fully connected network:
[0128] The feature vector P' of each data packet will pass through several fully connected layers in sequence. The output of the kth fully connected layer can be expressed as:
[0129] H k =σ(H k-1 W k +b k )
[0130] Among them, H k-1 is the output of the previous fully connected layer, is the weight matrix of the kth fully connected layer, b k is the bias term, and σ is the activation function. The first fully connected layer receives the input F' from the feature fusion, and the output dimension of the last fully connected layer will be the same as the number of categories of the classification task.
[0131] Softmax:
[0132] After being processed by several fully connected layers, the final output will pass through the Softmax activation function to map the feature vector of each data packet to the classification probability vector. The function of the Softmax function is to convert the output of the FFNN into a probability distribution. Its calculation formula is:
[0133]
[0134] where z j represents the score of the i-th category, and C represents the number of categories. After Softmax, the output is a vector for the classification of each data packet P' Where C is the number of categories in the classification task.
[0135] Final classification results:
[0136] For each data packet, the classification result can be obtained by selecting the category with the highest classification probability as the final predicted label:
[0137]
[0138] in, is the predicted classification label.
[0139] In addition, in Section 4.4.1 is the interpretable weight of each byte of the packet.
[0140] Comprehensive processing and intrusion detection results:
[0141] Finally, through the processing of byte level, packet level and flow level, the model will output the attack judgment results respectively, and obtain the final intrusion detection conclusion by integrating these results. In practical applications, this multi-level analysis can effectively improve the accuracy and robustness of detection, especially in complex attack scenarios.
[0142] Flow-level and packet-level fusion: Combines the temporal features of flow-level analysis with the spatial features of packet-level analysis to ensure a balance between efficiency and accuracy.
[0143] Complementarity between byte level, packet level and flow level: The byte level provides detailed explainability, the packet level strengthens feature extraction, and the flow level provides overall timing perception. The combination of these three can ensure that the intrusion detection model has stronger attack identification capabilities.
[0144] The data preprocessing is specifically as follows:
[0145] A: The data flow is divided through the five-tuple to ensure the independence of each connection flow.
[0146] B: Perform flow filtering and packet filtering to remove traffic not related to the ICS system and ensure that only relevant traffic is analyzed.
[0147] C: Data packet content cleaning, zero padding and truncation operations are performed to unify the input format and adapt it to the requirements of the deep learning model.
[0148] D: A sliding window mechanism is introduced to enhance the robustness of the model to changes in attack positions by dynamically adjusting the window size and step size.
[0149] The feature extraction is specifically as follows:
[0150] A: Byte-level features are processed through embedding and self-attention mechanisms;
[0151] B: Packet-level features are extracted through convolutional neural networks;
[0152] C: Stream-level features capture temporal features through a multi-head self-attention mechanism.
[0153] The present invention and its implementation methods are described above, and such description is not restrictive, and the actual structure is not limited thereto. In short, if a person skilled in the art is inspired by it, and does not deviate from the purpose of the invention, and does not creatively design a structure and implementation method similar to the technical solution, they should all fall within the protection scope of the present invention.
[0154] The specific contents of data preprocessing are as follows:
[0155] Data preprocessing is a key step to convert raw network traffic into data suitable for deep learning model input. In order to improve the accuracy and efficiency of detection, this paper designs a complete set of preprocessing processes, which includes six steps, from data flow diversion to the final processing of data format:
[0156] Step 1: Divert the data stream according to the five-tuple information to uniquely identify each network connection. The communication traffic between different devices in the industrial control network (such as PLC and HMI) usually has its own unique characteristics. Therefore, these different network connections can be effectively distinguished by five-tuple diversion. This method helps to improve detection accuracy because it allows for more detailed analysis of the traffic on each independent connection, especially in the process of traffic anomaly detection, which can more effectively identify abnormal behavior patterns.
[0157] Step 2: Perform flow filtering to retain communication traffic related to specific industrial control devices. Network traffic in industrial control systems usually follows specific communication patterns, such as instruction sets and communication cycles between specific devices and controllers. By filtering out traffic that is not related to ICS security, the dimension and complexity of the data can be significantly reduced, ensuring that the model focuses on the most relevant communication traffic. This step can reduce the incidence of false positives and improve detection efficiency.
[0158] Step 3: Packet filtering. Certain types of packets (such as DNS, SYN, ACK, FIN, etc.) are usually used to establish and maintain network connections rather than actual data transmission. Although some attacks may be hidden in these packets, in the attack scenario of this article, filtering out these packets helps model longer control flows and improves the detection capabilities of key communication flows. This strategy needs to be adjusted according to the actual attack scenario to ensure that important attack features are not filtered out.
[0159] Step 4: Clean up the packet content, which mainly includes masking the IP address and port information, and removing the Ethernet header. The purpose of this is to eliminate metadata that is not related to intrusion detection and retain only the payload information. In this way, the model avoids over-reliance on specific information about the network configuration, and instead focuses on learning the actual content features in the packet, thereby improving the generalization ability of the model and enabling it to operate effectively in different network environments.
[0160] Step 5: truncate and fill with zeros to ensure that the format of data input is uniform. For data packets that are too long, only the key parts are retained; for data packets that are insufficient in length, they are filled with zeros. Deep learning models have high requirements for the format of input data, and fixed data length is crucial for batch processing and video memory management. A uniform input size not only ensures the stability of the model training process, but also avoids computational complexity problems caused by inconsistent input lengths.
[0161] Step 6: Sliding window preprocessing scheme. This scheme gradually slides the window on the flow data and groups the processed data packets according to the window size N. Each time, N consecutive data packets are selected from the flow as input samples, and the window is slid with a step size s to generate new input samples. This method brings many advantages: First, since the industrial control system (ICS) data set is usually small, the sliding window technology can significantly expand the number of training samples, thereby improving the generalization ability of the model. Second, for single-packet attacks, the attack data packet may appear at any position in the flow. The simple grouping method may cause the position of the attack packet in the window to be fixed, so that the detection result depends on the specific position of the attack data packet, which in turn affects the detection performance of the model. The sliding window technology can avoid this problem because it allows samples to be trained in different positions and contexts, improving the robustness to the change of the attack packet position. Finally, in the real-time reasoning process, the sliding window technology enables the system to slide quickly in the traffic without waiting for the full window to be filled. The system only needs to wait for the data packet corresponding to the sliding distance to infer the classification result of the new data packet in real time, thereby significantly improving the detection efficiency.
[0162] Through this preprocessing solution, the original network traffic is successfully converted into a format that can be processed by the deep learning model, ensuring that key feature information related to ICS security is retained. This series of steps complement each other, not only significantly improving the accuracy and efficiency of detection, but also ensuring the applicability of the model in various complex industrial control network environments. In specific applications, especially the packet filtering and flow filtering steps, it may be necessary to adjust and optimize according to the actual attack scenario to ensure the best detection effect.
[0163] After adopting the above method, this application proposal has the following advantages:
[0164] 1. Byte-level interpretability:
[0165] Compared with rule-matching-based methods, it can automatically analyze private protocols. Compared with other deep learning-based methods, it achieves byte interpretability in the field of industrial control intrusion detection, saving manpower analysis costs while achieving reliable intrusion detection.
[0166] 2. Detection accuracy:
[0167] Compared with the deep packet inspection method based on CNN / CNN-RNN, the flow-level features are integrated into each data packet, so that each data packet can perceive the flow-level context, thereby effectively responding to single-packet attacks caused by fragment noise attacks, simulated positive atomic operations, and tampering with normal process operation sequences. Effectively improve the detection accuracy of these network attacks.
[0168] 3. Detection of particle size:
[0169] Compared with the RNN-based packet flow detection method, on the one hand, due to insufficient extraction of packet features, the accuracy of single-packet attacks is low. On the other hand, although the flow-level detection scheme can detect single-packet attacks caused by fragment noise attacks, simulated atomic operations, and tampering with normal process operation sequences, the flow-granularity detection cannot locate which specific data packets have abnormalities, which causes inconvenience for security personnel in subsequent analysis.
Claims
1. An interpretable fine-grained industrial control network intrusion detection method, characterized by: It includes model structure part and data processing part; The model structure part includes a byte-level interpretable module, a packet-level feature extraction module, a flow-level timing analysis module and a result output module; The data processing part includes byte-level, packet-level, and flow-level network models and data processing processes, and the data processing process includes data preprocessing, feature extraction, and final intrusion detection decision; The byte level, packet level, and flow level all include network model processing and result output. Ultimately, through processing at the byte level, packet level, and flow level, the model will output its own attack judgment results respectively, and obtain the final intrusion detection conclusion by fusing these results.
2. The interpretable fine-grained industrial control network intrusion detection method according to claim 1 is characterized by: The byte-level interpretable module includes the following contents: A: The byte sequence of the input data packet is converted into an embedding vector, including raw byte embedding, position embedding, and segment embedding, and the importance weight of each byte is calculated through the self-attention mechanism; B: Self-attention mechanism of shared query vector: This mechanism can calculate the contribution of each byte to the judgment of the entire data packet, improving the byte-level interpretability of the model.
3. The interpretable fine-grained industrial control network intrusion detection method according to claim 1 is characterized by: The packet-level feature extraction module includes the following contents: A: The byte features of the input data packet are processed through a convolutional neural network (CNN) to extract spatial features, capture the local dependencies between bytes, and then determine whether the data packet is malicious; B: This module extracts useful features from data packets through convolution operations, combines the pooling layer to reduce redundant information, and finally outputs packet-level intrusion detection results through the fully connected layer.
4. The interpretable fine-grained industrial control network intrusion detection method according to claim 1 is characterized by: The flow-level timing analysis module includes the following contents: A: When processing at the flow level, the model analyzes the temporal relationship between data packets through a multi-head self-attention mechanism to capture malicious disturbances in the flow; B: This module predicts whether there is abnormal behavior in the entire data flow through the timing dependency of the traffic, and is particularly suitable for detecting advanced attack patterns.
5. The interpretable fine-grained industrial control network intrusion detection method according to claim 1 is characterized by: The data preprocessing is specifically as follows: A: The data stream is divided through the five-tuple to ensure the independence of each connection stream; B: Perform flow filtering and packet filtering to remove traffic not related to the ICS system and ensure that only relevant traffic is analyzed; C: Data packet content cleaning, zero padding and interception operations make the input format unified and adapt to the requirements of deep learning models; D: A sliding window mechanism is introduced to enhance the robustness of the model to changes in attack locations by dynamically adjusting the window size and step size.
6. The interpretable fine-grained industrial control network intrusion detection method according to claim 1 is characterized by: The feature extraction is specifically as follows: A: Byte-level features are processed through embedding and self-attention mechanisms; B: Packet-level features are extracted through convolutional neural networks; C: Stream-level features capture temporal features through a multi-head self-attention mechanism.
7. The interpretable fine-grained industrial control network intrusion detection method according to claim 1 is characterized by: Flow-level and packet-level fusion: Combines the temporal features of flow-level analysis with the spatial features of packet-level analysis to ensure a balance between efficiency and accuracy.
8. The interpretable fine-grained industrial control network intrusion detection method according to claim 1 is characterized by: Complementarity between byte level, packet level and flow level: The byte level provides detailed explainability, the packet level strengthens feature extraction, and the flow level provides overall timing perception. The combination of these three can ensure that the intrusion detection model has stronger attack identification capabilities.
Citation Information
Patent Citations
Encrypted application traffic classification method and system based on local-global feature attention
CN116827873A
Incremental network traffic classification method and system based on pre-training representation
CN119030934A
CNN-GRU industrial control network attack detection method based on data packet and flow level feature fusion
CN119363386A
Removable emergency notification device
KR102773466B1
Cited By
Network intrusion detection method, system, device, medium and product
CN120811688A
Abnormal network intrusion detection system based on FPGA and artificial intelligence
CN121098574A
Intrusion detection method and system based on hierarchical interpretable behavior sequence modeling, and storage medium
CN121262000A
An intrusion detection method and system based on hierarchical interpretable behavior sequence modeling and a storage medium
CN121262000B
Ethernet intrusion detection method and system
CN121441558A