Deep detection method of abnormal traffic in industrial networks based on unsupervised learning

Through unsupervised learning recurrent neural network and autoencoder method, the detection problem of advanced unknown threats in industrial networks is solved, and high-precision abnormal traffic detection is achieved, which is suitable for industrial network security protection.

CN115766193BActive Publication Date: 2025-08-19ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211411407.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-11
Publication Date
2025-08-19
Estimated Expiration
2042-11-11

AI Technical Summary

Technical Problem

The existing industrial network traffic detection methods are difficult to effectively detect advanced unknown threats with high concealment and long incubation periods, and the unsupervised classification model is easy to overfit and has low detection accuracy.

Method used

Using an unsupervised learning method, the load word segmentation and autoencoder classification is performed through recurrent neural networks, and the loss function threshold interval is calculated using BiLSTM encoder and decoder to realize unsupervised load word segmentation and classification.

Benefits of technology

Without destroying semantic information, the accuracy of industrial network traffic detection is improved, the problem of malicious traffic and normal traffic imbalance is solved, and efficient abnormal traffic detection is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115766193B_ABST
    Figure CN115766193B_ABST
Patent Text Reader

Abstract

This paper proposes a method for deep detection of abnormal traffic in industrial networks based on unsupervised learning. The method first extracts the payload from the traffic data packet and converts it into corresponding characters according to the ASCII code to obtain a string. The string is segmented using an unsupervised word segmentation algorithm based on a recurrent neural network, and the segmentation result is output. The segmented string is converted into a vector and input into an unsupervised classification algorithm based on an autoencoder. The vector is encoded using an encoder based on a bidirectional long short-term memory network (BiLSTM) and the encoded value is output. The encoded value is decoded using a decoder based on the BiLSTM. The autoencoder is optimized by subtracting the encoder input from the decoder output as a loss function. The loss function threshold interval of normal traffic is obtained to complete the deep detection of abnormal traffic in industrial networks. The method provides a specific algorithm description using real network traffic data and obtains experimental results through a series of experiments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an industrial network traffic anomaly detection method, and in particular to an industrial network abnormal traffic in-depth detection method based on unsupervised learning. Background Art

[0002] Industrial networks provide a path for the digitalization, networking, and intelligent development of industry and even sectors. They are a key cornerstone of the Fourth Industrial Revolution and a vital guarantee for the stable operation of critical infrastructure. With the development of IoT technology and the widespread use of large-scale sensors in industrial control systems, industrial networks face significant risks from a loss of security perimeters and an increase in unknown threats. Therefore, improving the security of industrial networks is a critical issue that needs to be addressed.

[0003] Traffic detection is a key technology for preventing network attacks and unknown malicious threats, and is widely used in industrial network security. Currently, there are two main types of traffic detection methods: flow feature-based detection and payload-based detection.

[0004] Detection methods based on flow characteristics (such as the number of packets transmitted per second and the number of bytes per second) can determine whether traffic is normal or abnormal. However, this method can only detect attacks that significantly affect flow characteristics, such as denial of service (DoS) attacks and Nmap port scans. However, advanced unknown threats (APTs) have a particularly long incubation period, are highly concealed, and have little impact on flow characteristics, making this method difficult to detect.

[0005] Regardless of the stealthiness of an attack or the duration of its incubation period, the attacker's malicious behavior is always exposed in the packet payload. Therefore, payload-based detection methods are widely used in traffic detection, primarily consisting of word segmentation models and classification models. Word segmentation models analyze the strings within the payload and then, using methods such as sliding windows, break these strings into words. For example, if the original string is (thepointvalueis4), the resulting string is (thepointvalueeis4). However, existing word segmentation methods often destroy the semantic information within the payload, resulting in low detection accuracy for classification models. Therefore, developing a word segmentation method that preserves semantic information is crucial for improving the detection performance of classification models. In the field of natural language processing, word segmentation methods are divided into two categories: supervised and unsupervised. Supervised word segmentation requires a semantic library with pre-defined word segments as labels for training. However, a pre-defined semantic library for payload segments currently does not exist, making supervised methods unsuitable. Therefore, unsupervised word segmentation methods for payloads remain the only approach. Classification models can also use both unsupervised and supervised learning methods. However, the extreme imbalance between malicious and normal traffic samples in industrial networks makes supervised classification models prone to overfitting and poor performance. Therefore, studying how to perform unsupervised training of classification models using only normal samples is a critical challenge that must be addressed. Summary of the Invention

[0006] The purpose of this invention is to solve the problem of in-depth detection of abnormal industrial network traffic based on unsupervised learning. This paper addresses the research deficiencies in the field of industrial network traffic anomaly detection by proposing a comprehensive analytical method. The proposed payload segmentation and classification methods based on unsupervised learning provide guidance for solving the problems of payload segmentation and payload feature extraction performance.

[0007] The purpose of the present invention can be achieved by the following technical solution: A method for deep detection of abnormal traffic in industrial networks based on unsupervised learning, comprising the following steps:

[0008] (1) Obtain the data packet of industrial network traffic, extract the hexadecimal-encoded payload in the data packet, and convert the payload into corresponding characters according to the ASCII code to obtain a string;

[0009] (2) Setting the maximum segmentation length, dividing the string into segments of different lengths, and then dividing the segments into sub-segments of different lengths, where the maximum length of the segments and sub-segments does not exceed the set value;

[0010] (3) Input these sub-segments of different lengths into the recurrent neural network to obtain the corresponding occurrence probabilities;

[0011] (4) Based on the occurrence probabilities of sub-segments of different lengths, the occurrence probabilities of all word combinations of the entire string are calculated, and the sum is taken as the occurrence probability of the string; the recurrent neural network is trained using this occurrence probability as the loss function, and a trained recurrent neural network model is obtained;

[0012] (5) The sub-segments of different lengths obtained in step (2) are input into the trained recurrent neural network model again to obtain the corresponding occurrence probabilities; based on the occurrence probabilities of the sub-segments of different lengths, the word segmentation method is calculated to maximize the occurrence probability of the entire string, that is, the maximum probability path, and the string with the segmented words is obtained;

[0013] (6) Embed the word string into words, input it into the autoencoder, encode it through the encoder, and reconstruct it through the decoder. The difference between the encoder input and the decoder output is calculated as the loss function to train the autoencoder, obtain the loss function set of normal samples, and establish the loss function threshold range;

[0014] (7) The traffic data packet to be detected is input into the trained recurrent neural network model, and the trained recurrent neural network model outputs the word segmentation result; the word segmentation result is then input into the trained autoencoder to obtain the loss function value; if the loss function value of the current traffic data packet is within the loss function threshold range, the traffic data packet is judged to be normal; otherwise, it is judged to be abnormal.

[0015] As a preferred embodiment of the present invention, step (5) is specifically:

[0016] (5.1) Input the sub-segments of different lengths obtained in step (2) into the trained recurrent neural network model again to obtain the corresponding occurrence probability P_Z i ;

[0017] (5.2) Based on the sub-fragments of different strings, calculate how to segment the words so that the probability of the entire string is maximized, that is, the maximum probability path, and output the segmentation results. The formula is as follows:

[0018] S_P k =Combine{(P_Z1;P_Z2;…;P_Z I ;…;P_Z len(X)-SM+1 ) k}

[0019] Max_P=Max(S_P k )(k=1,2,…,M)

[0020] out_X=(S1,S2,…,S N )

[0021] Where Max_P is the maximum value in the probability set obtained by all word segmentation combinations; out_X is the output word segmentation result, the first sub-segment is S1, and the last one is S N .

[0022] As a preferred embodiment of the present invention, step (6) is specifically:

[0023] (6.1) The word-embedded string is embedded into the word vector and input into the autoencoder. The encoder is used to encode it and the decoder is used to decode it. The calculation formula is as follows:

[0024] X_V=Embedding(out_X)

[0025] E_V=E_BiLSTM(X_V)

[0026] D_V=D_BiLSTM(E_V)

[0027] Where X_V is the word embedding vector, E_BiLSTM() is the encoder based on BiLSTM; E_V is the output of the encoder; D_BiLSTM() is the decoder based on BiLSTM; D_V is the output of the decoder;

[0028] (6.2) Subtract the encoder input from the decoder output and use this as the loss function value to optimize the autoencoder. The loss function threshold range for normal traffic data packets is established. The calculation formula is as follows:

[0029] A_loss=D_V-X_V

[0030] T max =max(A_loss')

[0031] T min =min(A_loss')

[0032] T=(T min ,T max )

[0033] Where A_loss is the difference between the encoder input and the decoder output, and A_loss' is the loss function value obtained by inputting the training set of normal samples into the trained autoencoder; T max is the maximum value of A_loss', T max is its minimum value; T is the threshold interval of the loss function.

[0034] As a preferred embodiment of the present invention, step (7) is specifically:

[0035] Extract the payload from the traffic data packet to be detected and convert it into corresponding characters to obtain a string; divide the string into sub-segments of different lengths; input these sub-segments of different lengths into the trained recurrent neural network model to obtain a string with good word segmentation; then input the word segmentation result into the trained autoencoder to obtain its loss function value, and determine whether it is within the loss function threshold range; if so, it is judged as normal traffic; otherwise, it is judged as abnormal.

[0036] The beneficial effects of the present invention are:

[0037] (1) This paper proposes an unsupervised payload segmentation algorithm that completes the segmentation task without any prior knowledge without destroying the payload semantics, providing a high-quality data foundation for subsequent classification models.

[0038] (2) This paper proposes an autoencoder classification algorithm based on unsupervised learning. This algorithm trains the autoencoder using only normal traffic packets. The difference between the encoder input and the decoder output is calculated as the loss function. A threshold range for the loss function is established for normal samples to determine whether a traffic packet is abnormal. This method can address the extreme imbalance between malicious and normal traffic data in industrial networks. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is a flow chart of the method of the present invention;

[0040] Figure 2 This is the classification result diagram under different SM values;

[0041] Figure 3 This is a comparison chart of the results of different traffic anomaly detection methods. DETAILED DESCRIPTION

[0042] The present invention will be described in detail below with reference to the accompanying drawings, and the objects and effects of the present invention will become more apparent.

[0043] The industrial network traffic anomaly detection method based on unsupervised learning proposed in this paper can be referred to Figure 1 , mainly including the following steps:

[0044] Step 1: Obtain industrial network traffic data packets and extract their payloads. The calculation formula is as follows:

[0045] X'=extract(.pcap)

[0046] Where pcap is the collected industrial network traffic data packet, the extract() function can extract the payload in the data packet; X' is the extracted payload;

[0047] And convert the ASCII code into the corresponding character to get the string; the calculation formula is as follows:

[0048] X=ASCII(X')

[0049] Where, the ASCII() function converts the hexadecimal payload to the corresponding character; X is the converted string.

[0050] Step 2:

[0051] Set the maximum segmentation length SM and divide the string into segments of different lengths. The calculation formula is as follows:

[0052] Y={X 1~SM ;X 2~SM+1 ;…;X i~i+SM-1 ;…;X (len(X)-SM+1)~len(X)}

[0053] Y i =X i~i+SM-1

[0054] Where Y is the set of string segments divided from X, X i~i+SM-1 Represents a string segment consisting of the i-th character to the i+SM-1-th character in the string X; len(X) represents the length of the string,

[0055] Then divide each segment into sub-segments, and the calculation formula is as follows

[0056]

[0057]

[0058] Where Z i Represents X i~i+SM-1 The sub-fragment set of Z i j Indicates that in string X i~i+SM-1 Take the sub-segment consisting of the j-th to SM-th characters.

[0059] The maximum length of a segment and sub-segment does not exceed the set value; a segment can be considered as a word, and a string as a sentence;

[0060] Step 3: Sub-fragment set Z i For word embedding, the calculation formula is as follows

[0061] E_Z i =Word_Embedding(Z i )

[0062] Where, the Word_Embedding() function can convert a string sub-segment into a vector; E_Zi It's Z i The corresponding vector;

[0063] Then input the vector into the recurrent neural network (RNN) model to obtain the corresponding probability value. The calculation formula is as follows:

[0064] P_Z i =RNN(E_Z i )

[0065] Where, the RNN() function can calculate the probability of each string sub-segment; P_Z i It's Z i The corresponding probability.

[0066] Step 4: Based on the probabilities of words of different lengths, calculate the probability of all word combinations in the entire sentence and sum them as the probability of the sentence. Use this probability as the loss function to train the RNN and obtain the trained RNN model.

[0067] According to the occurrence probability of words of different lengths, the occurrence probability of all word combinations in the entire sentence is calculated and summed up. The formula is as follows:

[0068] S_P k =Combine{(P_Z1;P_Z2;…;P_Z I ;…;P_Z len(X)-SM+1 ) k}

[0069] Sum_P=Sum(S_P k )(k=1,2,…,M)

[0070] Where S_P k is the probability value of the string after using the k-th word combination: (P_Z1; P_Z2; ...; P_Z I ;…;P_Z len(X)-SM+1 ) k is the kth word combination of the entire sentence; the Combine{} function can calculate the probability value of the entire string of the kth word combination; Sum_P is the sum of the probability values of all word combinations in the entire string; Sum() can sum the probabilities, and M means there are M word combinations in total;

[0071] (4.2) The sum of the probabilities of all word combinations in the entire string is used as the loss function to optimize the recurrent neural network model. The formula is as follows:

[0072] Loss = -log(Sum_P)

[0073] In the formula, the loss function takes the opposite of the logarithm of Sum_P

[0074] Step 5: Input the sub-segments of different lengths obtained in step 2) into the trained recurrent neural network model again to obtain the corresponding probability value P_Z i ;

[0075] Based on the sub-fragments of different strings, calculate how to segment the words so that the probability of the entire string is maximized, that is, the maximum probability path, and output the word segmentation results. The formula is as follows:

[0076] S_P k =Combine{(P_Z1;P_Z2;…;P_Z I ;…;P_Z len(X)-SM+1 ) k}

[0077] Max_P=Max(S_P k )(k=1,2,…,M)

[0078] out_X=(S1,S2,…,S N )

[0079] Where Max_P is the maximum value in the probability set obtained by all word segmentation combinations; out_X is the output word segmentation result, the first sub-segment is S1, and the last one is S N .

[0080] Step 6: Embed the word string into words, input it into the autoencoder, encode it using the encoder, and then decode it using the decoder. The calculation formula is as follows:

[0081] X_V=Embedding(out_X)

[0082] E_V=E_BiLSTM(X_V)

[0083] D_V=D_BiLSTM(E_V)

[0084] Where X_V is the word embedding vector, E_BiLSTM() is the encoder based on BiLSTM; E_V is the output of the encoder; D_BiLSTM() is the decoder based on BiLSTM; D_V is the output of the decoder;

[0085] Subtract the encoder input from the decoder output and use this as the loss function value to optimize the autoencoder. Then, establish the loss function threshold range for normal traffic data packets. The calculation formula is as follows:

[0086] A_loss=D_V-X_V

[0087] T max =max(A_loss')

[0088] T min =min(A_loss')

[0089] T=(T min ,T max )

[0090] Where A_loss is the difference between the encoder input and the decoder output, and A_loss' is the loss function value obtained by inputting the training set of normal samples into the trained autoencoder; T max is the maximum value of A_loss', T max is its minimum value; T is the threshold interval of the loss function.

[0091] Step 7: Extract the payload from the traffic data packet to be tested and convert it into corresponding characters to obtain a string. Divide the string into sub-segments of different lengths. Input these sub-segments of different lengths into the trained RNN model to obtain a string with segmented words. Then input the segmentation result into the trained autoencoder to obtain its loss function value and determine whether it is within the loss function threshold range. If the loss function of the current traffic data packet is within the threshold range, the data packet is judged to be normal; otherwise, it is judged to be abnormal.

[0092] In a specific implementation of the present invention, each step of the present invention is described in more detail.

[0093] First, malicious traffic and normal traffic from the industrial network are collected. The normal traffic is divided into a training set and a test set at a ratio of 4:1, and all malicious traffic is divided into the test set, as shown in Table 1:

[0094] training set Test set Malicious traffic - 500 Normal traffic 64000 8000

[0095] The numbers in Table 1 represent the number of collected .pcap packets. For example, in the training set, the number of normal traffic packets is 64,000. Since this model is trained only on normal traffic, the training set does not contain malicious traffic.

[0096] Secondly, use the plug-in in Wireshark to extract the payload in the data packet and convert it into corresponding characters according to the ASCII code to obtain a string.

[0097] Next, set the maximum length of the word segmentation segment to divide the string into segments and sub-segments; the value set here is 4, 6, 8, and 10, a total of 4 different cases.

[0098] The RNN-based unsupervised word segmentation algorithm is trained using the sub-segments from the normal traffic training set, ultimately obtaining a trained RNN model. Next, the sub-segments from the normal traffic training set are fed into the trained RNN model, which outputs the word segmentation results, i.e., the segmented strings. For example, if the original string is (speedofrotoris3000), the segmented string will be (speed of rotor is 3000).

[0099] The segmentation results of the normal traffic training set, which has been segmented, are fed into an unsupervised classification algorithm based on an autoencoder. First, word embedding is performed to convert the word embedding into a vector. The vector is then encoded using a BiLSTM-based encoder and decoded using a BiLSTM-based decoder. The encoder input is subtracted from the decoder output, and this is used as the loss function to train the autoencoder. The number of epochs is set to 200, and after training, the trained autoencoder is obtained. Furthermore, the loss function value of the last epoch is used as a benchmark, with the maximum value as the upper bound and the minimum value as the lower bound to construct the threshold range for the normal traffic loss function. In this experiment, the threshold ranges obtained are shown in Table 2.

[0100] Table 2 Results of different threshold intervals

[0101] Maximum word segmentation length SM Threshold interval 4 (0.50%,0.61%) 6 (0.45%,0.62%) 8 (0.57%,0.66%) 10 (0.49%,0,70%)

[0102] Table 2 shows that when the maximum segmentation length SM is different, the final threshold interval results are also different, which will have a great impact on the subsequent classification results.

[0103] The unsupervised word segmentation algorithm uses the malicious and normal traffic in the test set as input and outputs the segmentation results. The segmentation results are then input into the trained autoencoder to obtain the loss function value, and the threshold interval obtained above is used to determine whether the sample is normal.

[0104] The evaluation indicators used here are as follows:

[0105]

[0106]

[0107]

[0108] Where TP is the number of abnormal samples correctly classified as abnormal samples; FP is the number of normal samples incorrectly classified as abnormal samples; FN is the number of abnormal samples incorrectly classified as normal samples; F1-score is the harmonic mean of Precision and Recall, and is also an indicator that balances the two. After the classifier meets the evaluation index requirements, it is used in the real-time abnormal traffic detection application in step (7).

[0109] Under different SM values, there will be different classification results, such as Figure 2 When SM is small, the word segmentation algorithm cannot extract sufficient semantic information. For example, if a word is 7 characters long, but the maximum word segmentation length is set to 4, only 4 of the 7 characters can be extracted, which destroys the semantic meaning. When SM is large, semantic information can be incorrectly extracted, such as mixing two words together as one. Experiments show that when SM = 8, the F1-score is the highest, reaching 94%.

[0110] At the same time, the N-Gram word segmentation algorithm, LSTM-based classification algorithm, and CNN-based classification algorithm are also compared. Figure 3 As shown in the figure, the F1-scores of the anomaly detection models based on N-Gram+LSTM and N-Gram+CNN are 0.75 and 0.71, respectively, both lower than the 0.94 obtained by the present invention. This also proves that the industrial network traffic anomaly detection method based on unsupervised learning proposed in this invention can fully extract the semantic relationship of traffic load and improve detection accuracy.

[0111] The above examples are merely specific embodiments of the present invention. Obviously, the present invention is not limited to the above examples, and many variations are possible. All variations that can be directly derived or imagined by a person skilled in the art from the disclosure of the present invention should be considered to be within the scope of protection of the present invention.

Claims

1. A method for deep detection of abnormal traffic in industrial networks based on unsupervised learning, characterized in that: The following steps are involved: (1) Obtain the data packet of industrial network traffic, extract the hexadecimal-encoded payload in the data packet, and convert the payload into corresponding characters according to the ASCII code to obtain a string; (2) Setting the maximum segmentation length, dividing the string into segments of different lengths, and then dividing the segments into sub-segments of different lengths, where the maximum length of the segments and sub-segments does not exceed the set value; (3) Input these sub-segments of different lengths into the recurrent neural network to obtain the corresponding occurrence probabilities; (4) Based on the occurrence probabilities of sub-segments of different lengths, the occurrence probabilities of all word combinations of the entire string are calculated, and the sum is taken as the occurrence probability of the string; the recurrent neural network is trained using this occurrence probability as the loss function, and a trained recurrent neural network model is obtained; (5) The sub-segments of different lengths obtained in step (2) are input into the trained recurrent neural network model again to obtain the corresponding occurrence probabilities; based on the occurrence probabilities of the sub-segments of different lengths, the word segmentation method is calculated to maximize the occurrence probability of the entire string, that is, the maximum probability path, and the string with the segmented words is obtained; (6) Embed the word string of the segmented words, input it into the autoencoder, encode it through the encoder, and reconstruct it through the decoder. The difference between the encoder input and the decoder output is calculated as the loss function to train the autoencoder, obtain the loss function set of normal samples, and establish the loss function threshold range; (7) Input the traffic data packet to be detected into the trained recurrent neural network model to obtain the word segmentation result; then input the word segmentation result into the trained autoencoder to obtain the loss function value; if the loss function value of the current traffic data packet is within the loss function threshold range, then the traffic data packet is determined to be normal; Otherwise, it is determined to be abnormal.

2. The method for in-depth detection of abnormal traffic in industrial networks based on unsupervised learning according to claim 1 is characterized in that: The step (1) is specifically as follows: (1.1) Obtain data packets of industrial network traffic; (1.2) Extract the load and calculate it as follows: X'=extract(.pcap) Where pcap is the collected industrial network traffic data packet, the extract() function can extract the payload in the data packet; X' is the extracted payload; (1.3) According to the ASCII code, the extracted payload is converted into the corresponding string. The calculation formula is as follows: X=ASCII(X') Where, the ASCII() function converts the hexadecimal payload to the corresponding character; X is the converted string.

3. The method for in-depth detection of abnormal traffic in industrial networks based on unsupervised learning according to claim 1 is characterized in that: The step (2) is specifically as follows: (2.1) Set the maximum segmentation length SM and divide the string into segments of different lengths. The calculation formula is as follows: Y={X 1~SM ;X 2~SM+1 ;…;X i~i+SM-1 ;…;X (len(X)-SM+1)~len(X) } AND i =X i~i+SM-1 Where Y is the set of string segments divided from X, X i~i+SM-1 Represents a string segment consisting of characters from the i-th character to the i+SM-1-th character in the string X; len(X) represents the length of the string; (2.2) Divide each segment into sub-segments, and the calculation formula is as follows Z i ={Y i 1~SM ;Y i 2~SM ;…;Y i j~SM ;…;Y i SM~SM }(i=1,2,…,len(X)-SM+1) Where Z i Represents X i~i+SM-1 The sub-fragment set of Z i j Indicates that in string X i~i+SM-1 Take the sub-segment consisting of the j-th to SM-th characters.

4. The method for in-depth detection of abnormal traffic in industrial networks based on unsupervised learning according to claim 1 is characterized in that: The step (3) is specifically as follows: (3.1) Sub-fragment set Z i For word embedding, the calculation formula is as follows E_Z i =Word_Embedding(Z i ) Where, the Word_Embedding() function can convert a string sub-segment into a vector; E_Z i It's Z i The corresponding vector; (3.2) Input the vector into the recurrent neural network model to obtain the corresponding probability of occurrence. The calculation formula is as follows: P_Z i =RNN(E_Z i ) In the formula, the RNN() function can calculate the probability of occurrence of each string sub-segment; P_Z i It's Z i The corresponding probability of occurrence.

5. The method for in-depth detection of abnormal traffic in industrial networks based on unsupervised learning according to claim 1 is characterized in that: The step (4) is specifically as follows: (4.1) Based on the occurrence probabilities of different string sub-fragments, calculate the occurrence probabilities of all word combinations of the entire string and sum them up. The formula is as follows: S_P k =Combine{(P_Z1;P_Z2;…;P_Z I ;…;P_Z len(X)-SM+1 ) k } Sum_P=Sum(S_P k )(k=1,2,…,M) Where S_P k is the probability of occurrence of the string after using the k-th word combination; (P_Z1; P_Z2; ...; P_Z I ;…;P_Z len(X)-SM+1 ) k is the kth word combination of the entire sentence; the Combine{} function can calculate the probability of occurrence of the entire string of the kth word combination; Sum_P is the sum of the probability of occurrence of all word combinations in the entire string; Sum() can sum the probabilities, and M means there are M word combinations in total; (4.2) The sum of the probabilities of all word combinations in the entire string is used as the loss function to optimize the recurrent neural network model. The formula is as follows: Loss = -log(Sum_P) In the formula, the loss function takes the opposite of the logarithm of Sum_P.

6. According to the method for in-depth detection of abnormal traffic in industrial networks based on unsupervised learning according to claim 1, the step (5) is specifically: (5.1) Input the sub-segments of different lengths obtained in step (2) into the trained recurrent neural network model again to obtain the corresponding occurrence probability P_Z i ; (5.2) Based on the sub-fragments of different strings, calculate how to segment the words so that the probability of the entire string is maximized, that is, the maximum probability path, and output the segmentation results. The formula is as follows: S_P k =Combine{(P_Z1;P_Z2;…;P_Z I ;…;P_Z len(X)-SM+1 ) k } Max_P=Max(S_P k )(k=1,2,…,M) out_X=(S1,S2,…,S N ) Where Max_P is the maximum probability of occurrence among all word segmentation combinations; out_X is the output word segmentation result; the first sub-segment is S1, and the last sub-segment is S N .

7. According to the method for in-depth detection of abnormal traffic in industrial networks based on unsupervised learning according to claim 1, the step (6) is specifically: (6.1) Perform word embedding on the word-partitioned string, obtain the word-embedded vector and input it into the autoencoder. Use the encoder to encode it, and then use the decoder to decode it. The calculation formula is as follows: X_V=Embedding(out_X) E_V=E_BiLSTM(X_V) D_V=D_BiLSTM(E_V) Where, Embedding() is a word embedding function that converts text into a vector; X_V is the vector after word embedding; E_BiLSTM() is an encoder based on BiLSTM; E_V is the output of the encoder; D_BiLSTM() is a decoder based on BiLSTM; D_V is the output of the decoder; (6.2) Subtract the encoder input from the decoder output and use this as the loss function value to optimize the autoencoder. The loss function threshold range for normal traffic data packets is established. The calculation formula is as follows: A_loss=D_V-X_V T max =max(A_loss') T min =min(A_loss') T=(T min ,T max ) Where A_loss is the difference between the encoder input and the decoder output, and A_loss' is the loss function value obtained by inputting the training set of normal samples into the trained autoencoder; T max is the maximum value of A_loss', T max is its minimum value; T is the threshold interval of the loss function.

8. According to the method for in-depth detection of abnormal traffic in industrial networks based on unsupervised learning according to claim 1, the step (7) is specifically as follows: Extract the payload from the traffic data packet to be tested and convert it into corresponding characters to obtain a string; divide the string into sub-segments of different lengths; input these sub-segments of different lengths into a trained recurrent neural network model to obtain a string with word segmentation; then input the string with word segmentation into a trained autoencoder to obtain its loss function value and determine whether it is within the loss function threshold range; If yes, it is judged as normal traffic; Otherwise it is considered abnormal.

Citation Information

Patent Citations

  • An industrial control system malicious sample generation method based on adversarial learning

    CN109902709A

  • Unsupervised adversarial learning electromagnetic spectrum abnormal signal detection method

    CN112924749A