Dynamic balance loss intrusion detection method for industrial control long-tail data

By dynamically adjusting the category weight and suppressing gradients, the problem of insufficient detection ability of tail categories in industrial control long-tail data is solved, and efficient detection effect is achieved under the distribution of long-tail data.

CN120579061APending Publication Date: 2025-09-02NORTHEASTERN UNIV CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510746437.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

When facing industrial control long-tail data, existing intrusion detection systems are difficult to effectively deal with data imbalance, resulting in insufficient detection capabilities of the model for tail-class attacks, affecting its detection performance and generalization capabilities in actual applications.

Method used

The dynamic focus loss algorithm (DFL) and dynamic balance loss algorithm (DBL) are used to adjust the category weight and suppress the gradient. By dynamically adjusting the category weight, the training attention is shifted to the indistinguishable tail categories, reducing the impact of the head category on the tail category, and improving the detection accuracy and accuracy of the model under the long-tail data distribution.

Benefits of technology

Under the distribution of industrial control long-tail data, the dynamic focus loss algorithm and the dynamic balance loss algorithm significantly improve the detection accuracy and accuracy of the model, especially the recall rate of tail categories, and improve the classification performance of the model in complex long-tail distribution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120579061A_ABST
    Figure CN120579061A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of network security, and discloses a dynamic balance loss intrusion detection method for industrial control long-tail data. Aiming at the problem of long-tail distribution existing in industrial control network real-time flow and intrusion detection data sets, different weights are set for different types of sample data from the aspect of improving the weight of a loss function, so that the accuracy of model detection is improved. The dynamic focus loss algorithm provided by the invention can effectively improve the dominant position of the head category, the training attention is transferred to the tail category which is difficult to distinguish by dynamically adjusting the category weight, and the detection accuracy and precision of the model under long-tail data distribution are improved. According to the dynamic balance loss algorithm provided by the invention, the gradient influence of the head category on the tail category is reduced by suppressing the negative sample gradient, and the detection capability of the model on the tail category is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technology, and in particular to a dynamic balance loss intrusion detection method for industrial control long-tail data. Background Art

[0002] With the development of information technology, the Industrial Internet of Things is gradually moving towards intelligence, digitization and networking. This development has brought unprecedented convenience and efficiency improvements to industrial automation, but it has also brought new challenges and risks in the field of network security.

[0003] Intrusion detection systems have become a key technology in the current field of network security. As a device or software program specially designed to monitor and analyze user and data activities, its main purpose is to detect and report any illegal intrusion behavior that may endanger system security. However, in the context of the widespread application of deep learning technology in intrusion detection systems, the long-tail distribution problem has also become a key obstacle to improving model performance. The amount of normal traffic data far exceeds the number of samples of various attack types, resulting in serious imbalance in the data set. This imbalance causes the deep learning model to tend to overfit the normal traffic data that appears more frequently during training, while ignoring less common attack types, making it impossible to effectively identify and respond to these threats in practical applications. It may also lead to insufficient generalization ability of the model when facing new or unknown attacks. In response to these challenges, the present invention proposes a dynamic balanced loss intrusion detection method for industrial control long-tail data.

[0004] The paper "Tan J, Lu X, Zhang G, et al. Equalization loss v2: A new gradient balance approach for long-tailed object detection[C] / / Proceedings ofthe IEEE / CVF conference on computer vision and pattern recognition. 2021:1685-1694." proposes a new gradient balancing method, Equalization Loss v2, to solve the problem of long-tail object detection. This method introduces a dynamic gradient balancing mechanism to adjust the gradients between different categories, so that the model pays more attention to samples in the tail categories, thereby reducing the gradient influence of the head categories on the tail categories and improving the detection performance of the model under long-tail distribution data. However, the effectiveness of this method is highly dependent on the characteristics and category distribution of the dataset. When the category distribution of the dataset is extremely unbalanced, or the number of samples in the tail categories is too small, this method may not be able to effectively improve the detection performance. In addition, this method requires dynamic adjustment of the gradient during implementation, which may increase the complexity and computational cost of the training process.

[0005] The paper "Wang J, Zhang W, Zang Y, et al. Seesaw loss for long-tailedinstance segmentation[C] / / Proceedings of the IEEE / CVF conference on computervision and pattern recognition. 2021: 9695-9704." introduces a Seesaw loss function for long-tailed instance segmentation. This method introduces two reweighting factors into the loss function to rebalance the positive and negative gradients of each category, respectively, so that the model can better handle data with long-tail distributions. During the training process, the Seesaw loss function dynamically assigns different weights to samples of different categories, increasing attention to tail categories, thereby improving the model's recognition ability for long-tail data. However, this method has high requirements for data preprocessing and feature extraction. If the data is not preprocessed properly or the feature extraction is insufficient, it may affect the model's learning effect on the tail categories. In addition, the reweighting strategy may cause the model to overfit to the tail categories in some cases, affecting its generalization ability.

[0006] In summary, existing methods are unable to effectively cope with data heterogeneity and generalization challenges when processing industrial control long-tail data, resulting in limited detection performance of the model in practical applications. The uneven distribution of data leads to insufficient detection capabilities of the model for tail category attacks. In large-scale industrial Internet of Things environments, the scalability and computational efficiency of existing methods are difficult to guarantee, affecting their application effect and reference value in practical scenarios. Summary of the Invention

[0007] The purpose of the present invention is to propose a dynamic balanced loss intrusion detection method for industrial control long-tail data. (1) In view of the long-tail distribution problem existing in the real-time traffic and intrusion detection data sets of industrial control networks, the present invention starts from improving the weight of the loss function and sets different weights for sample data of different categories, thereby improving the accuracy of model detection. (2) The dynamic focus loss algorithm proposed in the present invention can effectively improve the dominance of the head category. By dynamically adjusting the category weights, the training attention is shifted to the tail category that is difficult to distinguish, thereby improving the detection accuracy and precision of the model under the long-tail data distribution. (3) The dynamic balanced loss algorithm proposed in the present invention reduces the gradient influence of the head category on the tail category by suppressing the negative sample gradient, further improving the model's detection ability for the tail category.

[0008] The technical solution of the present invention is as follows: a dynamic balance loss intrusion detection method for industrial control long-tail data, which selects the dynamic focus loss algorithm DFL or the dynamic balance loss algorithm DBL to train the intrusion detection model according to the characteristics of the industrial control long-tail data; and cyclically trains until the intrusion detection model converges to obtain the final intrusion detection model.

[0009] When the data distribution of industrial control long-tail data is known, the dynamic focus loss algorithm DFL is used to train the intrusion detection model; when the distribution of industrial control long-tail data is unpredictable or it is real-time industrial control traffic data, the dynamic balance loss algorithm DBL is used to train the intrusion detection model.

[0010] The dynamic focus loss algorithm DFL is based on dynamic sample sampling and shifts training attention to the tail category by dynamically adjusting the category weights.

[0011] The dynamic focus loss algorithm DFL adopts a dynamic focus loss function To train:

[0012]

[0013]

[0014]

[0015]

[0016] in, It is One-hot encoding value of a category, if the sample belongs to the category ,but ,otherwise ;0 is the category number of normal network traffic; The intrusion detection model predicts that the sample belongs to the category probability; logits predicted by the intrusion detection model; is the weight of the i-th category; is a focusing parameter used to reduce the attention of easy-to-classify samples; is the dynamic weight; C is the number of categories of industrial control network traffic.

[0017] The dynamic weight The assignment is as follows:

[0018]

[0019] β is the proportion of long-tail data, e is the number of current training rounds, and E is the total number of training rounds.

[0020] At the beginning of each round of training, the dynamic balance loss algorithm DBL calculates the cumulative number of samples in each category, the balance factor of each category, and the probability distribution value after the softmax activation function. By suppressing the negative sample gradient, the balance factor is dynamically adjusted to reduce the gradient influence of the head category on the tail category.

[0021] The dynamic balance loss function of the dynamic balance loss algorithm DBL as follows:

[0022]

[0023] in, It is One-hot encoding value of each category;

[0024]

[0025] is the logits predicted by the intrusion detection model; C is the number of categories of industrial control network traffic;

[0026] The loss propagation gradient is:

[0027]

[0028] It is a balance factor of different categories, mainly composed of gradient suppression factor and compensation factor, and is calculated as follows:

[0029]

[0030] in, To control hyperparameters;

[0031]

[0032]

[0033] in, For the The number of class samples; No. The number of class samples;

[0034]

[0035]

[0036] is the inhibitory factor, according to the tail category and the head class The ratio of the number of samples between and is used to reduce the penalty for the tail category;

[0037]

[0038] The parameters Is a scaling factor used to control the degree of suppression. is a custom threshold;

[0039] is the compensation factor used to compensate for the gradient of misclassified samples:

[0040]

[0041] Among them, the parameters It is also the scaling factor, which is used to control the degree of compensation.

[0042] The intrusion detection model includes three parallel CNN layers, a bidirectional LSTM layer and an FC layer;

[0043] The input data is passed through three parallel CNN layers to extract hidden vectors. The hidden vectors of the relevant CNN layers are combined to generate u, which is input to the bidirectional LSTM layer and then passed through the FC layer to obtain the prediction result.

[0044]

[0045] is the first CNN layer, is the second CNN layer, This is the third CNN layer, and the convolution kernel sizes of the three are different; Indicates a combination operation;

[0046]

[0047] is a bidirectional LSTM layer, for layer.

[0048] Beneficial effects of the present invention: This invention addresses the long-tail data distribution problem in the industrial control field and proposes a dynamic balance loss intrusion detection method for industrial control long-tail data. The key points include the following three points:

[0049] (1) Aiming at the long-tail distribution problem of industrial control network real-time traffic and intrusion detection data sets, the present invention starts from improving the loss function weight and sets different weights for different categories of sample data, thereby improving the accuracy of model detection.

[0050] (2) A dynamic focal loss algorithm is proposed, which can improve the dominance of the head category by dynamically adjusting the category weights, shift the training attention to the tail category that is difficult to distinguish, and improve the detection accuracy and precision of the model under long-tail data distribution.

[0051] (3) A dynamic balance loss algorithm is proposed to reduce the gradient influence of the head category on the tail category by suppressing the gradient of negative samples, thereby further improving the model's detection ability for the tail category. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 This is a flowchart of the dynamic balance loss intrusion detection method for industrial control long-tail data;

[0053] Figure 2 This is the architecture diagram of the intrusion detection model; DETAILED DESCRIPTION

[0054] The specific implementation of the present invention is further described in detail below with reference to the accompanying drawings and examples.

[0055] This paper proposes a dynamic balanced loss intrusion detection method for industrial control long-tail data. This paper proposes two loss algorithms, namely the dynamic focus loss algorithm (DFL) and the dynamic balanced loss algorithm (DBL), to improve the dominance of the normal traffic header category and shift the training attention to the difficult-to-distinguish abnormal traffic category. Figure 1 As shown, the processing flow of the algorithm of the present invention includes the following steps:

[0056] Step 1: Preprocess the original industrial control traffic data to obtain data in a format that meets the input requirements of the intrusion detection model. The specific steps are as follows:

[0057] The character features in the original industrial control flow data are converted into numerical features using one-hot encoding.

[0058] All numerical features are standardized. The standardization formula is as follows:

[0059]

[0060] in is the original eigenvalue, is the average value of the feature, is the standard deviation of the feature, is the standardized eigenvalue.

[0061] Use the minimum-maximum normalization method to map the standardized values ​​to the [0,1] interval. The normalization formula is as follows:

[0062]

[0063] in is the minimum value of this feature, is the maximum value of this feature, is the normalized feature value, which is the input data of the intrusion detection model.

[0064] Step 2: Model initialization, the intrusion detection model used in this invention is as follows: Figure 2 As shown, it includes three parallel CNNs, a bidirectional LSTM and two FC layers.

[0065] CNN parameter initialization:

[0066]

[0067] For any input vector x, the convolutional neural network can slide and extract spatial local features at various scales through convolution kernels of different sizes.

[0068] LSTM parameter initialization:

[0069]

[0070] The hidden vectors of the relevant layers are combined to generate u, which is then fed into the LSTM layer. The LSTM layer is added primarily to process real-time traffic data with time series. After passing through the LSTM, it is fed into the FC layer, which ultimately produces a prediction result.

[0071] Step 3: Use the dynamic focus loss algorithm (DFL) or dynamic balance loss algorithm (DBL) for training;

[0072] In this paper, it is assumed that there are C categories of network traffic. To facilitate the correspondence with the dataset labels, 0 is defined as the category number of normal network traffic. The classification loss formula is as follows:

[0073]

[0074]

[0075] in, It is One-hot encoding value of a category, if the sample belongs to the category ,but ,otherwise . The model predicts that the sample belongs to the category probability. The logits predicted by the model.

[0076] Point loss is a loss function proposed to solve the long-tail distribution problem. It reduces the focus on easy-to-classify samples and increases the focus on difficult-to-classify samples, thereby encouraging the model to focus on improving the recognition ability of difficult-to-classify samples:

[0077]

[0078] in, is the weight of the i-th category, which can be adjusted according to the frequency or other importance of the category. is a focusing parameter used to reduce the attention of easy-to-classify samples.

[0079] Based on the focus loss, this paper proposes a dynamic focus loss algorithm. Using a progressive sampling method, as learning progresses, the sample balance probability will gradually transform into class balance. The specific implementation of the probability calculation is as follows:

[0080]

[0081] in,

[0082]

[0083]

[0084] For the The number of samples.

[0085] The dynamic focal loss algorithm uses dynamic focal loss in the training process, which is implemented as follows:

[0086]

[0087] in,

[0088]

[0089] e is the current training epoch, E is the total number of training epochs, and β is the proportion of long-tail data. λ decreases exponentially in the second half of training, thereby shifting the loss to the focal loss.

[0090] In some industrial control scenarios, data is real-time and its distribution is unpredictable, making it impossible to preprocess the training dataset. The dynamic focus loss algorithm cannot resample in such scenarios, resulting in suboptimal results. To address this issue, the present invention proposes another dynamic balance loss algorithm.

[0091] In this paper, it is assumed that there are C categories of network traffic. To facilitate the correspondence with the dataset labels, 0 is defined as the category number of normal network traffic. The classification loss formula is as follows:

[0092]

[0093]

[0094] in, It is One-hot encoding value of a category, if the sample belongs to the category ,but ,otherwise . The model predicts that the sample belongs to the category probability. The logits predicted by the model.

[0095] Each category The gradient passed by the loss is:

[0096]

[0097] Based on the cross entropy loss function based on softmax, in order to reduce the gradient suppression effect of other samples on difficult-to-classify samples, the present invention proposes the following loss function:

[0098]

[0099] in,

[0100]

[0101] It can be further deduced that the loss transfer gradient at this time is:

[0102]

[0103] There are different types of balance factors, which are mainly composed of inhibition factors and compensation factors. The calculation method is as follows:

[0104]

[0105] in, To control hyperparameters, the design is to not suppress the gradient when the prediction is wrong, and to calculate based on the number of samples when the prediction is correct.

[0106]

[0107] is the inhibitory factor, according to the tail category and the head class The ratio of the number of samples between and is used to reduce the penalty for the tail categories.

[0108]

[0109] When the head category The number of training samples is much higher than that of the tail category When the head category The samples will be of the tail category To improve the prediction probability of the tail category, it is necessary to reduce the The parameter Is a scaling factor used to control the degree of suppression. is a custom threshold.

[0110] A compensation factor is added on the basis of the suppression factor to compensate for the gradient of misclassified samples:

[0111]

[0112] The parameters It is also the scaling factor, which is used to control the degree of compensation.

[0113] The dynamic balance loss algorithm is applied to the intrusion detection model. At the beginning of each round of training, it is necessary to calculate the cumulative number of samples in each category, the balance factor of each category, and the probability distribution value after the softmax activation function.

[0114] This example evaluates the proposed algorithm on the Natural Gas Pipeline (NGP) (Table 1) and UNSW-NB15 (Table 2) intrusion detection datasets. The NGP dataset contains seven types of attack and benign samples, each consisting of 26 features and one label. The UNSW-NB15 dataset contains nine different types of attack and benign samples, each consisting of 42 features and one label. In both datasets, each attack type employs a unique strategy targeting a specific target. The distribution of samples across the two datasets is shown in Tables 1 and 2, respectively.

[0115] Table 1. Sample distribution of the natural gas pipeline dataset (NGP)

[0116]

[0117] Table 2 shows the sample distribution of the UNSW-NB15 dataset.

[0118]

[0119] To validate the effectiveness of our proposed method, we conducted a series of experiments comparing it with methods such as cross-entropy loss and focal loss on the natural gas pipeline dataset and the UNSW-NB15 dataset. The experiments employed various data distributions, including balanced, 100% imbalanced, 200% imbalanced, 500% imbalanced, and extremely imbalanced. In each distribution, 80% of the data was used for training and 20% for testing.

[0120] Experimental results, shown in Tables 3, 4, 5, 6, 7, 8, 9, 10, and 11, show that the dynamic focal loss and dynamic balanced loss methods outperform cross-entropy loss and focal loss in terms of accuracy, precision, recall, F1 score, and false alarm rate under different data distributions. On the natural gas pipeline dataset, DFL achieved an accuracy of 96.61%, a 2.36% improvement over CE, and a 16.2% increase in recall for tail categories, while only increasing training time by 4.5%. On the UNSW-NB15 data, DFL achieved an F1 score of 89.7% across sample categories (compared to 85.6% for CE). DFL's advantage lies in its segmented weight design: early in training, it prioritizes the head categories for rapid convergence, while later using focal loss to enforce tail boundaries.

[0121] Experiments demonstrate that the DBL algorithm, through its gradient suppression and compensation mechanisms, performs exceptionally well on extremely imbalanced industrial control data. In a natural gas pipeline dataset, DBL achieved an accuracy of 97.77% (compared to 94.25% for CE and 96.61% for DFL), improved the recall of tail classes to 85.3% (compared to 58.7% for CE), and reduced the false alarm rate to 0.81% (compared to 3.48% for CE). For the UNSW-NB15 data, DBL achieved a recall of 41.2% (compared to 22.8% for CE), demonstrating its strong adaptability to sparse data. DBL's core advantage lies in dynamically adjusting gradient weights to reduce the suppression of tail classes by head classes while compensating for misclassified examples, thereby achieving more balanced classification performance in complex, long-tail distributions.

[0122] Table 3 shows the performance of cross entropy loss (CE) in different distributions of natural gas pipeline dataset

[0123]

[0124] Table 4 shows the performance of focal loss (FL) on different distributions of natural gas pipeline datasets

[0125]

[0126] Table 5 shows the performance of dynamic focus loss (DFL) on different distributions of natural gas pipeline datasets

[0127]

[0128] Table 6 shows the performance of dynamic balance loss (DBL) in different distributions of natural gas pipeline datasets

[0129]

[0130] Table 7 shows the performance of cross entropy loss (CE) in different distributions of the UNSW-NB15 dataset

[0131]

[0132] Table 8 shows the performance of focal loss (FL) on different distributions of the UNSW-NB15 dataset

[0133]

[0134] Table 9 shows the performance of dynamic focus loss (DFL) on different distributions of the UNSW-NB15 dataset

[0135]

[0136] Table 10 shows the performance of dynamic balance loss (DBL) in different distributions of the UNSW-NB15 dataset

[0137]

[0138] Table 11 shows the time performance of different losses on each data set

[0139]

[0140] In summary, the present invention performs well under different data distribution conditions, effectively improving the accuracy and precision of the intrusion detection model, and verifying its superiority in processing long-tail distribution data.

Claims

1. A dynamic balance loss intrusion detection method for industrial control long-tail data, characterized by: According to the characteristics of industrial control long-tail data, the dynamic focus loss algorithm DFL or the dynamic balance loss algorithm DBL is selected to train the intrusion detection model; the training is repeated until the intrusion detection model converges to obtain the final intrusion detection model.

2. The dynamic balance loss intrusion detection method for industrial control long-tail data according to claim 1 is characterized in that: When the data distribution of industrial control long-tail data is known, the dynamic focus loss algorithm DFL is used to train the intrusion detection model; when the distribution of industrial control long-tail data is unpredictable or it is real-time industrial control traffic data, the dynamic balance loss algorithm DBL is used to train the intrusion detection model.

3. The dynamic balance loss intrusion detection method for industrial control long-tail data according to claim 1 is characterized in that: The dynamic focus loss algorithm DFL is based on dynamic sample sampling and shifts training attention to the tail category by dynamically adjusting the category weights.

4. The dynamic balance loss intrusion detection method for industrial control long-tail data according to claim 3 is characterized in that: The dynamic focus loss algorithm DFL adopts a dynamic focus loss function To train: , in, It is One-hot encoding value of a category, if the sample belongs to the category ,but ,otherwise ;0 is the category number of normal network traffic; Is the intrusion detection model predicting that the sample belongs to the category probability; logits predicted by the intrusion detection model; is the weight of the i-th category; is a focusing parameter used to reduce the attention of easy-to-classify samples; is the dynamic weight; C is the number of categories of industrial control network traffic.

5. The dynamic balance loss intrusion detection method for industrial control long-tail data according to claim 4 is characterized in that: The dynamic weight The assignment is as follows: , β is the proportion of long-tail data, e is the number of current training rounds, and E is the total number of training rounds.

6. The dynamic balance loss intrusion detection method for industrial control long-tail data according to claim 1 is characterized in that: At the beginning of each round of training, the dynamic balance loss algorithm DBL calculates the cumulative number of samples in each category, the balance factor of each category, and the probability distribution value after the softmax activation function. By suppressing the negative sample gradient, the balance factor is dynamically adjusted to reduce the gradient influence of the head category on the tail category.

7. The dynamic balance loss intrusion detection method for industrial control long-tail data according to claim 6 is characterized in that: The dynamic balance loss function of the dynamic balance loss algorithm DBL as follows: ,in, It is One-hot encoding value of each category; , is the logits predicted by the intrusion detection model; C is the number of categories of industrial control network traffic; The loss propagation gradient is: It is a balance factor of different categories, mainly composed of gradient suppression factor and compensation factor, and is calculated as follows: ,in, To control hyperparameters; , ,in, For the The number of class samples; No. The number of class samples; , , is the inhibitory factor, according to the tail category and the head class The ratio of the number of samples between and is used to reduce the penalty for the tail category; , where the parameters Is a scaling factor used to control the degree of suppression. is a custom threshold; is the compensation factor used to compensate for the gradient of misclassified samples: , where the parameter It is also the scaling factor, which is used to control the degree of compensation.

8. The dynamic balance loss intrusion detection method for industrial control long-tail data according to claim 1 is characterized in that: The intrusion detection model includes three parallel CNN layers, a bidirectional LSTM layer and an FC layer; The input data is passed through three parallel CNN layers to extract hidden vectors. The hidden vectors of the relevant CNN layers are combined to generate u, which is input to the bidirectional LSTM layer and then passed through the FC layer to obtain the prediction result. is the first CNN layer, is the second CNN layer, This is the third CNN layer, and the convolution kernel sizes of the three are different; Indicates a combination operation; is a bidirectional LSTM layer, for layer.