Industrial control system intrusion detection method and system based on Mamba-ReLU

The virtual attack samples are generated through the Mamba-ReLU method and combined with dynamic weight adjustment, the problem of insufficient detection of traditional industrial control system in complex attack mode is solved, and efficient and accurate intrusion detection of industrial control system is achieved.

CN120378182APending Publication Date: 2025-07-25张镇雄
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510589756.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

When facing advanced persistent threats and attack variants, the traditional industrial control system intrusion detection methods have insufficient detection capabilities, and there are problems with high false alarm rate, sample imbalance and feature selection limitations.

Method used

The intrusion detection method based on Mamba-ReLU is adopted to generate high-quality virtual attack samples by generating adversarial networks, and combined with the enhanced classification model and dynamic weight adjustment mechanism of stack integration, optimize computing efficiency and improve detection capabilities of complex attack modes.

Benefits of technology

It significantly improves the detection capabilities of zero-day attacks and variant attacks, reduces the risks of missed and false alarms, meets the real-time detection requirements of industrial scenarios, and improves detection coverage and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378182A_ABST
    Figure CN120378182A_ABST
Patent Text Reader

Abstract

The invention discloses an industrial control system intrusion detection method and system based on Mamba-ReLU, and relates to the field of industrial control system intrusion detection. The method comprises the steps of firstly obtaining network flow data of an industrial control system, and extracting a real sample from the network flow data; generating a virtual attack sample, screening the virtual attack sample, and combining the virtual attack sample with a real sample to form an extended sample; constructing an enhanced classification model based on stack integration, obtaining a feature vector X of the extended sample, inputting the feature vector X into the enhanced classification model, and outputting a first classification probability and a second classification probability; inputting the first classification probability and the second classification probability into a secondary classifier, and judging whether the network traffic data is attack traffic or not according to an output result of the secondary classifier; the calculation efficiency is optimized through a ReLU linear attention mechanism, and the output weight of the probability is dynamically adjusted through a Mamba model; according to the method, the risk of missing report and false report caused by attack mode change can be reduced, and the detection sensitivity of persistent attacks is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intrusion detection in industrial control systems, involves data analysis technology, and specifically is an intrusion detection method for industrial control systems based on Mamba-ReLU. Background Art

[0002] An industrial control system (ICS) is a business process control system composed of various automation control components and process control components for collecting and monitoring real-time data, ensuring the automated operation of industrial infrastructure, process control, and monitoring.

[0003] Traditionally, industrial control systems operate in a closed and physically isolated network environment, bringing a certain degree of natural security barrier. However, with the wave of modernization and technological progress, these systems are gradually integrating into the open Internet and thus facing diverse network threats such as advanced persistent threats, malware, and cross-domain network attacks.

[0004] Traditional machine learning methods usually rely on manual feature extraction and have poor performance in dealing with sample imbalance. In the case where attack traffic is scarce, it leads to insufficient recognition ability of the model for attack traffic.

[0005] The main problems of these traditional IDS methods include: poor detection ability for unknown attacks, high false alarm rate, sample imbalance, and limitations in feature selection. Summary of the Invention

[0006] Aiming at the deficiencies of the existing technology, the present invention provides an intrusion detection method and system for industrial control systems based on Mamba-ReLU.

[0007] To achieve the above purpose, the technical solution of the present invention is as follows:

[0008] In the first aspect, the present invention discloses an intrusion detection method for industrial control systems based on Mamba-ReLU, including the following steps:

[0009] Obtain the network traffic data of the industrial control system, and extract traffic data features from the network traffic data to form real samples;

[0010] Generate virtual attack samples Xfake through a GAN generator;

[0011] Screen the virtual attack samples Xfake through a GAN discriminator and merge them with real samples to form extended samples;

[0012] Construct an enhanced classification model based on stacked integration, where the enhanced classification model includes a first classifier based on a tree model, a second classifier based on a GAN discriminator, and a secondary classifier;

[0013] Obtain the feature vector X of the extended sample, input the feature vector X into the first classifier to output the first classification probability, and synchronously input it into the second classifier to output the second classification probability;

[0014] Optimize the calculation efficiency through the ReLU linear attention mechanism, and dynamically adjust the output weights of the first classifier and the second classifier through the Mamba model;

[0015] Input the first classification probability and the second classification probability into the secondary classifier according to the output weights, and determine whether the network traffic data is attack traffic based on the output result of the secondary classifier.

[0016] In a second aspect, the present invention discloses an industrial control system intrusion detection system based on Mamba-ReLU, including:

[0017] Traffic acquisition module: used to acquire the network traffic data of the industrial control system, and extract the traffic data features from the network traffic data to form real samples;

[0018] Sample generation module: used to generate virtual attack samples Xfake through a GAN generator, screen the virtual attack samples Xfake through a GAN discriminator, and merge them with real samples to form extended samples;

[0019] Model construction module: used to construct an enhanced classification model based on stacked integration, where the enhanced classification model includes a first classifier based on a tree model, a second classifier based on a GAN discriminator, and a secondary classifier;

[0020] Basic classification module: used to obtain the feature vector X of the extended sample, input the feature vector X into the first classifier to output the first classification probability, and synchronously input it into the second classifier to output the second classification probability;

[0021] Optimization and adjustment module: used to optimize the calculation efficiency through the ReLU linear attention mechanism, and dynamically adjust the output weights of the first classifier and the second classifier through the Mamba model;

[0022] Stacked integration module: used to input the first classification probability and the second classification probability into the secondary classifier according to the output weights, and determine whether the network traffic data is attack traffic based on the output result of the secondary classifier.

[0023] The present invention has the following beneficial effects:

[0024] 1. By introducing a generative adversarial network (GAN) to generate high-quality virtual attack samples and combining the dynamic screening mechanism of the discriminator, the attack sample library is expanded while preserving the time-series characteristics of industrial control traffic. This method solves the problem of insufficient model training caused by scarce attack samples in traditional methods and significantly improves the model's detection ability for zero-day attacks and variant attacks;

[0025] 2. By adopting a stacked ensemble model to fuse heterogeneous classifiers of a tree model (such as XGBoost) and a GAN discriminator, the parsing ability of the tree model for structured features and the deep feature extraction advantage of the discriminator are combined. The secondary classifier achieves decision-making complementarity through probability fusion, effectively improving the detection coverage rate and accuracy for complex attack patterns such as protocol tampering and time-series stealth attacks;

[0026] 3. Through the ReLU linear attention mechanism, the feature vector is block-processed and the matrix operation order is optimized, reducing the computational complexity from quadratic to linear level while preserving the key feature expression ability. This solution significantly reduces memory occupancy and computational latency under the resource constraints of embedded devices, meeting the real-time detection timeliness requirements of industrial scenarios;

[0027] 4. By introducing a weight dynamic adjustment mechanism based on the Mamba model, by analyzing the time dependence of the historical outputs of the classifier and combining the time span of the attack pattern, the classifier weights are adaptively allocated. This technology solves the problem that traditional static ensemble models are difficult to capture the time-series attack characteristics, reduces the risk of missed reports and false alarms caused by changes in attack patterns, and enhances the detection sensitivity to persistent attacks. Description of the Drawings

[0028] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0029] Figure 1 It is the overall method block diagram of Embodiment 1 of the present invention;

[0030] Figure 2 It is the method flow chart of Embodiment 1 of the present invention;

[0031] Figure 3 It is the flow chart of extracting feature vectors in Embodiment 1 of the present invention;

[0032] Figure 4 It is the overall system block diagram of Embodiment 2 of the present invention. Detailed Embodiments

[0033] It is easy to understand that according to the technical solution of the present invention, without changing the essential spirit of the present invention, those of ordinary skill in the art can propose various interchangeable structural ways and implementation ways. Therefore, the following specific embodiments and drawings are only exemplary descriptions of the technical solution of the present invention, and should not be regarded as all of the present invention or as a limitation or restriction on the technical solution of the present invention.

[0034] Application Overview:

[0035] Traditional intrusion detection systems (IDS) usually rely on two main methods: signature-based detection and anomaly-based detection. Signature-based detection relies on a predefined attack signature library to identify attacks by comparing traffic data with known attack signatures. The main problem with this method is that it can only detect known attacks and cannot identify new or unknown attacks, especially zero-day attacks and attack variants. Anomaly-based detection methods, on the other hand, identify traffic that deviates from normal behavior by building a behavior model of normal traffic. This method can detect unknown attacks, but faces a high false positive rate when network traffic fluctuates greatly or the normal traffic pattern changes. Traditional machine learning methods usually rely on manual feature extraction and perform poorly in dealing with sample imbalance. In the case where attack traffic is scarce, it results in insufficient ability of the model to identify attack traffic.

[0036] To solve the above problems, those skilled in the art first noticed the restriction of sample imbalance on model performance. By studying the application potential of generative adversarial networks in data augmentation, it was found that simply increasing the number of generated samples may introduce low-quality data with a large deviation from the real distribution. For this reason, a screening mechanism for collaborative optimization of the generator and discriminator was proposed, and the effectiveness of samples was controlled by setting a quality threshold. Regarding the characteristic that attack patterns evolve over time in industrial environments, it was found that traditional static ensemble models are difficult to capture the temporal correlation of attack behaviors. Thus, the feasibility of introducing state space models to analyze the time dependence of historical classification results was explored. At the feature calculation level, it was observed that the computational complexity of traditional attention mechanisms restricts the real-time detection efficiency, and then the feasibility of linear transformation and sparsification operations was studied.

[0037] Therefore, the present application proposes the following technical solutions.

[0038] Example 1:

[0039] As Figures 1-3Shown: An industrial control system intrusion detection method based on Mamba-ReLU, comprising the following steps: obtaining network traffic data of the industrial control system, and extracting traffic data features from the network traffic data to form real samples; generating virtual attack samples Xfake through a GAN generator; screening the virtual attack samples Xfake through a GAN discriminator, and merging them with the real samples to form extended samples; constructing an enhanced classification model based on stacked integration, the enhanced classification model including a first classifier based on a tree model, a second classifier based on a GAN discriminator, and a secondary classifier; obtaining a feature vector X of the extended samples, inputting the feature vector X into the first classifier, outputting a first classification probability, and synchronously inputting it into the second classifier, outputting a second classification probability; optimizing the calculation efficiency through a ReLU linear attention mechanism, and dynamically adjusting the output weights of the first classifier and the second classifier through a Mamba model; inputting the first classification probability and the second classification probability into the secondary classifier according to the output weights, and determining whether the network traffic data is attack traffic based on the output result of the secondary classifier.

[0040] Among them, the generative adversarial network generator refers to a deep neural network that generates virtual samples with a distribution similar to that of real attack samples through adversarial training. Specifically, a multi-layer long short-term memory network can be used to receive the joint input vector after splicing the noise vector and the real samples, and capture the time dependence of the traffic data through temporal unfolding processing. The generative adversarial network discriminator screening mechanism refers to retaining virtual samples that conform to the industrial control traffic characteristics by setting a authenticity threshold. Specifically, a method of mixing real samples and generated samples can be used, and the samples with discriminator output values higher than the threshold are calculated as effective expansion data. The stacked integration enhanced classification model refers to a multi-stage classification architecture that uses the output probabilities of heterogeneous classifiers as the input of the secondary classifier. Specifically, a combination method of using a tree model to process structured features and a generative adversarial network discriminator to extract deep features can be adopted. The linear attention mechanism optimization refers to reducing the computational complexity by block processing and matrix operation order adjustment. Specifically, a method of calculating the normalized cosine similarity after multi-head block division of the feature vector can be used. The state space model dynamic adjustment refers to allocating classifier weights based on the time correlation of historical classification results. Specifically, an adaptive allocation method of sliding window weighted analysis of the time span of the attack pattern can be adopted.

[0041] Specifically, the original industrial control network traffic is used to construct a real sample library through feature extraction, solving the limitations of manual feature design in traditional methods. The generator of the generative adversarial network fuses the features of real attack samples in the noise vector, and captures the temporal characteristics of industrial protocol sessions through the long short-term memory network to ensure that the generated samples conform to the industrial control traffic pattern. The discriminator sets a quality threshold to screen and retain high-confidence generated samples, avoiding the contamination of the training set by invalid data. The tree model classifier and the discriminator of the generative adversarial network form heterogeneous feature processing channels. The former captures traffic statistical features, and the latter uses the discrimination ability obtained through adversarial training. The secondary classifier fuses the probability outputs of the two types to achieve feature complementarity at the decision-making level. The linear attention mechanism performs block normalization on the feature vectors and reduces the memory occupancy through matrix operation decomposition. The state space model analyzes the temporal pattern of the historical output of the classifier, dynamically adjusts the integration weights according to the duration of the attack behavior, and enhances the detection sensitivity to persistent attacks.

[0042] Compared with the prior art, traditional data augmentation methods expand samples through simple sampling or noise injection, which easily destroys the temporal correlation of industrial traffic. This solution retains the traffic time-dependent features through the generator of the generative adversarial network, and the discriminator screening mechanism ensures the effectiveness of data augmentation. Conventional ensemble learning uses a fixed weight combination classifier, which is difficult to adapt to the dynamic changes of attack patterns in industrial environments. This solution introduces a state space model to analyze the temporal correlation of the historical performance of the classifier and realizes the dynamic optimization of weight allocation. Traditional attention mechanisms generate excessive computational loads when processing industrial long-sequence data. This solution improves the computational efficiency while retaining the key feature expression ability by linear block and matrix operation order adjustment.

[0043] Through the above technical solutions, this application effectively alleviates the problem of insufficient model training caused by the scarcity of attack samples in industrial control scenarios. The high-quality virtual samples generated by the generative adversarial network improve the class imbalance phenomenon. The stacked integration mechanism of heterogeneous classifiers combines the parsing ability of the tree model for structured features and the deep feature extraction advantages of the discriminator of the generative adversarial network, improving the recognition accuracy of complex attack patterns. The dynamic weight adjustment module enables the detection system to adapt to the temporal evolution characteristics of attack behaviors and reduces the false alarm risk caused by changes in attack patterns. The linear attention optimization significantly reduces the consumption of feature calculation resources and meets the timeliness requirements of real-time detection in industrial fields.

[0044] This application further proposes that the traffic data features include traffic volume, packet frequency, protocol type, moving window mean and variance, protocol field outliers, packet interval time series, and session duration.

[0045] Among them, the traffic volume refers to the amount of data transmitted per unit time, which can be specifically realized by byte count statistics and is used to quantify network load fluctuations; the packet frequency refers to the number of packets per unit time, which can be specifically realized by counter statistics and reflects the communication density; the protocol type refers to the category of communication protocols used in industrial control systems, and proprietary protocols such as Modbus or DNP3 can be specifically identified by deep packet parsing technology; the sliding window mean and variance refer to dynamic statistics based on a time window, which can be specifically realized by moving average calculation within a fixed-length window and are used to capture traffic anomalies in periodic instructions; the protocol field outlier refers to the situation where the key fields in industrial protocol messages deviate from the normal range, which can be specifically realized by comparing with a preset register address whitelist; the packet inter-arrival time series refers to the sequence of intervals between the arrival times of adjacent packets, which can be specifically generated by calculating the difference in microsecond-level timestamps; the session duration refers to the complete session duration from TCP handshake to connection termination, which can be specifically realized by tracking the session state machine.

[0046] Specifically, the protocol type feature identifies industrial control-specific protocols by parsing the header fields of packets and can detect the spoofing attack traffic of non-standard protocols. The sliding window mean and variance construct a dynamic statistical baseline in the time dimension. By calculating the moving average and dispersion degree of traffic metrics within consecutive windows, it can identify abnormal fluctuations in periodic control instructions. The protocol field outlier detection targets key protocol fields such as function codes and register addresses, and discovers parameter tampering behaviors by comparing with a preset compliance value range. The packet inter-arrival time series adopts a time series analysis method to capture abnormal interval patterns that do not conform to the industrial control time sequence logic, such as hidden attacks in high-frequency heartbeat packets. The session duration feature combines the long connection characteristics of industrial control systems and identifies abnormal interruptions or resource exhaustion attacks by statistically analyzing the differences in session lifecycle lengths. The combination of traffic volume and packet frequency constructs multi-dimensional space features, which can distinguish bursty service traffic from malicious flooding attacks. Each feature forms a complementarity from the three dimensions of protocol compliance, time sequence regularity, and spatial distribution, jointly constructing a multi-dimensional detection benchmark for industrial control traffic.

[0047] Compared with the existing technologies, traditional methods usually only extract single-dimensional information such as traffic statistics features or protocol types, lacking the joint analysis of time sequence patterns and protocol semantics. Existing feature engineering solutions do not consider the compliance verification of industrial protocol fields and cannot identify advanced attacks based on protocol field tampering. This solution captures periodic traffic anomalies through sliding window statistics and combines protocol field outlier detection to identify protocol layer attacks, solving the problem of insufficient detection capabilities of traditional methods for time sequence attacks and semantic attacks. The introduction of the packet inter-arrival time series fills the gap in microsecond-level time sequence features and can discover hidden attack patterns ignored by traditional statistical methods.

[0048] Through the above technical solutions, the present application can effectively identify malicious attacks based on protocol field tampering, detect abnormal traffic disguised as legitimate protocols, and reduce the false positive rate caused by a single feature dimension. The sliding window statistic enhances the detection sensitivity to periodic attacks, and the protocol compliance verification blocks the penetration attempts of illegal protocol fields. The timing interval analysis improves the ability to capture hidden timing attack patterns, and the multi-dimensional feature fusion significantly increases the detection coverage in complex attack scenarios.

[0049] The present application further proposes a method for generating virtual attack samples through a generative adversarial network, which specifically includes the following steps: The generator receives a random noise vector obeying a Gaussian distribution and concatenates it with the real attack sample features to form a joint input vector, and uses a multi-layer long short-term memory network to perform temporal unfolding processing on the input, and maps the hidden state to a virtual attack sample with the same dimension as the real sample through a fully connected layer.

[0050] Among them, the random noise vector of the Gaussian distribution refers to a multi-dimensional vector sampled from a normal distribution with a mean of 0 and a standard deviation of 1, which can be specifically implemented by using the torch.randn function in the PyTorch framework, and is used to inject random perturbations during the generation process to expand sample diversity. Among them, the feature dimension concatenation refers to concatenating the noise vector and the numerical features of the real sample in the same dimension direction, which can be specifically implemented by using the concatenate function in the Numpy library, and retains the protocol field outliers and session duration features of the real attack sample. Among them, the temporal unfolding processing of the long short-term memory network refers to processing the input sequence step by step through memory units, which can be specifically implemented by using the LSTM layer in the TensorFlow framework, and captures the dynamic change rules of the packet interval time series in the industrial control network traffic. Among them, the dimension mapping of the fully connected layer refers to converting the output of the hidden layer into a feature vector with the same dimension as the real sample through a linear transformation, which can be specifically implemented by using the Dense layer in the Keras framework, and ensures that the generated samples are aligned with the real data in the feature space distribution.

[0051] Specifically, after receiving the joint input vector containing noise and real samples, the generator performs temporal modeling on the traffic features through a multi-layer long short-term memory network. The memory units in the network layer pass the hidden state layer by layer in the time dimension, and learn the dynamic patterns of traffic size and sliding window mean during the industrial control protocol interaction process. In the hidden layer processing stage, the unit state update process at each time step is realized through the input gate, forget gate and output gate mechanisms, effectively capturing the long-term dependencies of protocol type conversion and packet frequency changes. After the preset number of time steps are unfolded, the final hidden state is passed to the fully connected layer for linear projection, and a virtual attack sample with a complete feature dimension is output, and this sample is statistically consistent with the real attack sample in terms of protocol field outliers and packet interval time series features.

[0052] Compared with the prior art, traditional generative adversarial networks usually directly use fully connected networks to generate samples in industrial control intrusion detection applications, without considering the unique temporal correlation characteristics of network traffic. When existing methods use convolutional neural networks to process traffic data, it is difficult to effectively model the continuous change rules of session duration and sliding window mean in industrial protocols. Through the introduction of long short-term memory networks for temporal unfolding processing, this solution can accurately capture the dynamic evolution process of protocol field outliers in the network traffic of industrial control systems. At the same time, combined with the splicing mechanism of noise vectors and real samples, it expands the diversity of generated samples while retaining the known attack feature distribution.

[0053] Through the above technical solution, this application effectively solves the problem of insufficient model training data caused by the scarcity of attack samples in industrial control systems. The generated virtual attack samples maintain a high degree of similarity with real attack samples in terms of protocol type distribution and packet frequency characteristics. At the same time, by injecting noise to increase sample diversity, it alleviates the classifier bias problem caused by sample class imbalance in traditional methods. The generated samples can accurately reflect the temporal change rules of protocol field outliers in industrial control network traffic, providing representative training data for subsequent classification model training, and significantly improving the recognition ability of intrusion detection systems for variant attacks and zero-day attacks.

[0054] This application further proposes a process for screening virtual attack samples through the discriminator of the generative adversarial network, including mixing virtual attack samples and real attack samples and inputting them into the discriminator, setting a authenticity threshold. If the output of the discriminator exceeds the threshold, the sample is retained, and the generator and discriminator are alternately trained until the JS divergence between the generated samples and the real samples converges to a preset range.

[0055] Among them, the authenticity threshold α refers to the critical value for the discriminator to judge the authenticity of samples. Specifically, a numerical range between 0.7 and 0.9 can be used to screen samples with high credibility considered by the discriminator and avoid interference from low-quality samples in the training process. The JS divergence is an index to measure the difference between the distribution of generated samples and the distribution of real samples. Specifically, it can be achieved by calculating the similarity of the two distributions. When the JS divergence converges to a preset range such as 0.05 to 0.15, it indicates that the distributions of generated samples and real samples tend to be consistent in the feature space. Alternate training means that the generator and discriminator are alternately optimized according to the rules of adversarial games. Specifically, it can be achieved by iteratively updating network parameters through the backpropagation algorithm, forcing the generated samples to continuously approach the distribution characteristics of real samples.

[0056] Specifically, in the screening stage, the virtual attack samples generated by the generator are mixed with the real attack samples and then input into the discriminator. The discriminator scores the authenticity of each generated sample. By setting a threshold α, only the samples with scores higher than the threshold are retained, thereby filtering out the noise data and ensuring the effectiveness of the samples in the extended dataset. In the training stage, the generator and the discriminator are alternately optimized: the generator improves the generation quality by minimizing the discriminator's ability to recognize the generated samples, while the discriminator maintains the screening criteria by maximizing the ability to distinguish between real samples and generated samples. During the training process, the JS divergence is continuously monitored. When this metric converges to a preset range, it indicates that the distribution of the generated samples has been sufficiently close to the real attack traffic. At this time, the training is stopped to avoid overfitting.

[0057] Compared with the prior art, existing methods usually directly use the unfiltered generated samples to expand the dataset, resulting in uneven data quality and a lack of quantitative control over the distribution consistency. This solution realizes the active quality control of the generated samples through a dynamic threshold screening mechanism and JS divergence monitoring. At the same time, it uses an adversarial training framework to establish a closed-loop optimization of generation and screening, solving the problem of insufficient sample effectiveness of traditional generative adversarial networks in the industrial control scenario.

[0058] Through the above technical solution, this application can significantly improve the generation quality of virtual attack samples, ensure that the extended dataset contains samples with the same characteristics as the real attack traffic, thereby alleviating the problem of insufficient model training caused by the scarcity of attack samples in industrial control system intrusion detection. At the same time, through adversarial training and distribution convergence control, the distribution deviation between the generated samples and the real samples is avoided, enhancing the stability and reliability of the data augmentation process.

[0059] This application further proposes a process of inputting the feature vector X into the first classifier and outputting the first classification probability. The tree model includes XGBoost, random forest or decision tree algorithms. The input feature vector X is passed along the tree structure of the tree model from the root node to the leaf node, and the branch path is selected according to the improved splitting rule. The output weight wk of the leaf node of the kth tree is determined by optimization in the training stage. After passing through K trees, the cumulative score S is obtained, and the cumulative score is converted into an output probability through the improved sigmoid function.

[0060] Among them, the improved splitting rule refers to a decision rule that enhances the sensitivity to sparse attack samples by dynamically adjusting the feature splitting threshold. Specifically, it can be implemented by adjusting the information gain calculation based on the sample weight distribution, and adaptively adjusting the selection strategy of the feature splitting point according to the sparsity of the attack samples during the training process. Among them, the output weight wk of the leaf node refers to the optimal leaf node prediction value determined by optimizing the objective function during the training stage of each decision tree. Specifically, it can be implemented by using the second-order derivative approximation method of the loss function in the gradient boosting framework, and constraining the weight size through a regularization term to prevent overfitting. Among them, the cumulative score S refers to the total score obtained by weighted summing the prediction results of multiple decision trees for the same feature vector. Specifically, it can be implemented by the cumulative tree-by-tree method in the additive model, and improving the overall classification performance by integrating the prediction results of multiple weak classifiers. Among them, the improved sigmoid function refers to a non-linear function that adjusts the probability transformation by introducing a curvature parameter. Specifically, it can be implemented by using a sigmoid deformation function with an adaptive temperature coefficient, and dynamically scaling the input value range according to the positive and negative sample ratios to correct the probability estimation deviation.

[0061] Specifically, when the feature vector X is transmitted in the tree model, the improved splitting rule dynamically adjusts the feature splitting threshold according to the distribution characteristics of the attack samples, making the tree structure more inclined to capture the key features of the attack traffic. Each tree determines the weight wk of the leaf node by optimizing the objective function during the training stage, and makes a multi-path decision on the feature vector and accumulates the prediction results of each tree to form the cumulative score S during the prediction stage. This score is non-linearly mapped through the improved sigmoid function, where the curvature parameter is dynamically adjusted according to the class distribution of the training data, effectively alleviating the probability estimation deviation problem caused by sample imbalance.

[0062] Compared with the prior art, the traditional tree model uses a fixed information gain or Gini coefficient as the splitting criterion, and it is difficult to accurately identify the splitting point of the key features when the attack samples are sparse, resulting in inaccurate classification path selection. The existing methods use the standard sigmoid function for probability transformation, and there is a systematic deviation in the output probability when the sample distribution is severely skewed. This application enhances the recognition ability of sparse samples through the dynamic splitting rule, and corrects the probability mapping relationship through the improved sigmoid function, so that accurate classification probabilities can still be generated when the positive and negative sample ratios are imbalanced.

[0063] Through the above technical solutions, the present application effectively solves the problem of distorted classification probabilities caused by scarce attack samples in industrial control system traffic detection, and improves the decision-making reliability of the tree model in the scenario of sample imbalance. The improved splitting rule enhances the model's ability to capture sparse attack features, the dynamic weight optimization mechanism improves the prediction accuracy of a single tree, and the improved probability conversion method corrects the estimation bias caused by skewed sample distribution, enabling the overall classifier to maintain stable detection performance when the attack samples are insufficient.

[0064] The process of inputting the feature vector into the second classifier and outputting the second classification probability in the present application includes: the feature vector is calculated layer by layer through the hidden layer of the generative adversarial network discriminator to output an unnormalized score; the score is converted into an output probability through the Sigmoid function, where the parameters of the Sigmoid function include the weights and biases of the output layer of the generative adversarial network discriminator.

[0065] Among them, the hidden layer of the generative adversarial network discriminator refers to a deep neural network structure composed of multiple fully connected layers, which can specifically be implemented by a deep neural network with residual connections. Each hidden layer contains an activation function and a normalization operation. The hierarchical calculation of the hidden layer can extract the non-linear features of network traffic data layer by layer, enabling the discriminator to learn the potential distribution pattern of attack traffic during the adversarial training process. The Sigmoid function refers to a mathematical function that maps the real number domain to the probability space, which can specifically be implemented by the standard Sigmoid function, and its output range is limited between 0 and 1. The weights and biases refer to the trainable parameter matrices of the output layer of the generative adversarial network discriminator, which can specifically be synchronously optimized through the backpropagation algorithm during the adversarial training process, enabling the discriminator to simultaneously improve the accuracy of the classification probability in the task of judging the authenticity of samples.

[0066] Specifically, after the feature vector is input into the generative adversarial network discriminator, it undergoes non-linear transformation through three hidden layers in sequence. Each hidden layer consists of 512 neurons, and uses the ReLU activation function and batch normalization operation. The feature vector output by the hidden layer is reduced in dimension through a fully connected layer to generate an unnormalized classification score. This score is converted into a probability value in the range of 0 to 1 through the Sigmoid function, where the dimension of the weight matrix is 256×1 and the bias term is a scalar parameter. During the adversarial training process, the discriminator updates its parameters by alternately receiving real attack samples and generated samples, enabling the collaborative optimization of the hidden layer feature extraction ability and the output layer probability calculation ability. The gradient penalty strategy is adopted during the training process to constrain the magnitude of the discriminator's weight update and ensure the training stability.

[0067] Compared with the prior art, the discriminator in the traditional generative adversarial network is only used to distinguish between generated samples and real samples, and the features extracted by its hidden layer are not directly used for classification tasks. In this solution, by reusing the hidden layer structure of the discriminator, while maintaining the stability of generative adversarial training, the deep feature expression ability of the discriminator is transformed into the ability to output classification probabilities. Existing methods require separate training of a classifier, resulting in the separation of feature extraction and classification tasks, while this solution realizes the dual reuse of the discriminator function and reduces the model complexity.

[0068] Through the above technical solution, this application effectively solves the defect that the discriminator of the generative adversarial network only serves as a sample screening tool in traditional applications and fails to fully utilize its deep feature expression ability. By reusing the hidden layer calculation structure of the discriminator and synchronously optimizing the classification probability output parameters during the adversarial training process, the discriminator can accurately output the classification confidence of attack traffic while completing the sample screening task. This method avoids the training overhead of an additional classifier, improves the model training efficiency, and at the same time ensures the synchronous improvement of the quality of generated samples and classification performance.

[0069] This application further proposes a process of dynamically adjusting the output weights of the first classifier and the second classifier through the Mamba model, including real-time collection of long-sequence network traffic data of the industrial control system; analyzing the time dependence of the historical classification results of the decision tree model and the GAN discriminator in the long-sequence network traffic data; according to the time dependence of the historical classification results, performing sliding window weighting on the output of the current classifier and adaptively allocating the output weights of each classifier; the window length is positively correlated with the time span of the attack pattern.

[0070] Among them, the Mamba model refers to a time series analysis module constructed based on the state space model, which can be specifically implemented by a sequence modeling algorithm with linear computational complexity and is used to capture the time evolution law of the historical output of the classifier. The sliding window weighting mechanism refers to intercepting a sequence of historical classification results of a fixed length in chronological order, which can be specifically implemented by exponential decay weighting or a dynamic convolution kernel and is used to enhance the weight influence of recent classification performance. The positive correlation between the window length and the attack pattern means setting the window size parameter according to the duration characteristics of the attack behavior. For example, a longer window period is set for APT attacks and a shorter window period is set for DoS attacks, and a mapping relationship between the attack type and the window length is established through a time series matching algorithm.

[0071] Specifically, the industrial control system network traffic forms a time - continuous detection data stream after real - time collection. The Mamba model performs state - space modeling on the historical classification results of the tree model and the GAN discriminator, and extracts the performance fluctuation characteristics of the classifier at different time periods through the hidden state transition equation. For the current detection moment, a sliding window is used to intercept the classification results of the previous N time steps, and the dynamic weight coefficients of each classifier are generated through weighted summation. For example, when detecting a persistent attack, the Mamba model identifies that the true positive rate of the GAN discriminator has been continuously higher than that of the tree model in the past 5 minutes, and then automatically increases the weight ratio of the discriminator. The window length parameter is dynamically adjusted according to the attack type. For example, for an APT attack with a latency of several weeks, a 30 - day window is used to capture the long - term behavior pattern; for a sudden DoS attack, a 5 - minute short window is used to quickly respond to instantaneous traffic changes.

[0072] Compared with the prior art, the traditional intrusion detection system adopts a classifier combination strategy with fixed weights and cannot adapt to the temporal variation characteristics of attack patterns. This solution models the time - dependence of the classifier output by introducing a state - space model and combines an adaptive sliding window mechanism to dynamically optimize the classifier weight allocation strategy according to the duration characteristics of attack behaviors. In the prior art, the size of the sliding window is mostly a fixed value, while this solution realizes the dynamic matching of detection parameters and attack characteristics by establishing an association model between the attack pattern and the window length.

[0073] Through the above technical solutions, this application effectively solves the problem of dynamic attack adaptability caused by static weight allocation of classifiers in industrial control system intrusion detection. By capturing the temporal evolution law of classifier performance through time - series modeling and combining the window adjustment mechanism adaptive to attack characteristics, the detection sensitivity of the integrated model to persistent attacks and instantaneous attacks is significantly improved. The dynamic weight allocation strategy enables the model to automatically adjust the contribution degree of the classifier according to the real - time detection scenario and maintain stable detection performance in the complex and changeable industrial network environment.

[0074] This application further proposes a process of optimizing the calculation efficiency through the ReLU linear attention mechanism, including block - processing of the traffic feature vector, matrix normalization, similarity calculation, and optimization of the calculation order.

[0075] Among them, multi-head partitioning means dividing a high-dimensional feature vector into multiple sub-blocks for parallel processing. Specifically, it can be achieved by using tensor deformation operations to convert the input vector into multi-channel sub-blocks, which is used to reduce the computational complexity of high-dimensional features. L2 normalization means normalizing the row vectors of a matrix to unit length. Specifically, it can be achieved by calculating the L2 norm of the row vectors and then performing a division operation, which is used to eliminate the interference of vector length on similarity calculation. Cosine similarity means measuring the feature correlation degree through the cosine value of the vector angle. Specifically, it can be achieved by using the dot product operation of the normalized vectors, which is used to capture the potential correlation between different sub-blocks. ReLU activation means applying a non-linear filter to the projection matrix. Specifically, it can be achieved by setting a threshold to set negative values to zero, which is used to enhance the sparsity of feature representation. Matrix multiplication associative law optimization means adjusting the operation order to reduce the dimension of the intermediate matrix. Specifically, it can be achieved by preferentially calculating the product of the key-value matrix and the feature vector, which is used to reduce memory occupancy.

[0076] Specifically, the traffic feature vector is first divided into multiple sub-blocks, and each sub-block is independently processed by L2 normalization. The normalized query matrix and key matrix calculate the cosine similarity between sub-blocks through dot product, and then the ReLU activation function is applied to filter out low-correlation features. By changing the calculation order of the traditional attention mechanism, the original QK^T operation is reconstructed into K^TV for priority calculation, significantly reducing the dimension of the intermediate matrix. The combination of this sub-block processing and calculation order optimization can reduce the computational complexity from quadratic to linear level while maintaining the feature expression ability.

[0077] In some specific embodiments, the number of sub-blocks can be dynamically adjusted according to the hardware parallel computing ability. For example, when the GPU video memory capacity is limited, the number of sub-blocks can be set to 8 or 16; the threshold parameter of the ReLU activation function can be adaptively determined according to the training data distribution; the matrix multiplication order optimization can be achieved by using tensor reshaping operations. For example, a three-dimensional tensor is reconstructed into a two-dimensional matrix for batch calculation.

[0078] Compared with the prior art, the traditional multi-head attention mechanism directly calculates the global feature correlation degree, which has the defects of high computational complexity and large memory consumption. This solution restricts the calculation range through sub-block processing, combines L2 normalization to eliminate dimensional bias, and uses ReLU activation to achieve feature selection. On the basis of retaining the local feature interaction ability, the theoretical calculation amount is reduced by about two orders of magnitude. Compared with the intermediate matrix that must be stored in the standard attention mechanism, the optimized calculation order only needs to retain the low-dimensional intermediate results, and the memory occupancy is reduced by more than 60%.

[0079] Through the above technical solutions, the present application effectively solves the problem of high computational load in the real-time detection scenario of industrial environments, and can achieve fast feature extraction of long-sequence traffic data under the resource constraints of embedded devices. The block processing mechanism enhances the model's ability to capture local attack features, the sparsification characteristic of ReLU activation suppresses noise interference, and the optimization of the calculation order significantly improves the system throughput, providing a feasible lightweight implementation solution for industrial control system intrusion detection.

[0080] The present application further proposes a process for determining whether network traffic data is attack traffic. The secondary classifier is a logistic regression or support vector machine model. The output probabilities of the first classifier and the second classifier are concatenated into a two-dimensional vector Fstack and input into the secondary classifier to output the final classification confidence S. If S ≥ θ (a preset threshold), it is determined as attack traffic; otherwise, it is determined as normal traffic.

[0081] Among them, the secondary classifier refers to an ensemble model used to comprehensively output the basic classifiers, and specifically, it can be implemented using logistic regression or support vector machine algorithms. Its role is to perform secondary discrimination on the basic classification results through linear or non-linear decision boundaries to solve the problem of insufficient generalization ability of a single classifier. The two-dimensional vector Fstack refers to a feature combination composed of the probability outputs of two basic classifiers, and specifically, it can be achieved through a horizontal concatenation operation. Its role is to fuse the discrimination information of different models into a unified feature representation to enhance the separability of the classification boundary. The classification confidence S refers to the probability value that the sample output by the secondary classifier belongs to the attack category, and specifically, it can be calculated through linear weighting or kernel function mapping. Its role is to provide a quantifiable decision basis for the final determination. The preset threshold θ refers to the critical value for determining attack traffic, and specifically, it can be set according to the balance requirements of the false alarm rate and the missed detection rate. Its role is to adapt to the detection requirements of different industrial control environments by dynamically adjusting the classification sensitivity.

[0082] Specifically, after the first classifier and the second classifier respectively output the attack probabilities, the results of the two are combined into a two-dimensional vector, which is input into the secondary classifier for final discrimination. The secondary classifier learns the complementary relationship between the basic classifiers during the training phase, and uses the linear decision of logistic regression or the kernel trick of support vector machine to construct a more robust classification boundary in the fused probability space. The finally output confidence S is compared with the preset threshold. When the confidence exceeds the threshold, an attack alarm is triggered; otherwise, it is marked as normal traffic.

[0083] In some specific embodiments, the weight parameters of the logistic regression model can be optimized through the cross-entropy loss function, the kernel function of the support vector machine model can be selected as the radial basis function, and the preset threshold θ can be determined by the optimal operating point of the ROC curve.

[0084] Compared with the prior art, existing methods usually use fixed weights or a single classifier for final determination, resulting in classification bias in dynamic attack modes. This solution effectively captures the temporal correlation features of attack behaviors by introducing a secondary classifier to perform secondary learning on the dynamic weight output of the basic model. At the same time, a probability splicing mechanism is used to enhance the feature expression ability, reducing missed detections or false alarms caused by misjudgments of a single model.

[0085] Through the above technical solution, this application can improve the discrimination accuracy of attack traffic in the industrial control environment, reduce the misjudgment risk caused by network traffic fluctuations or attack variants, and at the same time achieve balanced control of detection sensitivity and false alarm rate through a dynamic threshold mechanism, meeting the real-time detection requirements in complex industrial scenarios.

[0086] Embodiment 2:

[0087] As Figure 4 shown, the industrial control system intrusion detection system based on Mamba-ReLU includes:

[0088] Traffic acquisition module: used to acquire the network traffic data of the industrial control system, and extract traffic data features from the network traffic data to form real samples;

[0089] Sample generation module: used to generate virtual attack samples Xfake through a GAN generator, screen the virtual attack samples Xfake through a GAN discriminator, and merge them with real samples to form extended samples;

[0090] Model construction module: used to construct an enhanced classification model based on stacked integration, and the enhanced classification model includes a first classifier based on a tree model, a second classifier based on a GAN discriminator, and a secondary classifier;

[0091] Basic classification module: used to obtain the feature vector X of the extended sample, input the feature vector X into the first classifier, output the first classification probability, and synchronously input it into the second classifier to output the second classification probability;

[0092] Optimization and adjustment module: used to optimize the calculation efficiency through the ReLU linear attention mechanism, and dynamically adjust the output weights of the first classifier and the second classifier through the Mamba model;

[0093] Stacked integration module: used to input the first classification probability and the second classification probability into the secondary classifier according to the output weights, and determine whether the network traffic data is attack traffic based on the output result of the secondary classifier.

[0094] The above content is only an example and explanation of the structure of the present invention. Those skilled in the art of this technology can make various modifications or supplements to the described specific embodiments or use similar ways to replace them. As long as they do not deviate from the structure of the invention or exceed the scope defined by this claim book, they should all fall within the protection scope of the present invention.

[0095] In the description of this specification, the description referring to terms such as "an embodiment", "example", "specific example", etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0096] The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The preferred embodiments do not elaborate on all details and do not limit the invention to only the specific implementation manners. Obviously, according to the content of this specification, many modifications and changes can be made. This specification selects and specifically describes these embodiments in order to better explain the principle and practical application of the present invention, so that those skilled in the art of this technology can well understand and utilize the present invention. The present invention is only limited by the claim book and its full scope and equivalents.

Claims

1. An intrusion detection method for industrial control systems based on Mamba-ReLU, characterized in that, It includes the following steps: Obtain the network traffic data of the industrial control system, and extract traffic data features from the network traffic data to form real samples; Generate a virtual attack sample Xfake through a GAN generator; Screen the virtual attack sample Xfake through a GAN discriminator, and merge it with real samples to form an extended sample; Construct an enhanced classification model based on stacked integration. The enhanced classification model includes a first classifier based on a tree model, a second classifier based on a GAN discriminator, and a secondary classifier; Obtain the feature vector X of the extended sample, input the feature vector X into the first classifier, output the first classification probability, and synchronously input it into the second classifier to output the second classification probability; Optimize the calculation efficiency through the ReLU linear attention mechanism, and dynamically adjust the output weights of the first classifier and the second classifier through the Mamba model; Input the first classification probability and the second classification probability into the secondary classifier according to the output weights, and determine whether the network traffic data is attack traffic based on the output result of the secondary classifier.

2. The method according to claim 1, wherein The traffic data features include traffic size, packet frequency, protocol type, moving window mean and variance, protocol field outliers, packet interval time series, and session duration.

3. The method according to claim 1, wherein The process of generating a virtual attack sample through a GAN generator includes: The GAN generator receives a random noise vector Z that follows a Gaussian distribution; Concatenate the noise vector Z and the labeled real attack sample Xreal in the feature dimension to generate a joint input vector [Z; Xreal]; The GAN generator uses a multi-layer long short-term memory network (LSTM) to receive the concatenated joint input vector [Z; Xreal]; Perform temporal unfolding processing through a hidden layer containing N LSTM units to obtain the hidden state of the LSTM and capture the time dependence of the network traffic data; The GAN generator maps the hidden state of the LSTM to a virtual attack sample Xfake through a fully connected layer, and its dimension is the same as that of the real attack sample Xreal.

4. The method according to claim 1, characterized in that, The process of screening virtual attack samples through a GAN discriminator includes: Mix the virtual attack sample Xfake and the real attack sample Xreal and input them into the GAN discriminator. Set a authenticity threshold α. If the discriminator output D(Xfake)≥α, it is determined that the generated virtual attack sample Xfake is a high-quality sample and is retained; The GAN generator and the GAN discriminator are alternately trained until the Jensen-Shannon Divergence (JS divergence) between the distribution of the virtual attack sample Xfake and the real attack sample Xreal converges to a preset range.

5. The method according to claim 1, characterized in that The process of inputting the feature vector X into the first classifier and outputting the first classification probability includes: The tree model includes XGBoost, random forest, or decision tree algorithm; Input the feature vector X and transfer it along the tree structure of the tree model from the root node to the leaf node, and select the branch path according to the improved splitting rule; The output weight wk of the leaf node of the kth tree is determined by optimization in the training stage; After passing through K trees, the cumulative score S is obtained, and the cumulative score is converted into an output probability through an improved sigmoid function.

6. The method according to claim 1, wherein The process of inputting the feature vector X into the second classifier and outputting the second classification probability includes: The feature vector X passes through the hidden layer of the GAN discriminator for layer-by-layer calculation, and an unnormalized score is output. The score is converted into an output probability through the Sigmoid function p = σ(W * h + b), where h is the output of the last hidden layer, W and b are the weights and biases of the output layer of the GAN discriminator, and σ is the Sigmoid function.

7. The method according to claim 1, characterized in that, The process of dynamically adjusting the output weights of the first classifier and the second classifier through the Mamba model includes: Real-time collection of long-sequence network traffic data of the industrial control system; Analyze the time dependence of the historical classification results of the tree model and the GAN discriminator in the long-sequence network traffic data; According to the time dependence of the historical classification results, perform sliding window weighting on the output of the current classifier, and adaptively allocate the output weights of each classifier; The window length is positively correlated with the time span of the attack pattern.

8. The method according to claim 1, wherein The process of optimizing the calculation efficiency through the ReLU linear attention mechanism includes: Perform multi-head partitioning on the traffic feature vector X, and perform L2 normalization on the row vectors of the query matrix Q and the key matrix K respectively; Calculate the cosine similarity between the normalized blocks; Apply the ReLU activation on the projected query matrix Q and key matrix K; Using the associative law of matrix multiplication, decompose the calculation order into: Attention(Q, K, V) = Q(K T V).

9. The method according to claim 1, characterized in that The process of determining whether the network traffic data is attack traffic includes: The secondary classifier is a logistic regression or support vector machine model; Concatenate the output probabilities of the first classifier and the second classifier into a two-dimensional vector Fstack, and input it into the secondary classifier to output the final classification confidence S; If S ≥ θ (preset threshold), it is determined as attack traffic; Otherwise, it is determined as normal traffic.

10. An industrial control system intrusion detection system based on Mamba-ReLU, characterized in that, Including: Traffic acquisition module: used to acquire the network traffic data of the industrial control system, and extract traffic data features from the network traffic data to form real samples; Sample generation module: used to generate virtual attack samples Xfake through the GAN generator, screen the virtual attack samples Xfake through the GAN discriminator, and merge them with real samples to form extended samples; Model construction module: used to construct an enhanced classification model based on stacked integration, and the enhanced classification model includes a first classifier based on a tree model, a second classifier based on a GAN discriminator, and a secondary classifier; Basic classification module: used to obtain the feature vector X of the extended sample, input the feature vector X into the first classifier, output the first classification probability, and synchronously input it into the second classifier to output the second classification probability; Optimization and adjustment module: used to optimize the calculation efficiency through the ReLU linear attention mechanism, and dynamically adjust the output weights of the first classifier and the second classifier through the Mamba model; Stacked integration module: used to input the first classification probability and the second classification probability into the secondary classifier according to the output weights, and determine whether the network traffic data is attack traffic according to the output result of the secondary classifier.

Citation Information

Cited By

  • LLM-DoS attack protection method based on multi-level defense strategy and related device

    CN121309150A

  • Malicious domain name detection system fusing multi-modal embedding and dynamic weight Mama

    CN122226480A