Industrial internet intrusion detection method based on diffusion model

By applying data augmentation method and feature extraction technology based on diffusion model in industrial Internet intrusion detection, the problems of data category imbalance and high-dimensional timing data processing are solved, and efficient and accurate attack recognition and system robustness are achieved.

CN120050096APending Publication Date: 2025-05-27ZHENGZHOU UNIVERSITY OF LIGHT INDUSTRY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510200525.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the industrial Internet environment, traditional intrusion detection methods are difficult to effectively process high-dimensional and highly sequential network traffic data, especially when facing data categories imbalance, it is difficult to achieve efficient and accurate attack identification.

Method used

A method of intrusion detection based on diffusion model is proposed. Synthetic samples are generated through multi-layer diffusion models, data distribution is balanced, and feature extraction and dimensionality reduction are combined with 1D ResCNN and multi-head attention mechanism. Finally, multi-dimensional timing features are extracted through multi-layer bidirectional SRU module for intrusion detection.

Benefits of technology

The model's ability to identify a few types of attacks is improved, the system's robustness is enhanced, the feature extraction problem of high-dimensional network traffic data is solved, and the timing and efficient processing of network traffic data is realized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050096A_ABST
    Figure CN120050096A_ABST
Patent Text Reader

Abstract

The invention provides an industrial internet intrusion detection method based on a diffusion model. The method comprises the following steps: cleaning and standardizing original industrial internet flow data; taking the standardized industrial internet traffic data as input, generating a synthetic sample for the minority class of samples by adopting a multi-layer diffusion model, and sequentially carrying out anti-standardization and format adjustment processing and data balance sampling on the synthetic sample to obtain balance sample data; performing feature extraction and dimension reduction on the balance sample; the local time sequence features are processed through a multi-head attention mechanism, and global dependency features of the weighted key features are obtained; taking the global dependency feature of the weighted key feature as input, and extracting a multi-dimensional time sequence feature containing a long-term time dependency feature through a multi-layer bidirectional SRU module; and performing intrusion detection classification according to the multi-dimensional time sequence characteristics. According to the method, attack recognition can be efficiently and accurately carried out when high-dimensional and high-time-sequence network flow data and data category imbalance are carried out.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intrusion detection, and in particular to an industrial Internet intrusion detection method. Background Art

[0002] With the rapid development of the Industrial Internet, industrial equipment and control systems are widely connected, building a complex and dynamic ecosystem. While improving production efficiency and automation levels, it also brings severe network security challenges. The Industrial Internet involves a large amount of real-time data transmission and sensitive operations. Once it is attacked by a network, it may cause production interruption, equipment damage, and even threaten social security. Therefore, ensuring its security has become a research focus, and intrusion detection is a key protection method.

[0003] Traditional intrusion detection methods are often based on rules and feature engineering, which are difficult to meet the requirements of industrial Internet intrusion detection. First, the imbalance of industrial Internet data is a prominent problem. Normal traffic data is much larger than attack traffic data, which causes model training to be biased towards normal traffic, thereby reducing the attack detection rate. Secondly, the high dimensionality and complexity of industrial Internet data make its feature space huge, and traditional detection methods are difficult to effectively handle multiple attack modes. Finally, the industrial Internet intrusion detection system must have fast processing capabilities to detect and respond to potential attacks in a timely manner with a response speed of milliseconds to ensure the safety and continuity of the production process.

[0004] In response to the problem of unbalanced data in industrial control systems, A. Al-Abassi et al. proposed a DNN-DT deep learning model that combines deep neural networks (DNN) and decision trees (DT). This method can effectively improve the performance of the model in identifying minority attacks by generating new attack samples and enhancing samples of specific categories. In addition, in response to the emergence of new unknown attacks in the industrial Internet, Z. Fanyi proposed an intrusion detection method based on neighborhood filtering and stable learning, which enhances the model's detection ability for unknown attacks by generating new attack samples. In response to the problem of insufficient data samples in industrial control systems, the sample generation method combining active learning and LightGBM can selectively enhance samples based on existing data, thereby improving the accuracy and stability of the model for minority attacks. At the same time, the industrial control system network attack sample generation method combining 1D convolutional neural network (CNN) and generative adversarial network (WGAN) generates new attack samples through adversarial training, further enhancing the diversity of the data set and improving the robustness of the detection model.

[0005] Although the above methods have achieved remarkable results in improving the data imbalance problem and enhancing detection performance, they still face some challenges in practical applications. First, the generated samples may not fully capture the diversity of attack patterns in some cases, especially when dealing with unknown attacks. Traditional generative adversarial networks (GANs) can generate samples similar to the original data, but the quality and diversity of the generated samples are still limited to a certain extent. In addition, GAN-based generation methods are vulnerable to instability during the training process, resulting in the generated attack samples may not fully represent the real attack patterns. Summary of the Invention

[0006] Aiming at the technical problem that it is difficult to achieve efficient and accurate attack recognition for existing technologies in the face of high-dimensional and time-series network traffic data and data class imbalance, the present invention proposes an industrial Internet intrusion detection method based on a diffusion model, which balances the network traffic data in the industrial Internet environment, thereby improving the model's recognition ability for minority-class attacks and enhancing the robustness of the system; solves the problem of feature extraction for high-dimensional network traffic data; and realizes the efficient processing of the time series of network traffic data.

[0007] In order to achieve the above object, the technical solution of the present invention is realized as follows:

[0008] An industrial Internet intrusion detection method based on a diffusion model, comprising the steps of:

[0009] S1: Clean and standardize the original industrial Internet traffic data;

[0010] S2: Taking the standardized industrial Internet traffic data as input, generating synthetic samples for minority-class samples using a multi-layer diffusion model, and sequentially performing inverse standardization, format adjustment processing, and data balancing sampling on the synthetic samples to obtain balanced sample data;

[0011] S3: Extracting and reducing the dimensions of the balanced samples through 1D ResCNN to obtain local time-series features; processing the local time-series features through a multi-head attention mechanism to obtain global dependence features of weighted key features;

[0012] S4: Taking the global dependence features of weighted key features as input, extracting multi-dimensional time-series features containing long-term time dependence features through a multi-layer bidirectional SRU module;

[0013] S5: Classifying intrusion detection according to the multi-dimensional time-series features.

[0014] Further, the method for generating synthetic samples for minority-class samples using a multi-layer diffusion model is as follows: The initial input data is the standardized industrial Internet traffic data X. The minority-class samples are selected as the samples to be enhanced. At each diffusion layer of the multi-layer diffusion generation model, the output synthetic sample data of the previous diffusion layer is used as the input; the forward diffusion is performed through an encoder to generate noise samples The decoder is used to perform reverse denoising on the noise samples to generate synthetic sample data By generating representative minority-class attack samples, the data of each class in the training set is balanced, thereby improving the model's recognition ability for minority-class attacks and enhancing the robustness of the system.

[0015] Further, the process of the forward diffusion is as follows: The output synthetic sample data of the previous diffusion layer is used as the input, and noise is gradually added until the time step t = T ends, and finally the noise samples are obtained. The calculation method at each time step in the forward diffusion process is:

[0016]

[0017] where L represents the number of diffusion layers, is the noise-added data of the l-th diffusion layer at the t-th time step, is a specific noise term of each diffusion layer generated from a Gaussian distribution; the forward diffusion helps the model learn how to recover the original data features through reverse denoising, thereby providing a training signal for data generation;

[0018] The process of the reverse denoising is as follows: Starting from the noise samples and gradually removing the noise to generate new data consistent with the true data distribution The calculation method at each time step in the reverse denoising process is:

[0019]

[0020] where, is the cumulative noise attenuation, t is the current time step, is the noise intensity at the s-th time step, represents the noise attenuation factor, is the Gaussian noise term; the process of reverse denoising makes the data gradually return to the original distribution by reducing the noise at each step, achieving the purpose of data reconstruction.

[0021] Further, when training the multi-layer diffusion model, the mean squared error is used as the loss function to calculate the loss and update the model parameters; the calculation process is:

[0022] Calculate the loss function of the l-th diffusion layer:

[0023]

[0024] where x 0 is a minority-class sample, and E t represents the expectation;

[0025] Sum the losses of all layers with weights to obtain the total loss:

[0026]

[0027] where λ l is the weight coefficient of the loss of the l-th layer;

[0028] Update the parameters of all diffusion layers simultaneously through backpropagation. By minimizing the total loss L total , obtain the optimized multi-layer generation model, and use the optimized multi-layer generation model to obtain the final synthetic sample data By minimizing the total loss function, the model can learn how to recover high-quality samples from noise, thereby increasing the quantity and diversity of minority-class samples while maintaining the authenticity of the data.

[0029] Furthermore, the method for extracting features and reducing dimensions of balanced samples through 1D ResCNN to obtain local temporal features is as follows: Using the balanced sample data X p as the input, adopt multiple convolutional layers to extract features layer by layer. Each convolutional layer includes a 1D convolution operation, a Dropout operation, and a tanh activation function performed in sequence; perform residual connections between the outputs of some convolutional layers and the Dropout operations of other convolutional layers, and output the output of the last convolutional layer after a max-pooling operation as the local temporal feature sequence X Rs ; By effectively extracting meaningful features from temporal data while avoiding interference from redundant information;

[0030] The method for processing local temporal features through the multi-head attention mechanism to obtain the global dependence features of weighted key features is as follows: Using the local temporal feature sequence X Rs as the input, divide the local temporal feature sequence X Rs into h feature sequences X i , and each feature sequence X i corresponds to one head; Obtain the query matrix Query i of the feature sequence X i , the key matrix key i , and the value matrix Value i: Calculate the attention scores and normalized scores based on the query matrix, key matrix, and value matrix; multiply the normalized scores by the value matrix to obtain the attention values of each feature sequence X i ; Concatenate the attention values of all heads and form the global dependency feature X of the final weighted key features through a linear layer att ; By effectively extracting meaningful features from time series data while avoiding interference from redundant information, this mechanism is applied after 1D ResCNN and is mainly used to further focus on the global dependencies in the data and the long-range relationships between time series, enabling the model to focus on key time steps and assign higher weights to more important moments.

[0031] Furthermore, the method for extracting multi-dimensional features containing long-term time dependency features through a multi-layer bidirectional SRU module is as follows: perform word embedding on the global dependency feature X of the weighted key features att to obtain a low-dimensional continuous vector X in ={x 1 ,x 2 ,x t ......,x n}, where x t is the sequence feature at the t-th time step, n is the number of features, and perform forward SRU calculation and backward SRU calculation on the low-dimensional continuous vector X in respectively to obtain the forward hidden state and backward hidden state; merge the forward hidden state and backward hidden state into a comprehensive hidden state for representing multi-dimensional time series features, and extract more representative time series features from the enhanced industrial Internet data to improve the model's ability to capture the time series of attack behaviors.

[0032] Furthermore, the method for the forward SRU calculation is as follows: initialize the previous hidden state at the first time step as a zero vector, and starting from time step t = 1, obtain the forward hidden state through step-by-step SRU calculation;

[0033] The method for the backward SRU calculation is as follows: initialize the next hidden state at the last time step as a zero vector, and starting from time step t = n, obtain the backward hidden state through step-by-step SRU calculation;

[0034] The method for performing SRU calculation is as follows: in each SRU unit, perform a linear transformation on the input sequence feature x t to calculate the preliminary extracted features;

[0035]

[0036] where W is the linear transformation matrix;

[0037] Calculate the forget gate f t, determine the state vector C of the current time step according to the state vector C of the previous time step t-1 and the forget gate f t to determine the output of the state vector C of the current time step through an adaptive average t , the state vector C of the current time step t The calculation formula is as follows:

[0038] f t =σ·(W f x t +b f )

[0039]

[0040] where W f and b f are the weight and bias of the forget gate respectively, σ is the activation function, and ⊙ represents element-wise multiplication;

[0041] Calculate the reset gate r t , input the state vector C of the current time step t into the non-linear activation function, adaptively combine the input sequence feature x t and the output of the non-linear activation function according to the forget gate, and update the current hidden state h t , update the current hidden state h t The calculation formula is as follows:

[0042] r t =σ·(W r x t +b r )

[0043] h t =r t ⊙g(C t )+(1 - r t )⊙x t

[0044] where W r and b r are the reset gate weight and bias respectively, and g() is the non-linear activation function.

[0045] An industrial Internet intrusion detection model based on a diffusion model, including a data preprocessing module, a data enhancement module based on a multi-layer diffusion model, a feature extraction and dimensionality reduction module based on a ResCNN-Attention hybrid network, a time series modeling module based on BISRU, and an intrusion detection classification module connected in sequence;

[0046] Among them, the data preprocessing module is used to clean and standardize the original industrial Internet traffic data;

[0047] The data augmentation module based on the multi-layer diffusion model is used to take the standardized industrial Internet traffic data as input, generate synthetic samples for minority class samples using the multi-layer diffusion model, and perform inverse standardization, format adjustment processing, and data balancing sampling on the synthetic samples in sequence to obtain balanced sample data;

[0048] The feature extraction and dimensionality reduction module based on the ResCNN-Attention hybrid network is used to extract features and reduce the dimensionality of the balanced samples through 1D ResCNN to obtain local temporal features; process the local temporal features through the multi-head attention mechanism to obtain the global dependency features of weighted key features;

[0049] The temporal modeling module based on BISRU is used to take the global dependency features of weighted key features as input, and extract multi-dimensional temporal features including long-term time dependency features through the multi-layer bidirectional SRU module;

[0050] The intrusion detection classification module is used to perform intrusion detection classification according to the multi-dimensional temporal features.

[0051] Specifically, the data augmentation module based on the multi-layer diffusion model includes a diffusion model input layer, a multi-layer diffusion layer, a data synthesis layer, and a diffusion model output layer connected in sequence; the diffusion model input layer is connected to the data preprocessing module; each diffusion layer includes an encoder and a decoder connected in sequence, and the input of the encoder of the first diffusion layer is connected to the diffusion model input layer; in the middle diffusion layer and the last diffusion layer, the input of the encoder is connected to the output of the previous layer decoder, and the output of the decoder of the last diffusion layer is connected to the data synthesis layer;

[0052] The feature extraction and dimensionality reduction module based on the ResCNN-Attention hybrid network includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, and a max pooling layer connected in sequence. The first convolutional layer is connected to the diffusion model output layer. Each convolutional layer includes a 1D convolution operation, a Dropout operation, and a tanh activation function performed in sequence; the output of the first convolutional layer is connected to the Dropout operation result of the third convolutional layer for residual connection, and the output of the second convolutional layer is connected to the Dropout operation result of the fourth convolutional layer for residual connection; the output of the max pooling layer is connected to the multi-head attention module.

[0053] Specifically, the temporal modeling module based on BISRU includes an embedding layer, a multi-layer bidirectional SRU structure, and an output layer connected in sequence; the input of the embedding layer is connected to the multi-head attention module, and the output of the embedding layer is connected to the multi-layer bidirectional SRU module;

[0054] The described multi-layer bidirectional SRU module includes multiple layers of bidirectional SRU layers connected in sequence. Each layer of bidirectional SRU layer includes a forward SRU layer and a backward SRU layer that are processed in parallel. The forward SRU layer includes multiple SRU units connected in sequence in the forward direction.

[0055] The backward SRU layer includes multiple SRU units connected in sequence in the reverse direction. The SRU units in the forward SRU layer and the backward SRU layer of the first layer of bidirectional SRU layer are all connected to the output of the embedding layer. The SRU units in the forward SRU layer and the backward SRU layer of the intermediate layer of bidirectional SRU layers are all connected to the corresponding SRU units in the forward SRU layer and the backward SRU layer of the previous layer of bidirectional SRU layer. The SRU units in the forward SRU layer and the backward SRU layer of the last layer of bidirectional SRU layer are all connected to the intrusion detection classification module.

[0056] Advantages of the present invention:

[0057] An industrial Internet intrusion detection method based on a diffusion model is proposed. Combining a multi-layer diffusion generation model, by adopting the CRA (i.e., CNN-ResNet-Attention) structure, which integrates a convolutional neural network (CNN), a residual network (ResNet), and a multi-head attention mechanism (Attention), as well as BISRU (Bidirectional Interactive Sequential Residual Unit), it solves the technical problem that it is difficult to achieve efficient and accurate attack recognition in the industrial Internet environment when facing high-dimensional and strongly time-series network traffic data and data class imbalance.

[0058] An industrial Internet data augmentation method for a multi-layer diffusion generation model is proposed. By means of a diffusion mechanism, new and representative minority-class attack samples are generated, thereby realizing data augmentation and balancing the distribution of various samples in the training set. Through this data augmentation strategy, the robustness of the model in the face of data imbalance is significantly improved, and the detection ability of minority-class attacks is enhanced, thus effectively improving the accuracy and stability of intrusion detection.

[0059] Based on the above data augmentation method, a method for feature extraction and dimensionality reduction of industrial Internet intrusion detection data based on ResCNN-Attention is proposed. It integrates the convolutional neural network (CNN), residual network (ResNet), and multi-head attention mechanism (Attention) to efficiently extract the features of network traffic data and solves the problem of gradient disappearance in the training of deep networks through residual connections. The attention mechanism helps the model focus on key features and further improves the accuracy of feature extraction. Through this feature extraction and dimensionality reduction method, the system can not only improve the computational efficiency but also provide a more accurate and reliable feature representation for subsequent attack type detection, thus effectively supporting the accuracy of the intrusion detection task.

[0060] Based on the feature extraction and dimensionality reduction method, BISRU is used to process the forward and backward time series information in network traffic data, comprehensively capturing the long-term time-dependent features in the samples and improving the detection ability for abnormal behaviors. By integrating multi-level features, this method significantly enhances the processing ability for complex time series data and achieves high accuracy in the detection of various attack types. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0062] Figure 1 It is the flowchart of the method of the present invention.

[0063] Figure 2 It is the overall structure diagram of an industrial Internet intrusion detection model based on the diffusion model of the present invention.

[0064] Figure 3 It is the schematic diagram of the data augmentation module structure based on the multi-layer diffusion model of the present invention.

[0065] Figure 4 It is the diffusion process of the single-layer diffusion generation model of the present invention.

[0066] Figure 5 It is the schematic diagram of the feature extraction and dimensionality reduction module structure based on the ResCNN-Attention hybrid network of the present invention.

[0067] Figure 6 It is the network structure of the residual connection of the present invention.

[0068] Figure 7Schematic diagram of the multi-head attention mechanism of the present invention.

[0069] Figure 8 It is a schematic diagram of the structure of BISRU of the present invention. DETAILED DESCRIPTION

[0070] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0071] Example 1, an industrial Internet intrusion detection method based on a diffusion model, such as Figure 1 As shown, the steps include:

[0072] S1: Clean and standardize the raw industrial Internet traffic data.

[0073] Industrial Internet traffic data usually processed by industrial Internet intrusion detection systems are high-dimensional time series data, which may contain noise, missing values, and inconsistent feature distributions. Therefore, good data preprocessing can provide clean, consistent, and standardized inputs for subsequent model training to ensure the effectiveness of the generated model.

[0074] In the present invention, data preprocessing includes two steps: data cleaning and standardization. First, data cleaning is performed to remove duplicate samples, fill missing values, and remove outliers to ensure the integrity and consistency of the data set. Then, each feature is processed by data standardization, and its mean is adjusted to 0 and the standard deviation is adjusted to 1, thereby eliminating the influence of dimensional differences between different features and making each feature contribute equally to the model.

[0075] After data preprocessing, the standardized industrial Internet traffic data is passed as input to the diffusion generation model for subsequent data enhancement operations.

[0076] S2: Using the standardized industrial Internet traffic data as input, a multi-layer diffusion model is used to generate synthetic samples for minority samples, and the synthetic samples are de-standardized, format adjusted, and data balanced sampled in turn to obtain balanced sample data.

[0077] Specifically, Figure 3 The method of generating synthetic samples using a multi-layer diffusion model for minority class samples is as follows:

[0078] The initial input data is the standardized industrial Internet traffic data X (including majority class samples and minority class samples x0 ) Select a small number of samples as the samples to be enhanced. At each diffusion layer of the multi-layer diffusion generation model, synthesize sample data using the output of the previous diffusion layer as the input (the input of the first layer is the small number of samples x 0 );Generate noise samples through forward diffusion by the encoder

[0079] Furthermore, as Figure 4 shown, the process of the forward diffusion is: synthesize sample data using the output of the previous diffusion layer as the input, gradually add noise, and end at time step t = T, finally obtaining the noise sample The calculation method at each time step during the forward diffusion process is:

[0080]

[0081] where L represents the number of diffusion layers, which is 4 layers in this embodiment, is the noise-added data of the l-th diffusion layer at the t-th time step, is a parameter that controls the noise intensity, which is a random number between [0, 1], representing the noise intensity of the l-th diffusion layer at the t-th time step, and is designed as a decreasing sequence so that the noise gradually increases; is a specific noise term of each diffusion layer generated from the Gaussian distribution, making the data gradually approach pure noise. The forward diffusion helps the model learn how to restore the original data features through reverse denoising, thereby providing training signals for data generation.

[0082] Furthermore, use the decoder to perform reverse denoising on the noise sample to generate synthetic sample data

[0083] Furthermore, as Figure 4 shown, the process of the reverse denoising is: starting from the noise sample remove the noise step by step to generate new data consistent with the real data distribution. This denoising process restores the data to the original distribution by reducing the noise at each step, achieving the purpose of data reconstruction. The calculation method at each time step during the reverse denoising process is:

[0084]

[0085] where, For cumulative noise attenuation, it is used to represent the noise intensity cumulatively layer by layer at each step in the forward diffusion process. Here, t is the current time step, s represents each time step in the diffusion process (i.e., each diffusion operation), and s is an intermediate variable used in the cumulative product or summation, which is used to represent each time step from the initial time step to the current time step. is the noise intensity at the s-th time step, which is a random number between [0, 1]. represents the noise attenuation factor. is the Gaussian noise term.

[0086] The multi-layer diffusion mechanism can generate samples with rich diversity by independently introducing noise and performing denoising operations in multiple diffusion layers. The diffusion process of each layer has independent noise control to ensure that the samples generated by different layers have different characteristics and increase the diversity of data. Through multi-layer diffusion, the generative model can capture different levels of information in the data, and on the basis of maintaining the original characteristics of the data, increase the diversity of minority-class data, thereby effectively improving the classification ability of the model.

[0087] Furthermore, during the training of the multi-layer diffusion model, the mean squared error (MSE) is used as the loss function to calculate the loss and update the model parameters; the mean squared error is used to measure the difference between the synthesized samples and the real samples.

[0088] The calculation process of the loss function is as follows: Calculate the loss function of the l-th layer:

[0089]

[0090] where, x 0 is the real sample, that is, the minority-class sample in the standardized industrial Internet traffic data, and E t represents the expectation.

[0091] The total loss is obtained by weighted summation of the losses of all layers:

[0092]

[0093] where, L is the number of layers of the multi-layer diffusion generative model, and λ l is the weight coefficient of the loss of the l-th layer.

[0094] Furthermore, the parameters of all diffusion layers are updated simultaneously through backpropagation, so that each layer collaboratively improves the generation quality. By minimizing the total loss, an optimized multi-layer generative model is obtained, and the final synthesized sample data is obtained by using the optimized multi-layer generative model By minimizing the total loss function, the model can learn how to recover high-quality samples from noise, so as to increase the quantity and diversity of minority-class samples while maintaining the authenticity of the data.

[0095] Further, perform inverse normalization and format adjustment processing: Obtain the mean and standard deviation from the standardization process of the original industrial Internet traffic data. These two parameters are usually calculated and saved when standardizing the original industrial Internet traffic data. According to the inverse normalization formula, perform an inverse transformation on the synthetic samples to restore them to the same numerical range as the original samples; According to the target format determined by the format of the samples in the original industrial Internet traffic data, adjust the number of decimal places of the inverse-normalized samples; Pay special attention to the number of decimal places during format adjustment to determine the target format to which it needs to be adjusted. For example, if the values in the original dataset retain two decimal places, then the generated samples also need to be adjusted to retain two decimal places. Inverse normalization and format adjustment ensure the usability of the generated samples, enabling them to be seamlessly integrated into the original dataset. The inverse normalization formula is as follows:

[0096]

[0097] where, X orign is the data after inverse normalization, σ h is the standard deviation, and μ h is the mean. Through this formula, the standardized data can be restored to the same numerical range as the original data.

[0098] Further, perform data balancing sampling, merge the data after inverse normalization and format adjustment into the standardized industrial Internet traffic data X, and perform sample quantity limitation to obtain balanced sample data X p ; For example, for various types of attack samples, uniformly set a limit of 1800 samples. Through this balancing strategy, the model can learn with a more balanced sample distribution during the training process, thereby enhancing the recognition accuracy of various attack behaviors.

[0099] S3: Through 1D ResCNN, extract features and reduce the dimension of the balanced samples to obtain local temporal features; Process the local temporal features through the multi-head attention mechanism to obtain the global dependency features of the weighted key features; This method constructs a combined model of a convolutional neural network (ResCNN) and an attention mechanism (Attention) to automatically filter and compress redundant features in the network traffic data while retaining the key information crucial for attack detection. Through this feature dimension reduction method, the system can not only improve the computing efficiency but also provide a more accurate and reliable feature representation for subsequent attack type detection, thereby effectively supporting the accuracy of the intrusion detection task.

[0100] Further, as Figure 5 shown, the method of extracting features and reducing the dimension of the balanced samples through 1D ResCNN to obtain local temporal features is specifically as follows:

[0101] Using the balanced sample data X p as the input, multiple convolutional layers are used to extract features layer by layer. Each convolutional layer includes a 1D convolutional operation, a 20% Dropout operation, and a tanh activation function performed in sequence. The 1D convolutional operation uses 64 filters of size 3. The output of some convolutional layers is connected in residual with the output of the Dropout operation of another part of the convolutional layers. The output of the last convolutional layer is output as the local temporal feature sequence X after passing through the max pooling operation (MaxPooling). Rs ; The 20% Dropout operation is used to randomly ignore the outputs of some neurons to prevent overfitting, retain key feature information, and enhance the generalization ability of the model. The tanh activation function is used to increase the non-linear expression ability of the model. In industrial Internet intrusion detection, network traffic data is usually time series data, and its feature dimension is relatively high, containing various types of information. In order to effectively extract meaningful features from time series data and avoid interference from redundant information, the present invention uses 1D ResCNN for feature extraction and dimensionality reduction.

[0102] In this embodiment, 4 convolutional layers are used. Through the residual connection (as Figure 6 shown), the output of the first convolutional layer is added to the result of the Dropout operation of the third convolutional layer; through the residual connection, the output of the second convolutional layer is added to the result of the Dropout operation of the fourth convolutional layer; reducing the feature loss in the information transmission process and improving the gradient transmission efficiency.

[0103] In 1D convolution, each convolution kernel (filter) slides only in one dimension, thereby extracting local features in time series data. By stacking convolutional layers in multiple layers, the network can gradually learn feature representations from low-level to high-level. The residual connection further solves the problems of gradient disappearance and training difficulties that may occur when the network depth increases. By directly adding the input data to the output of the convolutional layer, it ensures that information can directly flow through the network layer, thereby accelerating the training of the model and increasing the training depth. In time series data processing, this structure can retain the key information in the input data while extracting effective time series features through convolution.

[0104] Furthermore, as Figure 7 shown, the method for processing the local temporal features through the multi-head attention mechanism to obtain the global dependent features of the weighted key features is specifically as follows:

[0105] As Figure 7 shown, using the local temporal feature sequence X RsTake the input and divide the local temporal feature sequence X Rs into h (h = 2 in this embodiment) feature sequences X i , and each feature sequence X i corresponds to a head.

[0106] Furthermore, calculate the attention values of each feature sequence through self-attention:

[0107] For each feature sequence X i , obtain the query matrix Query i , key matrix key i , and value matrix Value i of the feature sequence X i through linear transformation. The calculation formulas are:

[0108] Query i = W q X i

[0109] key i = W k X i

[0110] Value i = W v X i

[0111] Furthermore, calculate the attention score and softmax score according to the query matrix, key matrix, and value matrix. The calculation formulas are:

[0112] Attention Score i = Query.Key

[0113]

[0114] where d k represents the vector dimension of the key matrix, and softmax() represents the normalization function.

[0115] Furthermore, multiply the softmax score by the value matrix to obtain the attention value of each feature sequence X i . The calculation formula is:

[0116] Attention Value i = Value i * Softmax Score i

[0117] Furthermore, the attention values of all heads are concatenated, and a global dependency feature of the final weighted key feature is formed through a linear layer (Concat). The calculation formula of the global dependency feature is:

[0118] X att = concat(Attention Value 1 , Attention Value 2 ,..., Attention Value n )

[0119] By introducing the multi-head attention mechanism, the model can not only better capture long-range temporal dependencies but also enhance the ability to focus on key events. This is crucial for identifying abnormal temporal behaviors in industrial Internet intrusion detection, especially for the recognition of minority class attack behaviors.

[0120] S4: Using the global dependency feature of the weighted key feature as the input, multi-dimensional temporal features containing long-term time dependency features are extracted through a multi-layer bidirectional SRU module. The multi-dimensional temporal features not only include the short-term dependency relationships of local time series but also the long-term dependency information within a relatively long time range. By combining the local temporal features extracted by 1D ResCNN and the global dependency features captured by the multi-head attention mechanism, the model can comprehensively use the multi-dimensional temporal features for intrusion detection classification and identify potential attack patterns.

[0121] In the intrusion detection task of the industrial Internet, network traffic data has significant temporal dependency characteristics. Traditional recurrent neural networks (RNNs) have certain limitations in capturing complex long-sequence dependency relationships. Especially when dealing with long temporal relationships, it is difficult to efficiently extract key information. To solve this problem, the present invention proposes to use a four-layer BISRU module to extract more representative temporal features from the enhanced industrial Internet data to improve the model's ability to capture the temporal information of attack behaviors.

[0122] Furthermore, as Figure 8 shown, the calculation process for extracting multi-dimensional features containing long-term time dependency features is:

[0123] Perform word embedding on the global dependency feature X at of the weighted key feature to obtain a low-dimensional continuous vector X in = {x 1 , x 2 , x t ......, x n}, where x t is the sequence feature at the t-th time step (the time step t here does not have the same meaning as the time step t in the diffusion model), and n is the number of features. The low-dimensional continuous vector Xin Perform forward SRU calculation and backward SRU calculation respectively to obtain the final forward hidden state and the final backward hidden state; merge the final forward hidden state and the final backward hidden state into a comprehensive hidden state.

[0124] Furthermore, the method for performing forward SRU calculation is: initialize the intermediate state C of the previous time step at the first time step to a zero vector. After that, starting from time step t = 1, obtain the forward hidden state by gradually performing SRU calculations. A forward hidden state h is calculated at each time step. t-1 This state is obtained through SRU calculation from the intermediate state C of the previous time step t and the input sequence feature x of the current time step. t-1 t

[0125] Furthermore, the method for performing backward SRU calculation is: initialize the intermediate state C of the next time step at the last time step to a zero vector. After that, starting from time step t = n, obtain the backward hidden state by gradually performing SRU calculations. A backward hidden state t+1 is calculated at each time step. This state is obtained through SRU calculation from the intermediate state C of the next time step and the input sequence feature x of the current time step. t+1 t

[0126] Furthermore, taking the SRU calculation in the forward SRU calculation as an example, the method for performing SRU calculation is:

[0127] In each SRU unit, perform a linear transformation on the input sequence feature x t to calculate the preliminary extraction feature

[0128]

[0129] where W is the linear transformation matrix.

[0130] Furthermore, calculate the forget gate f t , and determine the output of the intermediate state vector C of the current time step according to the adaptive average of the intermediate state vector C of the previous time step t-1 and the forget gate f t . The calculation formula is: t

[0131] f t = σ·(W f x t + b f )

[0132]

[0133] Among them, W f and b f are the weight and bias of the forget gate respectively, σ is the activation function, and the tanh function is adopted here. ⊙ represents element-wise multiplication.

[0134] The forget gate f t depends on the intermediate state vector C t-1 at the previous time step. How much past information to retain depends on the obtained f t . To reduce parameters, (1 - f t ) is directly used to control the current information. By calculating the intermediate state vector C t Using lightweight recurrence in each SRU enhances the ability to process data in parallel.

[0135] Furthermore, calculate the reset gate r t . Input the adjusted intermediate state vector C t at the current time step into the non-linear activation function, and adaptively combine the input sequence feature x t and the output of the non-linear activation function to update the hidden state h t at the current time step. The calculation formula is:

[0136] r t = σ·(W r x t + b r )

[0137] h t = r t ⊙ g(C t )+(1 - r t )⊙ x t

[0138] Among them, W r and b r are the weight and bias of the reset gate respectively. The non-linear activation function g(C t ), and the sigmoid function is adopted here. (1 - r t )⊙ x t is a skip connection, which allows the gradient to be directly propagated to the previous layer. By calculating the hidden state h t Using a highway network in each SRU for gradient training to improve the model performance.

[0139] Furthermore, by means of concatenation, the forward hidden state h t and the backward hidden state Merge, merge the hidden states of the two into a complete comprehensive hidden state h t ′ is used to represent multi-dimensional time series features for subsequent classification output.

[0140] The four-layer BISRU module adopted by the present invention has a bidirectional structure, enabling BISRU to recognize the time dependence relationship of attack behaviors in industrial Internet traffic, thereby more accurately characterizing the time series features of different attack types. Through the time series information extraction ability of the bidirectional structure, the BISRU module effectively enhances the model's recognition and classification performance of attack behaviors. The simplified gating mechanism of BISRU retains key state information, improving the model's processing efficiency and accuracy for long sequence data.

[0141] S5: Perform classification of intrusion detection based on multi-dimensional time series features.

[0142] Finally, input the extracted multi-dimensional time series features h = {h 1 ′, h 2 ′,..., h t ′,..., h n ′} into the classification layer, perform classification through the fully connected layer and the Softmax layer to obtain the probabilities of eight attack types, and output the category with the highest probability as the prediction result.

[0143] Embodiment 2

[0144] An industrial Internet intrusion detection model based on a diffusion model, as Figure 2 shown, includes a data preprocessing module, a data augmentation module based on a multi-layer diffusion model, a feature extraction and dimensionality reduction module based on a ResCNN-Attention hybrid network, a time series modeling module based on BISRU, and an intrusion detection classification module connected in sequence.

[0145] Among them, the data preprocessing module is used to clean and standardize the original industrial Internet traffic data;

[0146] The data augmentation module based on the multi-layer diffusion model is used to take the standardized industrial Internet traffic data as input, generate synthetic samples for minority class samples using the multi-layer diffusion model, and perform inverse standardization, format adjustment processing, and data balancing sampling on the synthetic samples in sequence to obtain balanced sample data;

[0147] The feature extraction and dimensionality reduction module based on the ResCNN-Attention hybrid network is used to extract and reduce the dimensions of the balanced samples through 1D ResCNN to obtain local time series features; process the local time series features through the multi-head attention mechanism to obtain the global dependence features of the weighted key features;

[0148] The time series modeling module based on BISRU is used to take the global dependent features of weighted key features as input and extract multi-dimensional time series features containing long-term time-dependent features through a multi-layer bidirectional SRU module;

[0149] The intrusion detection classification module is used to classify intrusion detection according to the multi-dimensional time series features.

[0150] Specifically, as Figure 3 shown, the data augmentation module based on the multi-layer diffusion model includes a diffusion model input layer, a multi-layer diffusion layer, a data synthesis layer, and a diffusion model output layer connected in sequence. The diffusion model input layer is used to obtain the standardized industrial Internet traffic data X and select a small number of class samples as the samples to be augmented. The multi-layer diffusion layer is used to generate synthetic samples for the small number of class samples using the multi-layer diffusion model, and use the mean square error as the loss function to calculate the loss and update the model parameters. The data synthesis layer is used to perform inverse standardization, format adjustment processing, and data balancing sampling on the synthetic samples in sequence. The diffusion model output layer is used to output balanced sample data.

[0151] Specifically, the diffusion model input layer is connected to the data preprocessing module; each diffusion layer includes an encoder and a decoder connected in sequence. The encoder is used to perform forward diffusion to generate noise samples, and the decoder is used to perform reverse denoising on the noise samples. The input of the encoder of the first diffusion layer is connected to the diffusion model input layer; in the middle multi-layer diffusion layer and the last diffusion layer, the input of the encoder is connected to the output of the decoder of the previous layer, and the output of the decoder of the last diffusion layer is connected to the data synthesis layer. In this embodiment, the number of diffusion layers is 4.

[0152] Specifically, the feature extraction and dimensionality reduction module based on the ResCNN-Attention hybrid network includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, and a max pooling layer connected in sequence. The convolutional layer is used for feature extraction and dimensionality reduction, and the max pooling layer is used for dimensionality reduction processing. The input of the first convolutional layer is connected to the diffusion model output layer. Each convolutional layer includes a 1D convolutional operation, a 20% Dropout operation, and a tanh activation function performed in sequence. The output of the first convolutional layer is connected to the result of the Dropout operation of the third convolutional layer in a residual connection. The output of the second convolutional layer is connected to the result of the Dropout operation of the fourth convolutional layer in a residual connection. The output of the max pooling layer is connected to the multi-head attention module, and the multi-head attention module is used to process the local time series features to obtain the global dependent features of the weighted key features.

[0153] Specifically, as Figure 8As shown in the figure, the BISRU-based time series modeling module includes an embedding layer (Word Embedding), a multi-layer bidirectional SRU module, and an output layer connected in sequence; the input of the embedding layer is connected to the multi-head attention module, and the output of the embedding layer is connected to the multi-layer bidirectional SRU module. The embedding layer is used to perform word embedding on the global dependence features of the weighted key features to obtain low-dimensional continuous vectors; the multi-layer bidirectional SRU module is used to extract multi-dimensional time series features containing long-term time dependence features; the output layer is used to output multi-dimensional time series features;

[0154] Specifically, the multi-layer bidirectional SRU module includes multi-layer bidirectional SRU layers (4 layers in this embodiment) connected in sequence. Each bidirectional SRU layer includes a forward SRU layer and a backward SRU layer for parallel processing. The forward SRU layer is used to perform forward SRU calculations, and the backward SRU layer is used to perform backward SRU calculations; the forward SRU layer includes multiple SRU units connected in sequence in the forward direction, and the backward SRU layer includes multiple SRU units connected in sequence in the reverse direction. The SRU unit is used to perform SRU calculations; the SRU units in the forward SRU layer and the backward SRU layer of the first bidirectional SRU layer are all connected to the output of the embedding layer. The SRU units in the forward SRU layer and the backward SRU layer of the intermediate bidirectional SRU layer are all connected to the corresponding SRU units in the forward SRU layer and the backward SRU layer of the previous bidirectional SRU layer; the SRU units in the forward SRU layer and the backward SRU layer of the last bidirectional SRU layer are all connected to the intrusion detection classification module.

[0155] Specifically, the intrusion detection classification module includes a fully connected layer and a Softmax layer connected in sequence. The fully connected layer is connected to the SRU units in the forward SRU layer and the backward SRU layer of the last bidirectional SRU layer; the fully connected layer maps the output features of BISRU to the classification space of each category to generate scores for each category; then, the Softmax layer converts these scores into probability values between 0 and 1 and ensures that the sum of the probabilities of all categories is 1, thereby realizing the probabilistic discrimination of eight attack types. Finally, the model outputs the category with the highest probability as the prediction result.

[0156] In order to verify the effectiveness and authenticity of the method of the present invention, experimental analysis and experimental comparison of the method of the present invention are carried out in this embodiment.

[0157] Experimental environment: The experiment was conducted on a microcomputer with an Intel Core i7 processor, 32.0 GB of memory, an NVIDIA GeForce GTX745 GPU, and a Windows 10 operating system. Python 3.7 was used for programming, and the Keras 2.3.1 and TensorFlow 1.15.0 frameworks, as well as the PyTorch 1.10.0 framework, were used for model training.

[0158] Experimental dataset: In the present invention, the Gas Pipeline standard industrial dataset was used. This dataset was provided by Mississippi State University in 2014 and was designed specifically for industrial Internet intrusion detection research. It simulates the natural gas pipeline operation scenario in an industrial control system environment and includes normal flow and seven different types of malicious attack flows. All data in the dataset record the communication information, status parameters, and operation behaviors of the industrial control system, and have high representativeness of the industrial environment. Therefore, it has important research value for industrial Internet intrusion detection systems. The dataset contains 27 columns of data, with 26 columns being data features and the last column being the label. The data features cover multiple aspects of flow communication and device operation in the industrial control system, including: command_address, response_address, command_memory, response_memory, command_memory_count, response_memory_count, etc. The label column is used to identify the data category, with normal data being 0 and the labels for the seven attack behaviors being 1 - 7 respectively. However, the Gas Pipeline dataset has a high degree of imbalance, with significant differences in the data volume of different attack categories. The specific distribution of the dataset flow categories is shown in Table 1. During the experiment, 20% of it was randomly selected as the test set for testing, and the rest was the training set.

[0159] 1. Table 1. Dataset description

[0160]

[0161]

[0162] Compared with traditional intrusion detection datasets (such as the NSL-KDD dataset), the Gas Pipeline dataset can better reflect the complex and diverse attack types and high-dimensional features in the industrial Internet. It not only contains common malicious traffic types but also includes the state and behavior characteristics of simulated industrial devices. Using this dataset, the recognition performance of the intrusion detection model under different attack types can be verified, and at the same time, it can effectively meet the actual needs of high-dimensional data and large traffic in the industrial Internet scenario.

[0163] Data preprocessing: Before the input layer, the original industrial Internet of Things traffic data was cleaned, normalized, and standardized to ensure the data was clean and consistent. Through the low-variance filtering method, 17 important features were selected to form a unified data input into the input layer.

[0164] Feature selection:

[0165] The Gas Pipeline dataset contains multiple feature sequences. To improve the efficiency of the detection system, the present invention uses the low-variance feature filtering method for feature selection. First, the variance of each feature is calculated to identify features with small variances and limited information content. On this basis, the 9 features with the smallest variances are selected and removed to ensure that the remaining features have larger variances, enabling significant differentiation of different samples. Finally, a 17-dimensional effective feature dataset is obtained, and these features contribute significantly to sample differentiation.

[0166] Data standardization:

[0167] The Gas Pipeline dataset has the characteristic of high-dimensional features, and the range of feature values has a large span. This scale difference will interfere with model training. To reduce the impact of numerical differences between different features on model training, the present invention uses the min-max normalization method to map all feature values to the range [0,1]. Normalization processing can improve the convergence speed of the model and make different features receive equal attention during training. The min-max normalization formula is as follows:

[0168] One-Hot Encoding

[0169] The Gas Pipeline dataset contains unordered discrete features. Directly inputting them into the classifier will reduce the classification effect. To solve this problem, the present invention uses the One-Hot Encoding method to convert the labels of each category into binary feature columns. For example, the dataset of the present invention contains eight classification labels (Normal, NMRI, CMRI, MSCI, MPCI, MFCI, DOS, and Recon). After One-Hot Encoding, the labels will be represented in the form of 8-dimensional binary vectors. For example, the Normal category is represented as (1,0,0,0,0,0,0,0), NMRI is represented as (0,1,0,0,0,0,0,0), and so on. Through One-Hot Encoding, the discrete labels are converted into numerical data that is more easily processed by machine learning models, improving the prediction performance of the classifier.

[0170] Model evaluation metrics:

[0171] To comprehensively evaluate the performance of the model, the following evaluation metrics are used in the experiments of the present invention:

[0172] Accuracy: Used to measure the overall classification accuracy of the model for all attack categories. The calculation formula is:

[0173]

[0174] Precision: Used to evaluate the accuracy of the model when predicting a specific attack category, that is, the proportion of samples actually belonging to a certain category among the samples predicted by the model as a certain category.

[0175]

[0176] Recall: Used to measure the recognition ability of the model for samples of a specific category, indicating the proportion of samples actually belonging to a certain category that are correctly recognized.

[0177]

[0178] F1 Score: The F1 score is the harmonic mean of precision and recall, used to balance precision and recall in the case of data imbalance. The formula is:

[0179]

[0180] In the formula: TP is the number of normal data correctly classified by the model; TN is the number of abnormal data correctly classified by the model; FP is the number of data in the abnormal class misclassified as normal data by the model; FN is the number of data in the normal class misclassified as abnormal data by the model.

[0181] Analysis of experimental results:

[0182] In this invention, by testing the Gas Pipeline dataset, the effectiveness of the multi-layer diffusion generation model in data augmentation is verified. For the two minority classes MSCI and MFCI in the dataset, corresponding synthetic samples are generated to balance the number of samples of each category in the training set. The specific number of generated samples is shown in Table 2. To better evaluate the performance of the model in data augmentation, 14,400 samples are selected from the Gas Pipeline dataset in the experiment, with 1,800 samples for each category, and they are divided into a training set and a test set according to a ratio of 8:2.

[0183] Table 2 Number of generated samples in the training set

[0184]

[0185] Analysis of model performance:

[0186] To verify the superiority of the method proposed in this invention, the experiment compares the performance of multiple intrusion detection methods including traditional methods, deep learning, and fusion algorithms. The experimental results are shown in Table 3.

[0187] Table 3 Performance Comparison of Different Methods

[0188]

[0189] It can be seen from the experimental data that although the deep learning-based models (such as CNN-GRU and BLSTM RNN) show certain advantages in capturing temporal features, their performance significantly degrades in the data environment with class imbalance, especially in the key metrics for measuring the overall performance of the model such as recall rate and F1 value. In addition, the fusion algorithms (such as CNN+BiSRU and 1D CWGAN-BiSRU) have improved in overall performance by optimizing the network structure and introducing data augmentation techniques, but the highest recall rate is only 95.9%. Although it has improved compared with the traditional methods, there is still a gap compared with 96.87% of the method of the present invention, indicating the possibility of missed detection when capturing unknown attack samples, especially under the challenges of class imbalance and attack diversity, showing the limitations of generalization ability.

[0190] The method proposed by the present invention has achieved significant improvements in various performance metrics by combining feature dimensionality reduction, data augmentation, and temporal dependence modeling techniques. Specifically, this method has reached 99.42%, 98.74%, 96.87%, and 97.76% in terms of accuracy, precision, recall rate, and F1 value respectively, all significantly superior to the comparative methods. This shows that the method of the present invention can not only accurately capture known attack samples, but also show strong detection ability and robustness in the identification of minority classes and unknown attacks, thus effectively coping with the multiple challenges of data class imbalance, strong traffic temporality, and unknown attack detection in industrial Internet intrusion detection, providing an efficient technical solution for improving the security guarantee level of the industrial Internet.

[0191] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An industrial Internet intrusion detection method based on a diffusion model, characterized in that: Includes steps: S1: Clean and standardize the original industrial Internet traffic data; S2: Using the standardized industrial Internet traffic data as input, a multi-layer diffusion model is used to generate synthetic samples for minority samples, and the synthetic samples are de-standardized, format adjusted, and data balanced sampled in turn to obtain balanced sample data; S3: Perform feature extraction and dimensionality reduction on balanced samples through 1D ResCNN to obtain local time series features; process local time series features through multi-head attention mechanism to obtain global dependency features of weighted key features; S4: Taking the global dependency features of weighted key features as input, multi-dimensional temporal features containing long-term temporal dependency features are extracted through a multi-layer bidirectional SRU module; S5: Classification of intrusion detection based on multi-dimensional temporal features.

2. The industrial Internet intrusion detection method based on the diffusion model according to claim 1 is characterized in that: The method for generating synthetic samples using a multi-layer diffusion model for minority class samples is as follows: the initial input data is the standardized industrial Internet traffic data X, the minority class samples are selected as the samples to be enhanced, and in each diffusion layer of the multi-layer diffusion generation model, the output of the previous diffusion layer is used to synthesize the sample data is the input; forward diffusion is performed through the encoder to generate noise samples The decoder processes the noise samples Perform inverse denoising to generate synthetic sample data 3. The industrial Internet intrusion detection method based on the diffusion model according to claim 2 is characterized in that: The forward diffusion process is as follows: the output of the previous diffusion layer synthesizes the sample data As input, noise is gradually added until time step t = T ends, and the final noise sample is The calculation method for each time step in the forward diffusion process is: Where L represents the number of diffusion layers, is the noise data of the lth diffusion layer at the tth time step, is a noise term specific to each diffusion layer generated from a Gaussian distribution; The reverse denoising process is as follows: Starting from, gradually remove the noise and generate new data consistent with the real data distribution The calculation method for each time step of the inverse denoising process is: in, is the cumulative noise attenuation, t is the current time step, is the noise intensity at the sth time step, represents the noise attenuation factor, is the Gaussian noise term.

4. The industrial Internet intrusion detection method based on a diffusion model according to any one of claims 1 to 3 is characterized in that: When training the multi-layer diffusion model, the mean square error is used as the loss function to calculate the loss and update the model parameters; the calculation process is: Calculate the loss function of the lth diffusion layer: Among them, x0 is a minority class sample, E t express expectations; The total loss is obtained by weighted summing the losses of all layers: Among them, λ l is the weight coefficient of the loss of the lth layer; The parameters of all diffusion layers are updated simultaneously through back-propagation by minimizing the total loss L total , obtain the optimized multi-layer generation model, and use the optimized multi-layer generation model to obtain the final synthetic sample data 5. The industrial Internet intrusion detection method based on the diffusion model according to claim 4 is characterized in that: The method of extracting features and reducing dimensions of balanced samples through 1D ResCNN to obtain local time series features is: To balance the sample data X p As input, multiple convolutional layers are used to extract features layer by layer. Each convolutional layer includes 1D convolution operation, Dropout operation and tanh activation function in sequence. The output of some convolutional layers is connected with the Dropout operation of another convolutional layer through residual connection. The output of the last convolutional layer is output as the local temporal feature sequence X after the maximum pooling operation. Rs ; The method of processing the local time series features through the multi-head attention mechanism to obtain the global dependency features of the weighted key features is: taking the local time series feature sequence X Rs As input, the local time series feature sequence X Rs Divide into h feature sequences X i , each feature sequence X i Corresponding to a head; obtain the feature sequence X through linear transformation i The query matrix Query i , key matrix key i Sum value matrix Value i : Calculate the attention score and normalized score based on the query matrix, key matrix and value matrix; multiply the normalized score with the value matrix to obtain each feature sequence X i The attention value of all heads is concatenated and the final weighted key feature global dependency feature X is formed through a linear layer. att .

6. The industrial Internet intrusion detection method based on a diffusion model according to any one of claims 1-3 and 5, characterized in that: The method for extracting multidimensional features including long-term time-dependent features through a multi-layer bidirectional SRU module is as follows: the global dependency feature X of the weighted key feature at Perform word embedding to obtain a low-dimensional continuous vector X in ={x1,x2,x t ......,x n }, x t is the sequence feature of the t-th time step, n is the number of features, and the low-dimensional continuous vector X in Perform forward SRU calculation and reverse SRU calculation respectively to obtain forward hidden state and reverse hidden state; merge the forward hidden state and reverse hidden state into a comprehensive hidden state for representing multi-dimensional temporal features.

7. The industrial Internet intrusion detection method based on diffusion model according to claim 6 is characterized in that: The forward SRU calculation method is as follows: the previous hidden state of the first time step is initialized as a zero vector, and starting from time step t=1, the forward hidden state is obtained by performing SRU calculation step by step; The reverse SRU calculation method is as follows: the next hidden state of the last time step is initialized to a zero vector, Starting from time step t=n, obtain the reverse hidden state by performing SRU calculation step by step; The method for performing SRU calculation is: In each SRU unit, the input sequence feature x t Perform linear transformation and calculate the preliminary extracted features; Where W is the linear transformation matrix; Calculate the forget gate f t , according to the state vector C of the previous time step t-1 and forget gate f t The adaptive average value determines the output state vector C of the current time step t , the state vector C of the current time step t The calculation formula is: f t =σ·(W f x t +b f ) Among them, W f and b f are the weight and bias of the forget gate, σ is the activation function, and ⊙ represents element-by-element multiplication; Calculate the reset gate r t , the state vector C of the current time step t Input into the nonlinear activation function, adaptively combining the input sequence features x according to the forget gate t And the output of the nonlinear activation function, update the current hidden state h t , update the current hidden state h t The calculation formula is: r t =σ·(W r x t +b r ) h t =r t ⊙g(C t )+(1-r t )⊙x t Among them, W r and b r are the reset gate weights and biases respectively, and g() is the nonlinear activation function.

8. An industrial Internet intrusion detection model based on a diffusion model, used to implement the industrial Internet intrusion detection method based on a diffusion model according to any one of claims 1 to 7, characterized in that: It includes a sequentially connected data preprocessing module, a data enhancement module based on a multi-layer diffusion model, a feature extraction and dimension reduction module based on a ResCNN-Attention hybrid network, a BISRU-based time series modeling module, and an intrusion detection classification module; Among them, the data preprocessing module is used to clean and standardize the original industrial Internet traffic data; The data enhancement module based on the multi-layer diffusion model is used to take the standardized industrial Internet traffic data as input, generate synthetic samples for minority samples using the multi-layer diffusion model, and perform de-standardization and format adjustment processing and data balance sampling on the synthetic samples in turn to obtain balanced sample data; The feature extraction and dimensionality reduction module based on the ResCNN-Attention hybrid network is used to extract features and reduce dimensions of balanced samples through 1D ResCNN to obtain local time series features; the local time series features are processed through the multi-head attention mechanism to obtain the global dependency features of weighted key features; The BISRU-based temporal modeling module is used to extract multidimensional temporal features including long-term temporal dependency features through a multi-layer bidirectional SRU module, taking the global dependency features of weighted key features as input; The intrusion detection classification module is used to classify intrusion detection based on multi-dimensional time series features.

9. The industrial Internet intrusion detection model based on the diffusion model according to claim 8 is characterized in that: The data enhancement module based on the multi-layer diffusion model comprises a diffusion model input layer, a multi-layer diffusion layer, a data synthesis layer and a diffusion model output layer connected in sequence; the diffusion model input layer is connected to the data preprocessing module; each diffusion layer comprises an encoder and a decoder connected in sequence, and the input of the encoder of the first diffusion layer is connected to the diffusion model input layer; in the middle diffusion layer and the last diffusion layer, the input of the encoder is connected to the output of the previous decoder, and the decoder output of the last diffusion layer is connected to the data synthesis layer; The feature extraction and dimensionality reduction module based on the ResCNN-Attention hybrid network includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer and a maximum pooling layer connected in sequence. The first convolutional layer is connected to the diffusion model output layer, and each convolutional layer includes a 1D convolution operation, a Dropout operation and a tanh activation function performed in sequence; the output of the first convolutional layer is residually connected to the Dropout operation result of the third convolutional layer, and the output of the second convolutional layer is residually connected to the Dropout operation result of the fourth convolutional layer; the output of the maximum pooling layer is connected to the multi-head attention module.

10. The industrial Internet intrusion detection model based on the diffusion model according to claim 9 is characterized in that: The BISRU-based temporal modeling module includes an embedding layer, a multi-layer bidirectional SRU structure and an output layer connected in sequence; the input of the embedding layer is connected to the multi-head attention module, and the output of the embedding layer is connected to the multi-layer bidirectional SRU module; the multi-layer bidirectional SRU module includes multi-layer bidirectional SRU layers connected in sequence, each bidirectional SRU layer includes a forward SRU layer and a reverse SRU layer processed in parallel, the forward SRU layer includes a plurality of SRU units connected in sequence in the forward direction, and the reverse SRU layer includes a plurality of SRU units connected in sequence in the reverse direction; the SRU units in the forward SRU layer and the reverse SRU layer of the first bidirectional SRU layer are both connected to the output of the embedding layer, the SRU units in the forward SRU layer and the reverse SRU layer of the middle bidirectional SRU layer are both connected to the corresponding SRU units in the forward SRU layer and the reverse SRU layer of the previous bidirectional SRU layer; the SRU units in the forward SRU layer and the reverse SRU layer of the last bidirectional SRU layer are both connected to the intrusion detection classification module.

Citation Information

Cited By

  • Network traffic anomaly detection method and device based on diffusion model, and medium

    CN120896801A

  • A network traffic anomaly detection method, device and medium based on a diffusion model

    CN120896801B