Background flow generation method based on diffusion model and bidirectional flow characteristics
Through the background traffic generation method based on diffusion model and bidirectional flow characteristics, the lack of traffic generation in the existing technology in complex network environments is solved, and high-fidelity and diverse background traffic generation is achieved, and network security simulation and intrusion detection training is supported.
Patent Information
- Application Number
- CN202510921208.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-09-05
AI Technical Summary
Existing traffic generation technologies are difficult to meet the multi-dimensional needs of traffic authenticity, structural integrity and controllability in complex network environments, especially in mixed background scenarios where multiple categories coexist.
The background flow generation method based on diffusion model and bidirectional flow characteristics is adopted, including the spatiotemporal feature extraction module, the conditional control module and the diffusion generation module. Through the spatiotemporal feature extraction, conditional control and diffusion generation process, high-fidelity and diverse background flow are generated.
It realizes high-quality generation of multi-category hybrid network background traffic, simulates timing characteristics and bidirectional alternating structures in the real network, supports network security simulation, intrusion detection training and complex traffic modeling, and provides strong data support.
Smart Images

Figure CN120602198A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and in particular to a background traffic generation method based on a diffusion model and bidirectional flow characteristics. Background Art
[0002] With the rapid evolution of information infrastructure and the continuous expansion of network scale, the network environment is becoming increasingly complex, and bandwidth pressure, business diversity and security threats have increased significantly. Background traffic, as the basic data flow generated by conventional communication protocols and legitimate applications in the network, does not directly participate in attack behavior in most cases, but plays an important role in network performance evaluation, intrusion detection system (IDS) testing and simulation environment construction. There is a large amount of background traffic in business scenarios such as web browsing, file transfer, video playback, and voice communication. It exhibits characteristics such as diversity, asymmetry and time-varying in the network. Accurately generating representative background traffic is of great practical significance for improving the detection and generalization capabilities of network defense systems. However, existing traffic generation technologies still cannot meet the multi-dimensional requirements for traffic authenticity, structural integrity and controllability in complex network environments.
[0003] Traditional background traffic generation methods primarily rely on statistical modeling or simulation replay techniques. The former typically assumes that traffic conforms to a certain probability distribution (such as Poisson or normal distribution) and lacks the ability to model the dynamic changes and heterogeneous characteristics of real networks. The latter, while achieving high fidelity, achieves traffic reproduction by replaying historical traffic. However, due to limitations in the representativeness of the collected data, the temporal granularity, and the generalization of scenarios, it is difficult to support flexible configuration and customized generation. Furthermore, deep learning methods such as generative adversarial networks (GANs) have been introduced to the field of traffic generation in recent years, but they suffer from problems such as unstable training, poor controllability of generated samples, and difficulty capturing the characteristics of bidirectional alternating flows. They perform particularly poorly when dealing with mixed background scenarios with multiple classes. Summary of the Invention
[0004] (1) Technical problems solved
[0005] In view of the deficiencies of the prior art, the present invention provides a background flow generation method based on a diffusion model and bidirectional flow characteristics.
[0006] (2) Technical solution
[0007] To achieve the above-mentioned object, the present invention provides the following technical solutions: The background traffic generation method based on the diffusion model and bidirectional flow characteristics of the present invention includes a spatiotemporal feature extraction module, a condition control module and a diffusion generation module;
[0008] The spatiotemporal feature extraction module is used to pre-process the original network traffic, extract the packet length, arrival time and flow direction information of the data packet, and convert it into a spatiotemporal image representation with bidirectional features;
[0009] The condition control module is used to set the generation ratio according to the generation requirements of different types of traffic, and convert the category ratio vector into a control vector through embedding;
[0010] The diffusion generation module receives the spatiotemporal image and the category control vector as input, gradually reconstructs the traffic image through a condition-guided diffusion process, and then maps it back to the background traffic data, outputting high-fidelity background traffic that is close to the actual network structure.
[0011] Preferably, the spatiotemporal feature extraction module includes:
[0012] Data cleaning: filtering invalid or damaged data packets;
[0013] Bidirectional flow coding: The flow direction information is added to the packet length to construct a two-dimensional histogram of packet length and time. The flow direction is distinguished by the color channel to generate a spatiotemporal feature image.
[0014] Further preferably, the condition control module includes:
[0015] Conditional vector generation: Generate a vector condition with the same dimension as the number of categories based on the preset category ratio vector ;
[0016] Embedding mapping: by embedding function f embed Convert the conditional vector to a high-dimensional embedding vector condition emb , with spatiotemporal image stitching as input to the diffusion model.
[0017] Again preferably, the conditional vector generation formula of the conditional control module is:
[0018]
[0019] Among them, ratio i Represents the generation ratio of each traffic type, condition vector The final result is a vector containing multiple elements, which is used to adjust the proportion of each type of sample in the generation process. The embedding process is implemented by the following formula:
[0020] condition emb =f embed (condition vector )
[0021] Among them, f embed Is an embedding function that converts the conditional vector into an embedding vector condition with the specified dimension emb , in the generation phase, the generation model embeds the condition embAs an additional input, we control the type of traffic generated and its proportion. In each generation step, the model uses the conditional vector to adjust the generated samples to ensure that the generated samples meet the expected traffic proportion.
[0022] x t =f gen (x t-1 ,condition emb )
[0023] Among them, x t is the sample generated at time step t, f gen is the generating function, condition emb Affects the sample generation process.
[0024] Preferably, the specific process of the diffusion generation module includes: inputting the spatiotemporal feature image into the Transformer encoder, extracting multi-scale features, combining the conditional embedding vector, enhancing the temporal dependency through the cross-modal attention mechanism, and gradually generating the denoised flow image in the reverse denoising process, and restoring it to the flow data through inverse mapping;
[0025] The diffusion generation module includes two stages: a forward diffusion process and a backward diffusion process. In the forward diffusion process, the original data is gradually transformed into pure noise by adding noise multiple times, while in the backward diffusion process, the model gradually restores the structure of the data based on the noise, and finally generates samples similar to the training data.
[0026] Further preferably, the forward diffusion process of the diffusion generation module asymptotically transforms the input real image x0 into a pure Gaussian noise image x by adding T times noise to the initial image. T ,In each step of noise addition, x t-1 Add a Gaussian noise to generate a new latent variable x t ,The single noise adding process is expressed as:
[0027]
[0028] in: represents the mean of the Gaussian distribution, β t Ι represents the variance of Gaussian distribution, β t is a hyperparameter that gradually increases with t, and Ι represents the unit matrix with the same dimension as the input sample x0.
[0029] Again preferably, the reverse denoising process of the diffusion generation module is expressed as:
[0030]
[0031] Among them, α t , βt is the diffusion scheduling parameter, ∈ θ is the noise prediction network, z is the standard Gaussian noise, condition emb is the conditional embedding vector.
[0032] Preferably, in the residual module of the diffusion generation module, the conditional information is divided into two parts, which are processed by the convolution layer respectively to generate different control signals, which are finally fused with the input signal. The formula is:
[0033] x t =σ(g)·tanh(f)
[0034] Among them, g and f are the gate signal and filter signal after conditional convolution processing, σ is the Sigmoid activation function, and tanh is the hyperbolic tangent activation function.
[0035] Further preferably, the diffusion generation module introduces diffusion step embedding technology, and maps the step information to a high-dimensional space through a fully connected layer and an activation function. The formula is as follows:
[0036] e t =Projection2(Silu(Projection1(e′ t )))
[0037] where e' t It is the original embedding generated based on the step size, Projection1 and Projection2 are the fully connected layers, and Silu is the activation function.
[0038] Again preferably, the diffusion generation module adopts a cross-modal attention mechanism, which enables the model to effectively combine input data and conditional information, encode and decode input features through multiple self-attention layers, and capture temporal dependencies using the Transformer structure. The calculation formula of the cross-modal attention mechanism is:
[0039]
[0040] Among them, Q, K, and V are generated by linear mapping of spatiotemporal features and conditional embedding vectors respectively, and d is the feature dimension;
[0041] The spatiotemporal feature image is in Flowpic format, where the packet length is represented by color depth and the flow direction is distinguished by color channels.
[0042] (3) Beneficial effects
[0043] Compared with the existing technology, the present invention provides a background traffic generation method based on the diffusion model and bidirectional flow characteristics, which has the following beneficial effects:
[0044] This technical solution enables high-quality generation of multi-category mixed network background traffic, fully simulating the temporal characteristics and bidirectional alternating structure of real networks. It utilizes a spatiotemporal feature extraction module to extract key features from traffic sessions and encodes them as image information for input into a diffusion model. Through a step-by-step denoising approach, it generates highly reproducible and structurally consistent traffic samples. Furthermore, combined with a conditional control mechanism, it can adjust the proportions of different traffic types and guide their generation, providing strong data support for network security simulation, intrusion detection training, and complex traffic modeling, with broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 This is a flow chart of the background flow generation method based on the diffusion model and bidirectional flow characteristics of the present invention;
[0046] Figure 2 This is a diagram of the solution architecture of the present invention;
[0047] Figure 3 Generate a sample frame diagram for the diffusion model of the present invention. DETAILED DESCRIPTION
[0048] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0049] See also Figure 1-3 ,The background traffic generation method based on the diffusion model and bidirectional flow characteristics of the present invention includes a spatiotemporal feature extraction module, a condition control module and a diffusion generation module;
[0050] The spatiotemporal feature extraction module is used to pre-process the original network traffic, extract the packet length, arrival time and flow direction information of the data packet, and convert it into a spatiotemporal image representation with bidirectional features;
[0051] The condition control module is used to set the generation ratio according to the generation requirements of different types of traffic, and convert the category ratio vector into a control vector through embedding;
[0052] The diffusion generation module receives the spatiotemporal image and the category control vector as input, gradually reconstructs the traffic image through a condition-guided diffusion process, and then maps it back to the background traffic data, outputting high-fidelity background traffic that is close to the actual network structure.
[0053] This technical solution aims to address the shortcomings of existing background traffic generation technologies in terms of authenticity modeling, temporal feature restoration, and multi-category traffic ratio control. By introducing an improved spatiotemporal feature extraction module, a conditional control module, and a diffusion generation module, this method achieves the automatic generation of high-fidelity and diverse background traffic in complex network environments.
[0054] Step 1. The spatiotemporal feature extraction module uses the Flowpic format to add flow direction to image features, thereby capturing bidirectional flow features and integrating three types of information: packet length, packet time, and packet direction. In a session, traffic typically takes the form of alternating bidirectional flows. The improved method further introduces packet direction information. The server-to-client direction is defined as positive, while the client-to-server direction is defined as negative. The positive and negative signs containing the direction information are superimposed on the packet length, ultimately doubling the amount of information in the input matrix. The improved Flowpic adds more traffic information to the neural network input, including the direction, packet length, and arrival time of bidirectional flows. The direction and packet length distribution are expressed through color and color depth.
[0055] Step 2. Conditional Control Module. The conditional control module introduces probability weighting to set specific generation ratios for different traffic types. This control mechanism ensures that the generated traffic samples align with the traffic distribution in real-world network environments, better meeting the diverse needs of traffic generation and providing more representative traffic samples when simulating diverse real-world application scenarios. In the proposed model, the main purpose of conditional control is to control the generation ratio of each traffic type through conditional vectors when generating different types of traffic samples.
[0056] Step 2-1. Generate a conditional vector. Generate a conditional vector based on the given traffic type ratios. This conditional vector represents the control ratio of each traffic type when generating samples. There are multiple traffic types, and flow_ratios specifies the ratio of each traffic type. The conditional vector is generated based on these ratios.
[0057]
[0058] Among them, ratio i Represents the generation ratio of each traffic type, condition vector The final result is a vector containing multiple elements, which is used to adjust the proportion of each type of sample in the generation process.
[0059] Step 2-2. Conditional embedding: The generated conditional vector will be converted into an embedded representation in a high-dimensional space through the embedding layer, which is convenient for combining with other features and affecting the generation process.
[0060] condition emb =f embed(condition vector )
[0061] Among them, f embed Is an embedding function that converts the conditional vector into an embedding vector condition with the specified dimension emb .
[0062] Step 2-3. Application in the generation process. In the generation phase, the generation model embeds the condition into the condition emb As an additional input, we control the type of traffic generated and its proportion. In each generation step, the model uses the conditional vector to adjust the generated samples to ensure that the generated samples meet the expected traffic proportion.
[0063] x t =f gen (x t-1 ,condition emb )
[0064] Among them, x t is the sample generated at time step t, f gen is the generating function, condition emb Affects the sample generation process.
[0065] Step 3. Diffusion generation module. Figure 3 As shown in the figure, based on the improved spatiotemporal feature extraction module, the bidirectional flow features of network traffic are extracted as image features; then, the diffusion model is used to learn the potential distribution of these features, where the image features are used as the conditional input of the Transformer encoder, and a probability control term is added to the conditional control part to control the generation ratio; the features encoded by the Transformer are further passed to the attention module, and the attention module focuses on specific important parts of the original noisy traffic. The residual block is used to optimize the feature extraction process, alleviate the gradient vanishing problem in deep neural networks, and accelerate the convergence of the network.
[0066] Step 3-1. Use the bidirectional flow image features as conditional input. The convolutional layer projects the original input to the target number of channels, performing a preliminary feature mapping on the input data. To adapt the model to different diffusion step sizes, the model introduces a diffusion embedding technique. This module encodes the diffusion step size and maps the features at different step sizes into a high-dimensional space. The formula is as follows:
[0067] e t =Projection2(Silu(Projection1((e' t )))
[0068] where e' tIt is the original embedding generated based on the step size, Projection1 and Projection2 are the fully connected layers, and Silu is the activation function.
[0069] Step 3-2. The input is fed into a series of residual modules for further processing. At each layer, the data is weighted using a conditional control mechanism. Within each residual module, the conditional information is split into two parts, each processed through a convolutional layer to generate different control signals, which are then fused with the input signal.
[0070] x t =σ(g)·tanh(f)
[0071] Among them, g and f are the gate signal and filter signal after conditional convolution processing, σ is the Sigmoid activation function, and tanh is the hyperbolic tangent activation function.
[0072] Step 3-3. The cross-modal attention mechanism enables the model to effectively combine input data and conditional information. Input features are encoded and decoded through multiple self-attention layers, and temporal dependencies are captured using the Transformer architecture. Furthermore, to further enhance the model's learning of time steps, temporal features are incorporated into the residuals of each layer.
[0073] Detailed workflow
[0074] Spatiotemporal feature extraction
[0075] Data cleaning: Ultrasonic cleaning of raw network traffic to remove invalid or damaged data packets.
[0076] Bidirectional flow coding: Converts the packet length, arrival time, and flow direction information of a data packet into a spatiotemporal image representation with bidirectional features. The specific steps include:
[0077] Split bidirectional streams by time windows;
[0078] Map the direction information into symbols and superimpose them on the packet length;
[0079] Construct a packet length-time two-dimensional histogram and distinguish the spatiotemporal distribution of forward and reverse traffic by color depth.
[0080] Conditional Control
[0081] Conditional vector generation: Generates conditional vectors based on preset traffic class ratios.
[0082] Embedding mapping: The conditional vector is converted into a high-dimensional embedding vector through the embedding function and concatenated with the spatiotemporal image as the input condition of the diffusion model.
[0083] Diffusion Generation
[0084] Transformer encoder: The spatiotemporal feature image is input into the Transformer encoder to extract multi-scale features, and combined with the conditional embedding vector, the temporal dependency is enhanced through the cross-modal attention mechanism.
[0085] The cross-modal attention mechanism enables the model to effectively combine input data and conditional information. Input features are encoded and decoded through multiple self-attention layers, and temporal dependencies are captured using the Transformer structure. Furthermore, to further enhance the model's learning of time steps, temporal features are introduced into the residuals of each layer. The calculation formula for the cross-modal attention mechanism is:
[0086]
[0087] Among them, Q, K, and V are generated by linear mapping of spatiotemporal features and conditional embedding vectors respectively, and d is the feature dimension
[0088] Reverse denoising: gradually generate denoised traffic images and restore them to traffic data through inverse mapping. The specific steps include:
[0089] Forward diffusion process: gradually transforming the initial image into pure noise by adding noise multiple times;
[0090] The forward diffusion process of the diffusion generation module adds T times of noise to the initial image, and asymptotically transforms the input real image x0 into a pure Gaussian noise image x T ,In each step of noise addition, x t-1 Add a Gaussian noise to generate a new latent variable x t ,The single noise adding process is expressed as:
[0091]
[0092] in: represents the mean of the Gaussian distribution, β t Ι represents the variance of Gaussian distribution, β t is a hyperparameter that increases with t, and Ι represents the unit matrix with the same dimension as the input sample x0
[0093] Backward diffusion process: The model gradually restores the structure of the data based on the noise, and eventually generates samples similar to the training data.
[0094] The reverse denoising process of the diffusion generation module is expressed as:
[0095]
[0096] Among them, α t , β tis the diffusion scheduling parameter, ∈ θ is the noise prediction network, z is the standard Gaussian noise, condition emb is the conditional embedding vector.
[0097] Model training and evaluation
[0098] JS divergence is used to characterize the distance between the generated traffic data of different generation models and the real traffic data, so as to measure the authenticity of the data generated by different models.
[0099] The experimental dataset verifies that the improved model is significantly superior to the traditional GAN method in terms of generation quality and error control, proving its effectiveness and advancement in traffic generation tasks.
[0100] JS divergence is a method to measure the distance between different random distributions. The formula is defined as follows:
[0101]
[0102] Where P represents the distribution of real data; Q represents the distribution of generated data; M = (P + Q) / 2; KL divergence, also known as relative entropy or information divergence, can be used to measure the difference between two probability distributions. Given two probability distributions P and Q, the KL divergence between them is defined as:
[0103]
[0104] Among them, p(x) and q(x) are the probability density functions of P and Q respectively. Expanding KL(P||Q) yields:
[0105]
[0106] Where H(P) is the entropy and H(P,Q) is the cross entropy of P and Q. In information theory, the entropy H(P) represents the minimum number of bytes required to encode a random variable from P, and the cross entropy H(P,Q) represents the number of bytes required to encode a variable from P using a code based on Q.
[0107] Result Output
[0108] Ultimately, multi-category traffic samples that are close to the actual network environment distribution are generated, achieving high-fidelity and high-diversity network background traffic generation, and providing high-quality data support for application scenarios such as network security simulation, intrusion detection training, and complex traffic simulation.
[0109] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. Background traffic generation method based on diffusion model and bidirectional flow characteristics, characterized by: It includes spatiotemporal feature extraction module, condition control module and diffusion generation module; The spatiotemporal feature extraction module is used to pre-process the original network traffic, extract the packet length, arrival time and flow direction information of the data packet, and convert it into a spatiotemporal image representation with bidirectional features; The condition control module is used to set the generation ratio according to the generation requirements of different types of traffic, and convert the category ratio vector into a control vector through embedding; The diffusion generation module receives the spatiotemporal image and the category control vector as input, gradually reconstructs the traffic image through a condition-guided diffusion process, and then maps it back to the background traffic data, outputting high-fidelity background traffic that is close to the actual network structure.
2. The background traffic generation method based on the diffusion model and bidirectional flow characteristics according to claim 1 is characterized in that: The spatiotemporal feature extraction module includes: Data cleaning: filtering invalid or damaged data packets; Bidirectional flow coding: The flow direction information is added to the packet length to construct a two-dimensional histogram of packet length and time. The flow direction is distinguished by the color channel to generate a spatiotemporal feature image.
3. The background traffic generation method based on the diffusion model and bidirectional flow characteristics according to claim 2 is characterized in that: The condition control module includes: Conditional vector generation: Generate a vector condition with the same dimension as the number of categories based on the preset category ratio vector ; Embedding mapping: by embedding function f embed Convert the conditional vector to a high-dimensional embedding vector condition emb , with spatiotemporal image stitching as input to the diffusion model.
4. The background traffic generation method based on the diffusion model and bidirectional flow characteristics according to claim 3 is characterized in that: The conditional vector generation formula of the conditional control module is: Among them, ratio i Represents the generation ratio of each traffic type, condition vector The final result is a vector containing multiple elements, which is used to adjust the proportion of each type of sample in the generation process. The embedding process is implemented by the following formula: condition emb =f embed (condition vector ) Among them, f embed Is an embedding function that converts the conditional vector into an embedding vector condition with the specified dimension emb , in the generation phase, the generation model embeds the condition emb As an additional input, we control the type of traffic generated and its proportion. In each generation step, the model uses the conditional vector to adjust the generated samples to ensure that the generated samples meet the expected traffic proportion. x t =f gen (x t-1 ,condition emb ) Among them, x t is the sample generated at time step t, f gen is the generating function, condition emb Affects the sample generation process.
5. The background traffic generation method based on the diffusion model and bidirectional flow characteristics according to claim 4 is characterized in that: The specific process of the diffusion generation module includes: inputting the spatiotemporal feature image into the Transformer encoder, extracting multi-scale features, combining them with the conditional embedding vector, enhancing the temporal dependency through the cross-modal attention mechanism, and gradually generating the denoised traffic image during the reverse denoising process, and restoring it to traffic data through inverse mapping; The diffusion generation module includes two stages: a forward diffusion process and a backward diffusion process. In the forward diffusion process, the original data is gradually transformed into pure noise by adding noise multiple times, while in the backward diffusion process, the model gradually restores the structure of the data based on the noise, and finally generates samples similar to the training data.
6. The background traffic generation method based on the diffusion model and bidirectional flow characteristics according to claim 5 is characterized in that: The forward diffusion process of the diffusion generation module asymptotically transforms the input real image x0 into a pure Gaussian noise image x by adding T times noise to the initial image. T ,In each step of noise addition, x t-1 Add a Gaussian noise to generate a new latent variable x t ,The single noise adding process is expressed as: in: represents the mean of the Gaussian distribution, β t Ι represents the variance of Gaussian distribution, β t is a hyperparameter that gradually increases with t, and Ι represents the unit matrix with the same dimension as the input sample x0.
7. The background traffic generation method based on the diffusion model and bidirectional flow characteristics according to claim 5 is characterized in that: The reverse denoising process of the diffusion generation module is expressed as: Among them, α t , β t is the diffusion scheduling parameter, ∈ θ is the noise prediction network, z is the standard Gaussian noise, condition emb is the conditional embedding vector.
8. The background traffic generation method based on the diffusion model and bidirectional flow characteristics according to claim 6 is characterized in that: In the residual module of the diffusion generation module, the conditional information is divided into two parts, which are processed by the convolution layer respectively to generate different control signals, which are finally fused with the input signal. The formula is: x t =σ(g)·tanh(f) Among them, g and f are the gate signal and filter signal after conditional convolution processing, σ is the Sigmoid activation function, and tanh is the hyperbolic tangent activation function.
9. The background traffic generation method based on the diffusion model and bidirectional flow characteristics according to claim 6 is characterized in that: The diffusion generation module introduces the diffusion step embedding technology, which maps the step information to a high-dimensional space through a fully connected layer and an activation function. The formula is as follows: yes t =Projection2(Force(Projection1(e′ t ))) where e' t It is the original embedding generated based on the step size, Projection1 and Projection2 are the fully connected layers, and Silu is the activation function.
10. The background traffic generation method based on the diffusion model and bidirectional flow characteristics according to claim 6, characterized in that: The diffusion generation module adopts a cross-modal attention mechanism, which enables the model to effectively combine input data and conditional information, encode and decode input features through multiple self-attention layers, and use the Transformer structure to capture temporal dependencies. The calculation formula of the cross-modal attention mechanism is: Among them, Q, K, and V are generated by linear mapping of spatiotemporal features and conditional embedding vectors respectively, and d is the feature dimension; The spatiotemporal feature image is in Flowpic format, where the packet length is represented by color depth and the flow direction is distinguished by color channels.