High-simulation-degree communication flow generation method based on diffusion model
Through the diffusion model of the Nistrom attention mechanism and the trend decomposition module, the problems of high computational complexity and low data quality in communication traffic generation are solved, and the communication traffic generation with high simulation is achieved, and the prediction ability of the model in long sequence generation tasks is improved.
Patent Information
- Application Number
- CN202510833050.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-08-12
AI Technical Summary
In the field of communication, it is difficult to generate high-quality traffic time series data, the existing models are inefficient in training and inference, it is difficult to capture the daily periodicity and long-tail characteristics of communication traffic, and the standard diffusion model has high computational complexity and is difficult to apply in practice.
Using a diffusion model-based method, through the Nistrom attention mechanism and trend decomposition module, the communication traffic is decomposed into trend terms and fluctuation terms, and the timestamp information is used for training to generate high-simulation communication traffic data.
It significantly reduces the computational complexity, improves the quality of generated data, and is more in line with the real traffic characteristics, improves the Freshette initial distance and correlation score of generated data, and improves the model's prediction ability in long-sequence generation tasks.
Smart Images

Figure CN120474942A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of traffic generation, and in particular to a method for generating high-simulation communication traffic based on a diffusion model. Background Art
[0002] The rapid development of machine learning requires not only algorithmic advancements and increased computing power, but also a vast, high-quality dataset as a foundation. In time series fields such as finance, meteorology, transportation, and health, reliable time series data are a prerequisite for various tasks, including prediction, interpolation, and classification. However, in emerging communications fields like the Industrial Internet within 5G / 6G, collecting real-world data can sometimes be difficult, or real traffic data cannot be used due to concerns about user privacy. This results in a severe shortage of training samples, becoming a bottleneck that restricts model performance. Therefore, in the communications field, generating high-quality traffic time series data is crucial for data augmentation, model training, and privacy protection.
[0003] The task of time series generation aims to generate new sequences that conform to temporal dependencies by learning patterns from historical data. In recent years, this field has made significant progress in model architecture innovation. Variational autoencoders, as a generative framework that fuses autoencoders with probabilistic graphical models, achieve model optimization by jointly minimizing reconstruction error and KL divergence. However, KL divergence can lead to over-regularization of the latent space, blurring the generated samples. Furthermore, the encoder-decoder structure is insufficient in capturing long-term dependencies in time series. Generative adversarial networks (GANs) consist of a generator and a discriminator. The generator receives random noise and generates fake data with the goal of deceiving the discriminator. The discriminator distinguishes between real and generated data and outputs probability values to guide generator optimization. During adversarial training, GANs optimize the objective function through a minimax game. However, GAN model training suffers from issues such as vanishing gradients, pattern collapse, and a potential lack of diversity in generated data.
[0004] In recent years, generative methods based on diffusion models have attracted widespread attention in the field of computer vision. Due to their high generation quality and strong stability, the advantages of such methods have extended to the field of time series analysis. However, to improve the quality of generated data, these models use large-scale deep neural networks as their backbone, significantly increasing the spatiotemporal complexity of training and inference.
[0005] Communication traffic has significant daily cyclical characteristics, such as morning and evening peaks and midday fluctuations. Short time series cannot fully capture these patterns. Burst traffic has a long tail effect, requiring long time series to distinguish short-term fluctuations from long-term trends in order to generate realistic traffic data. However, when using the standard attention mechanism transformer model as the backbone network for diffusion models in long time series generation, training and inference efficiency are low, making it difficult to apply in practice. Furthermore, the large number of low-weight values contained in the standard self-attention matrix dilutes the effective information, making it difficult to improve model performance. Summary of the Invention
[0006] The purpose of the present invention is to provide a method for generating highly simulated communication traffic based on a diffusion model in order to solve the problems existing in the above-mentioned prior art.
[0007] The purpose of the present invention can be achieved by the following technical solutions:
[0008] A method for generating highly realistic communication traffic based on a diffusion model, comprising:
[0009] Step S1: Acquire an original communication traffic data set, wherein the original communication traffic data set consists of a plurality of first sequences, each of which includes a plurality of continuous original communication traffic data;
[0010] Step S2: Based on the pre-configured noise variance, noise is added to the original communication traffic data according to the forward diffusion formula, and the decomposition module is used to decompose the first sequence after the noise addition into a decoder trend term and a decoder fluctuation term as the initial input of the decoder;
[0011] Step S3: The timestamp of the first sequence and the diffusion step of the noise are combined to obtain the external information embedding of the diffusion generation model;
[0012] Step S4: using the first sequence after adding noise, and the trend term and fluctuation term obtained by decomposition to train the diffusion generation model;
[0013] Step S5: Gaussian noise is input into the trained diffusion generation model to obtain a denoised result, and a communication traffic data sequence is generated through a sampling formula.
[0014] The step S1 comprises:
[0015] Step S1-1: Obtain a continuous communication traffic signal with a length of L and a dimension of d;
[0016] Step S1-2: extracting subsequences as first sequences based on the continuous communication traffic signal through a sliding window of length S, and combining all first sequences to obtain an original communication traffic data set.
[0017] The pre-configured noise variance increases with the increase of diffusion steps. The data after adding noise at any time is:
[0018]
[0019]
[0020] Where: x t is the communication traffic data after t diffusion steps and noise addition, is the cumulative noise coefficient, x0 is the original communication traffic data, ∈ is the added noise, I is the unit matrix, is a Gaussian distribution, β s is the noise variance corresponding to the diffusion step s.
[0021] The decoder trend term and decoder fluctuation term are respectively:
[0022]
[0023] Where: Tr is the decoder trend term, is the first sequence after noise addition, Padding(·) is the constraint to keep the length of the input sequence unchanged, AvgPool(·) is the average pooling operation, and Fl is the decoder fluctuation term.
[0024] The step S3 comprises:
[0025] The learnable intra-week sequential feature embedding dictionary is defined as The time embedding dictionary is defined as The diffusion step embedding dictionary is defined as Among them, N w =7 represents the number of days in a week, N d represents the number of time steps in a day, T is the total number of diffusion steps, and d is the embedding dimension;
[0026] For a traffic flow time series of length L, use its timestamp as the index to retrieve the corresponding weekly sequence embedding from the embedding dictionary and time embedding
[0027] Concatenate the weekly sequence embedding and time embedding to get the timestamp encoding
[0028] Retrieve T based on the current diffusion step t diff , get the current noise step embedding
[0029] The timestamp encoding and the noise step number embedding are concatenated through a broadcast mechanism to obtain the external information embedding of the diffusion generation model.
[0030] The step S4 comprises:
[0031] Step S4-1: Input the noise-added first sequence obtained in step S2 into the encoder, and pass it through the Nystrom attention mechanism, decomposition module, feedforward neural network, decomposition module, and adaptive layer normalization in sequence in the encoder to obtain the output of each layer encoder;
[0032] Step S4-2: The trend term and fluctuation term obtained by decomposing the first sequence after adding noise are input into the decoder, and the denoised communication traffic data is obtained by passing through the Nystrom attention mechanism, decomposition module, Nystrom attention mechanism, decomposition module, feedforward neural network, decomposition module, and adaptive layer normalization in sequence in the decoder.
[0033] The step S4-1 specifically includes:
[0034] Obtaining an encoder fluctuation initial term based on the first sequence after adding noise, and inputting the encoder fluctuation initial term into the n-layer encoder;
[0035] The internal dependency of the sequence is obtained through the Nystrom attention mechanism, and after using the residual connection, the fluctuation component is extracted through the decomposition module to obtain the first fluctuation term of the encoder layer l
[0036]
[0037] in: is the first fluctuation term of layer l, is the fluctuation item output of the previous encoder, is the Nystrom attention mechanism, Decomp(·) is the decomposition module, and _ is the discarded part;
[0038] The results will be extracted Through the feedforward neural network and decomposition module, and using residual connection, the second fluctuation term of the encoder layer l is obtained
[0039]
[0040] in: is the second fluctuation term of layer l, FFN(·) is the feedforward neural network operator, and Decomp(·) is the decomposition module operator;
[0041] The results will be extracted The adaptive layer normalization module is fed into the adaptive layer normalization module to control the distribution characteristics of the encoder output to obtain the l-th layer encoder output, wherein the adaptive layer normalization module embeds the obtained external information into the TST embThe range and offset of the adaptive normalization module are obtained through two different fully connected layers, so that the distribution of the generated communication sequence is controlled by both the timestamp and the diffusion time step. The range and offset of the adaptive normalization module and the output of the encoder of the lth layer are specifically:
[0042] Scale=FC1(TST emb )
[0043] Shift=FC2(TST emb )
[0044]
[0045] Where: Scale is the range of the adaptive normalization module, FC1(·) is the first fully connected layer, Shift is the offset of the adaptive normalization module, FC2(·) is the operator of the second fully connected layer, X out is the output of the l-th layer encoder, and LayerNorm(·) is the layer normalization.
[0046] The step S4-2 includes:
[0047] The decoder fluctuation initial term and the decoder trend initial term obtained by the decomposition module are input into the m-layer decoder. The internal dependency of the sequence is obtained through the Nystrom attention mechanism, and the residual connection is used. The fluctuation component and trend classification are then extracted through the decomposition module to obtain the first fluctuation term of the decoder l layer and the first trend term of the decoder l layer:
[0048]
[0049] in: is the first fluctuation term of the decoder layer l, is the first trend item of the decoder layer l, It is the fluctuation item output of the previous layer decoder;
[0050] The first fluctuation item of the decoder layer l and the first trend item of the decoder layer l are input into the Nystrom attention module as queries, and the fluctuation item output by the last layer of the encoder is As the key and value, the residual connection and decomposition module are used to extract the fluctuation component and trend classification, and the second fluctuation item of the decoder layer l and the second trend item of the decoder layer l are obtained:
[0051]
[0052] in: is the second fluctuation term of the decoder layer l, is the second trend item of decoder layer l;
[0053] The second fluctuation term of the decoder layer l and the second trend term of the decoder layer l are passed through the feedforward neural network and decomposition module, and residual connection is used to obtain the third fluctuation term and the third trend term of the decoder layer l:
[0054]
[0055] in: is the third fluctuation term of decoder layer l, is the third trend item of decoder layer l;
[0056] After the adaptive layer normalization process, the output fluctuation term of the decoder layer l is obtained to match the target distribution:
[0057]
[0058] in: is the output fluctuation term of the lth layer, and AdaLN(·) is the adaptive normalization processing under the external information embedding constraint;
[0059] The trend items after each decomposition module are finally mapped to obtain the trend output of the decoder layer:
[0060]
[0061] in: is the trend item of decoder layer l, is the trend term of the decoder layer l-1, W l,i is the projection matrix corresponding to the trend item of the ith decoder in the lth layer, is the i-th trend item of decoder layer l;
[0062] The trend item output by the last layer of the m-layer decoder and fluctuation terms Add them together to get the communication traffic data sequence after model denoising.
[0063] A device for generating high-simulation communication traffic based on a diffusion model comprises a memory, a processor, and a program stored in the memory. When the processor executes the program, the method described above is implemented.
[0064] A storage medium stores a program, which implements the above method when executed.
[0065] Compared with the prior art, the present invention has the following beneficial effects:
[0066] 1. By using the Nystrom attention mechanism instead of the standard self-attention mechanism, the computational complexity can be significantly reduced, redundant calculations can be filtered out, and the quality of generated data can be improved. Compared with advanced models such as DiffusionTS, TimeVAE, and DDPM, the Frechette inception distance (Context-FID) and relevance score (C-Score) of generated communication data are improved.
[0067] 2. The trend decomposition module used can decompose communication time series into trend terms and fluctuation terms, which is more reasonable than traditional direct modeling methods. At the same time, this application also embeds the timestamp information of communication traffic into the model to ensure that the generated sequence is more consistent with the characteristics of actual business scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 Schematic diagram of the main steps of the method of the present invention;
[0069] Figure 2 This is a schematic diagram of the model principle of the forward diffusion and reverse denoising method proposed in the present invention;
[0070] Figure 3 This is a diagram of the backbone network architecture of the method proposed in the present invention;
[0071] Figure 4 This is a visualization diagram of the probability distribution of the generation effect of the proposed method and the comparison algorithm on the Taiwan dataset (the dotted line is the generated data, and the solid line is the original data). DETAILED DESCRIPTION
[0072] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.
[0073] A method for generating high-simulation communication traffic based on diffusion model, such as Figure 1 and Figure 2 As shown, including:
[0074] Step S1: obtaining an original communication traffic data set, wherein the original communication traffic data set consists of a plurality of first sequences, each of which includes a plurality of continuous original communication traffic data;
[0075] Wherein, step S1 includes:
[0076] Step S1-1: Obtain a continuous communication traffic signal with a length of L and a dimension of d
[0077] Step S1-2: extracting subsequences as first sequences based on the continuous communication traffic signal through a sliding window of length S, and combining all first sequences to obtain an original communication traffic data set.
[0078] For sliding windows in, is the observation value at the i-th time step. This constructs the communication traffic data set The goal of the communication traffic generation task is to design a diffusion generation model Gaussian noise Mapping to synthetic traffic sequence Thus, we can obtain A synthetic communication traffic dataset with consistent distribution
[0079] Step S2: Based on the pre-configured noise variance, noise is added to the original communication traffic data according to the forward diffusion formula, and the decomposition module is used to decompose the first sequence after the noise addition into a decoder trend term and a decoder fluctuation term as the initial input of the decoder;
[0080] In the diffusion model, the forward diffusion is based on the set variance β t By adding Gaussian noise to the original communication traffic data, the following Markov chain can be constructed:
[0081]
[0082] Through this process, the original communication traffic data can be gradually converted into isotropic Gaussian noise x T . Define the cumulative noise figure:
[0083]
[0084] Through the reparameterization technique, it is easier to directly sample the data state after adding noise at any time. The preconfigured noise variance increases with the increase of the diffusion step. The data after adding noise at any time is:
[0085]
[0086] Where: x t is the communication traffic data after t diffusion steps and noise addition, is the cumulative noise coefficient, x0 is the original communication traffic data, ∈ is the added noise, obtained by Gaussian distribution sampling, I is the unit matrix, is a Gaussian distribution, β s is the noise variance corresponding to the diffusion step s.
[0087] The interweaving of trends, cycles, and noise in traffic sequences makes it difficult to effectively model and extract key temporal patterns such as inter-cycle dependencies or long-range correlations. Therefore, a decomposition module, denoted as Decomp(·), is used. This design uses a sliding average to separate traffic sequences into trend and fluctuation terms. This progressive decomposition addresses the problem of mixed sequence patterns in long-term predictions. The decoder trend term and decoder fluctuation term are:
[0088]
[0089] Where: Tr is the decoder trend term, is the first sequence after noise addition, Padding(·) is the constraint to keep the length of the input sequence unchanged, AvgPool(·) is the average pooling operation, and Fl is the decoder fluctuation term.
[0090] The first sequence after noise addition is decomposed into decoder fluctuation initial terms through the decomposition module and decoder trend initialization term And input both into the decoder.
[0091] Step S3: The timestamps of the first sequence are combined with the diffusion steps of the noise to obtain the external information embedding of the diffusion generation model. The external information embedding integrates the noise step number information and the timestamp information of the communication traffic. During the timestamp encoding process, in this application, the position of each time step of the communication traffic in the day and its corresponding day of the week are considered.
[0092] Specifically, step S3 includes:
[0093] The learnable intra-week sequential feature embedding dictionary is defined as The time embedding dictionary is defined as The diffusion step embedding dictionary is defined as Among them, N w =7 represents the number of days in a week, N d represents the number of time steps in a day, T is the total number of diffusion steps, and d is the embedding dimension;
[0094] For a traffic flow time series of length L, use its timestamp as the index to retrieve the corresponding weekly sequence embedding from the embedding dictionary and time embedding
[0095] Concatenate the weekly sequence embedding and time embedding to get the timestamp encoding
[0096] Retrieve T based on the current diffusion step t diff , get the current noise step embedding
[0097] The timestamp encoding and the noise step embedding are spliced together through the broadcast mechanism to obtain the external information embedding of the diffusion generation model
[0098] Step S4: Using the first sequence after adding noise, and the trend term and fluctuation term obtained by decomposition, the diffusion generation model is trained, including:
[0099] Step S4-1: The first sequence after noise addition obtained in step S2 is input into the encoder. In the encoder, it passes through the Nystrom attention mechanism, decomposition module, feedforward neural network, decomposition module, and adaptive layer normalization in sequence to obtain the output of each layer encoder. The structure of the diffusion generation model is as follows: Figure 3 As shown,
[0100] Specifically, they include:
[0101] Obtaining an encoder fluctuation initial term based on the first sequence after adding noise, and inputting the encoder fluctuation initial term into the n-layer encoder;
[0102] The internal dependency of the sequence is obtained through the Nystrom attention mechanism, and after using the residual connection, the fluctuation component is extracted through the decomposition module to obtain the first fluctuation term of the encoder layer l
[0103]
[0104] in: is the first fluctuation term of layer l, is the fluctuation item output of the previous encoder, is the Nystrom attention mechanism, Decomp(·) is the decomposition module, and _ is the discarded part;
[0105] The results will be extracted Through the feedforward neural network and decomposition module, and using residual connection, the second fluctuation term of the encoder layer l is obtained
[0106]
[0107] in: is the second fluctuation term of layer l, FFN(·) is the feedforward neural network operator, and Decomp(·) is the decomposition module operator;
[0108] The results will be extracted The adaptive layer normalization module is sent to the adaptive layer normalization module to control the distribution characteristics of the encoder output to obtain the output of the first layer encoder, where the adaptive layer normalization module embeds the obtained external information into the TST embThe range and offset of the adaptive normalization module are obtained through two different fully connected layers, so that the distribution of the generated communication sequence is controlled by both the timestamp and the diffusion time step. The range and offset of the adaptive normalization module and the output of the encoder of the lth layer are specifically:
[0109] Scale=FC1(TST emb )
[0110] Shift=FC2(TST emb )
[0111]
[0112] Where: Scale is the range of the adaptive normalization module, FC1(·) is the first fully connected layer, Shift is the offset of the adaptive normalization module, FC2(·) is the operator of the second fully connected layer, X out is the output of the l-th layer encoder, and LayerNorm(·) is the layer normalization.
[0113] Generally, you can use X out =AdaLN(X in ,TST emb ) to express the above steps, it can be expressed as:
[0114]
[0115] Step S4-2: The trend term and fluctuation term obtained by decomposing the first sequence after adding noise are input into the decoder, and the denoised communication traffic data is obtained by passing through the Nystrom attention mechanism, decomposition module, Nystrom attention mechanism, decomposition module, feedforward neural network, decomposition module, and adaptive layer normalization in sequence in the decoder.
[0116] Specifically, step S4-2 includes:
[0117] The decoder fluctuation initial term and the decoder trend initial term obtained by the decomposition module are input into the m-layer decoder. The internal dependency of the sequence is obtained through the Nystrom attention mechanism, and the residual connection is used. The fluctuation component and trend classification are then extracted through the decomposition module to obtain the first fluctuation term of the decoder l layer and the first trend term of the decoder l layer:
[0118]
[0119] in: is the first fluctuation term of the decoder layer l, is the first trend item of the decoder layer l, It is the fluctuation item output of the previous layer decoder;
[0120] The first fluctuation item of the decoder layer l and the first trend item of the decoder layer l are input into the Nystrom attention module as queries, and the fluctuation item output by the last layer of the encoder is As the key and value, the residual connection and decomposition module are used to extract the fluctuation component and trend classification, and the second fluctuation item of the decoder layer l and the second trend item of the decoder layer l are obtained:
[0121]
[0122] in: is the second fluctuation term of the decoder layer l, is the second trend item of decoder layer l;
[0123] The second fluctuation term of the decoder layer l and the second trend term of the decoder layer l are passed through the feedforward neural network and decomposition module, and residual connection is used to obtain the third fluctuation term and the third trend term of the decoder layer l:
[0124]
[0125] in: is the third fluctuation term of decoder layer l, is the third trend item of decoder layer l;
[0126] After the adaptive layer normalization process, the output fluctuation term of the decoder layer l is obtained to match the target distribution:
[0127]
[0128] in: is the output fluctuation term of the lth layer, and AdaLN(·) is the adaptive normalization processing under the external information embedding constraint;
[0129] The trend items after each decomposition module are finally mapped to obtain the trend output of the decoder layer:
[0130]
[0131] in: is the trend item of decoder layer l, is the trend term of the decoder layer l-1, W l,i is the projection matrix corresponding to the trend item of the ith decoder in the lth layer, is the i-th trend item of decoder layer l;
[0132] The trend item output by the last layer of the m-layer decoder and fluctuation terms Add together to get the communication traffic data sequence after model denoising
[0133] Step S5: Input the Gaussian noise into the trained diffusion generation model to obtain the denoised result, and generate the communication traffic data sequence through the sampling formula.
[0134] During the training process, the neural network directly predicts the flow data before noise addition based on the noisy signal after t diffusion steps. Still Gaussian distribution, variance Usually fixed as To reduce the complexity of model learning. However, the mean calculation depends on the predicted
[0135] According to the forward diffusion formula It can be seen that:
[0136]
[0137] Substituting into the original diffusion model mean formula, we get:
[0138]
[0139] in: is the original data predicted by the model based on the noised data. The loss function can be simplified to:
[0140]
[0141] The sampling formula for generation is as follows, where μ θ (x t ,t) by the predicted Calculation yields:
[0142]
[0143] 1. Experimental parameter settings
[0144] The model proposed in this application is referred to as CommDiff. The generation effect of the method of this application is tested on three real communication network traffic data sets, namely AIIA, Taiwan and Milan. In the AIIA data set, there are 3 nodes, 10,000 time step data, a sampling interval of 60 minutes, and 48 generated points. In the Taiwan data set, there are 6 nodes, 10,000 time step data, a sampling interval of 5 minutes, and 576 generated points. In the Milan data set, there are 10 nodes, 4,320 time step data, a sampling interval of 10 minutes, and 288 generated points.
[0145] At the same time, in order to verify the effectiveness of the proposed method, the most advanced algorithms were selected for comparison, including DiffusionTS, TimeVAE and DDPM.
[0146] 2. The effect of the algorithm in this application on improving prediction accuracy
[0147] For each dataset, a time series covering two consecutive days was generated. For example, for the Milan dataset, the number of points was generated. The prediction results were measured using two widely adopted metrics: the Context-Aware Fréchette Inception Distance (Context-FID) and the Correlation Score (C-Score). These metrics provide a comprehensive quality assessment of the model generation results based on the similarity of the feature distribution between the generated data and the real data, and the consistency of global temporal dependencies, respectively. Lower values indicate better model generation.
[0148] As can be seen from the table, the model achieved the best results across all datasets, particularly on the long sequence generation task in the Taiwan dataset, where it significantly outperformed other algorithms. The Diffusion-TS model performed second best, performing well on AIIA and Milan but struggling with long sequence generation. While the inherent latent space continuity constraint in TimeVAE facilitates noise reduction, it weakens the ability to represent complex temporal dynamics. DDPM was unable to fully leverage the characteristics of time series to improve the overall quality of generated data when faced with a high-variable generation task like the Milan dataset.
[0149] 4. Impact of each component on prediction accuracy
[0150] In order to systematically evaluate the effectiveness of each key module in the model of this application, a series of ablation experiments were designed to compare the performance of different variant models to gain an in-depth understanding of the impact of each component on the overall performance. Specifically, this application constructed and tested the following five variant models: replacing Nystrom attention with standard self-attention, removing the adaptive layer normalization module, removing the decomposition module, and removing the encoder module. When the standard attention mechanism is used, the model performs close to the Nystrom attention method in short-term traffic generation tasks, but the performance drops significantly in long-term traffic generation scenarios. This shows that Nystrom attention can effectively improve the model's predictive ability in long sequence tasks by removing redundant information in long sequences and extracting key features. We can also see from Table 1 that adaptive layer normalization effectively improves the quality of the data generated by the model, the decomposition module successfully extracts the internal patterns of the communication sequence, and the encoder is crucial in mining the complex relationships between long sequences.
[0151] Table 1
[0152]
[0153] 5. Operational efficiency of this application method
[0154] To further evaluate the model performance, Figure 4 As shown in Table 2 and Table 3, the memory usage and running efficiency of CommDiff and the standard attention mechanism model in the training and generation stages are compared and analyzed on the Taiwan dataset.
[0155] Table 2
[0156]
[0157] Table 3
[0158]
[0159] Experimental results show that, thanks to the application of Nystrom's attention mechanism, CommDiff demonstrates significant advantages in both computational efficiency and memory optimization. In particular, during the generation phase, its computational time and memory usage are reduced by 66.0% and 46.5%, respectively.
[0160] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
Claims
1. A method for generating high-fidelity communication traffic based on a diffusion model, characterized in that: include: Step S1: Acquire an original communication traffic data set, wherein the original communication traffic data set consists of a plurality of first sequences, each of which includes a plurality of continuous original communication traffic data; Step S2: Based on the pre-configured noise variance, noise is added to the original communication traffic data according to the forward diffusion formula, and the decomposition module is used to decompose the first sequence after the noise addition into a decoder trend term and a decoder fluctuation term as the initial input of the decoder; Step S3: The timestamp of the first sequence and the diffusion step of the noise are combined to obtain the external information embedding of the diffusion generation model; Step S4: using the first sequence after adding noise, and the trend term and fluctuation term obtained by decomposition to train the diffusion generation model; Step S5: Gaussian noise is input into the trained diffusion generation model to obtain a denoised result, and a communication traffic data sequence is generated through a sampling formula.
2. The method for generating high-simulation communication traffic based on a diffusion model according to claim 1, characterized in that: The step S1 comprises: Step S1-1: Obtain a continuous communication traffic signal with a length of L and a dimension of d; Step S1-2: extracting subsequences as first sequences based on the continuous communication traffic signal through a sliding window of length S, and combining all first sequences to obtain an original communication traffic data set.
3. The method for generating high-simulation communication traffic based on a diffusion model according to claim 1, characterized in that: The pre-configured noise variance increases with the increase of diffusion steps. The data after adding noise at any time is: Where: x t is the communication traffic data after t diffusion steps and noise addition, is the cumulative noise coefficient, x0 is the original communication traffic data, ∈ is the added noise, I is the unit matrix, is a Gaussian distribution, β s is the noise variance corresponding to the diffusion step s.
4. The method for generating high-simulation communication traffic based on a diffusion model according to claim 3, characterized in that: The decoder trend term and decoder fluctuation term are respectively: Where: Tr is the decoder trend term, is the first sequence after noise addition, Padding(·) is the constraint to keep the length of the input sequence unchanged, AvgPool(·) is the average pooling operation, and Fl is the decoder fluctuation term.
5. The method for generating high-simulation communication traffic based on a diffusion model according to claim 1, characterized in that: The step S3 comprises: The learnable intra-week sequential feature embedding dictionary is defined as The time embedding dictionary is defined as The diffusion step embedding dictionary is defined as Among them, N w =7 represents the number of days in a week, N d represents the number of time steps in a day, T is the total number of diffusion steps, and d is the embedding dimension; For a traffic flow time series of length L, use its timestamp as the index to retrieve the corresponding weekly sequence embedding from the embedding dictionary and time embedding Concatenate the weekly sequence embedding and time embedding to get the timestamp encoding Retrieve T based on the current diffusion step t diff , get the current noise step embedding The timestamp encoding and the noise step number embedding are concatenated through a broadcast mechanism to obtain the external information embedding of the diffusion generation model.
6. The method for generating high-simulation communication traffic based on a diffusion model according to claim 1, characterized in that: The step S4 comprises: Step S4-1: Input the noise-added first sequence obtained in step S2 into the encoder, and pass it through the Nystrom attention mechanism, decomposition module, feedforward neural network, decomposition module, and adaptive layer normalization in sequence in the encoder to obtain the output of each layer encoder; Step S4-2: The trend term and fluctuation term obtained by decomposing the first sequence after adding noise are input into the decoder, and the denoised communication traffic data is obtained by passing through the Nystrom attention mechanism, decomposition module, Nystrom attention mechanism, decomposition module, feedforward neural network, decomposition module, and adaptive layer normalization in sequence in the decoder.
7. The method for generating high-simulation communication traffic based on a diffusion model according to claim 6, characterized in that: The step S4-1 specifically includes: Obtaining an encoder fluctuation initial term based on the first sequence after adding noise, and inputting the encoder fluctuation initial term into the n-layer encoder; The internal dependency of the sequence is obtained through the Nystrom attention mechanism, and after using the residual connection, the fluctuation component is extracted through the decomposition module to obtain the first fluctuation term of the encoder layer l in: is the first fluctuation term of layer l, is the fluctuation item output of the previous encoder, is the Nystrom attention mechanism, Decomp(·) is the decomposition module, and _ is the discarded part; The results will be extracted Through the feedforward neural network and decomposition module, and using residual connection, the second fluctuation term of the encoder layer l is obtained in: is the second fluctuation term of layer l, FFN(·) is the feedforward neural network operator, and Decomp(·) is the decomposition module operator; The results will be extracted The adaptive layer normalization module is fed into the adaptive layer normalization module to control the distribution characteristics of the encoder output to obtain the l-th layer encoder output, wherein the adaptive layer normalization module embeds the obtained external information into the TST emb The range and offset of the adaptive normalization module are obtained through two different fully connected layers, so that the distribution of the generated communication sequence is controlled by both the timestamp and the diffusion time step. The range and offset of the adaptive normalization module and the output of the encoder of the lth layer are specifically: Scale=FC1(TST emb ) Shift=FC2(TST emb ) Where: Scale is the range of the adaptive normalization module, FC1(·) is the first fully connected layer, Shift is the offset of the adaptive normalization module, FC2(·) is the operator of the second fully connected layer, X out is the output of the l-th layer encoder, and LayerNorm(·) is the layer normalization.
8. The method for generating high-simulation communication traffic based on a diffusion model according to claim 7, characterized in that: The step S4-2 includes: The decoder fluctuation initial term and the decoder trend initial term obtained by the decomposition module are input into the m-layer decoder. The internal dependency of the sequence is obtained through the Nystrom attention mechanism, and the residual connection is used. The fluctuation component and trend classification are then extracted through the decomposition module to obtain the first fluctuation term of the decoder l layer and the first trend term of the decoder l layer: in: is the first fluctuation term of the decoder layer l, is the first trend item of the decoder layer l, It is the fluctuation item output of the previous layer decoder; The first fluctuation item of the decoder layer l and the first trend item of the decoder layer l are input into the Nystrom attention module as queries, and the fluctuation item output by the last layer of the encoder is As the key and value, the residual connection and decomposition module are used to extract the fluctuation component and trend classification, and the second fluctuation item of the decoder layer l and the second trend item of the decoder layer l are obtained: in: is the second fluctuation term of the decoder layer l, is the second trend item of decoder layer l; The second fluctuation term of the decoder layer l and the second trend term of the decoder layer l are passed through the feedforward neural network and decomposition module, and residual connection is used to obtain the third fluctuation term and the third trend term of the decoder layer l: in: is the third fluctuation term of decoder layer l, is the third trend item of decoder layer l; After the adaptive layer normalization process, the output fluctuation term of the decoder layer l is obtained to match the target distribution: in: is the output fluctuation term of the lth layer, and AdaLN(·) is the adaptive normalization processing under the external information embedding constraint; The trend items after each decomposition module are finally mapped to obtain the trend output of the decoder layer: in: is the trend item of decoder layer l, is the trend term of the decoder layer l-1, W l,i is the projection matrix corresponding to the trend item of the ith decoder in the lth layer, is the i-th trend item of decoder layer l; The trend item output by the last layer of the m-layer decoder and fluctuation terms Add them together to get the communication traffic data sequence after model denoising.
9. A device for generating high-simulation communication traffic based on a diffusion model, comprising a memory, a processor, and a program stored in the memory, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 8 is implemented.
10. A storage medium having a program stored thereon, characterized in that: When the program is executed, the method according to any one of claims 1 to 8 is implemented.