Communication network flow data generation method based on mechanism knowledge and generative model
By combining mechanism knowledge and generative models, using CcGAN and LSTM structures to generate network traffic data, the problem that the existing technology cannot generate realistic network traffic data is solved, and high-precision data support is achieved, which is suitable for communication network traffic management and the construction of digital twin networks.
Patent Information
- Application Number
- CN202510222215.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-06
AI Technical Summary
Existing business modeling methods cannot generate a large amount of realistic network traffic data, it is difficult to simulate the continuity and periodicity in real business data, and it cannot effectively support traffic prediction and the construction of digital twin networks.
The communication network traffic data generation method based on mechanism knowledge and generative model is adopted, and the trend terms, season terms and random terms are obtained through STL timing decomposition, and a data generation model based on CcGAN is constructed. The LSTM structure is used to capture the time and statistical characteristics of the traffic data to generate a network traffic data sequence that is approximated to the true distribution.
It realizes the generation of a large number of network traffic data sequences that are approximate to the true distribution at a specified time point on demand, retains the time characteristics of the data sequence, reduces the cost of acquiring service data, and supports the communication network traffic management, service load prediction and the construction of digital twin networks.
Smart Images

Figure CN120105104A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of wireless communications, and in particular to a method for generating communication network traffic data based on mechanism knowledge and a generative model. Background Art
[0002] With the rapid development of mobile communication technology, users' demand for networks continues to increase, the application scenarios of communication networks are becoming more diverse, and emerging services such as the Internet of Things, virtual reality, and autonomous driving are constantly emerging. The number of mobile communication users is growing at an exponential rate. Therefore, the service traffic in the network continues to increase at an unprecedented scale. At the same time, more diversified service scenarios such as enhanced mobile broadband (eMBB), ultra-reliable low-latency communication (URLLC), and massive machine type communication (mMTC) have stricter requirements on network performance such as data rate, reliability, latency, and availability. In order to better meet users' extreme and diverse service needs, it is necessary to study network traffic management and service prediction. By predicting users' traffic characteristics in advance, resources can be scheduled on demand in a preset manner, thereby reducing transmission delays, enhancing user service experience, and improving resource utilization. In recent years, with the extensive application of machine learning technology, accurate network traffic prediction requires a large amount of data to ensure the training and verification of algorithms. In actual applications, due to issues such as user privacy, it is difficult to obtain a large amount of real public service data. An accurate traffic generation model can make up for this shortcoming.
[0003] At the same time, with the research and development of digital twin technology, digital twin networks can help operators effectively manage and maintain physical networks, and model users and various network elements in the virtual domain to form twin networks. Through simulation, prediction, emulation and decision-making in the twin network, real-time monitoring of network status, user behavior prediction and communication strategy, resource allocation strategy optimization, etc. can be achieved, thereby achieving a high degree of network autonomy without affecting the actual physical network. Among them, accurate business traffic modeling is an important part of building a digital twin network.
[0004] In the future, the types of communication services will be diverse, and network traffic will show multiple periodic patterns in time, with unique changes every day and every week. Traditional mechanism-based modeling methods can only build models for individual business scenarios, and it is difficult to adapt to the complex and diverse business types in future mobile communication scenarios. In addition, the mechanism model only captures the statistical significance of business data, and does not characterize the periodicity and correlation of business data in time. In actual applications, due to factors such as user behavior habits and seasonal characteristics, network business traffic often shows multiple periodicities in time and large-scale correlation. The traffic data generated by the traditional three-layer business model is independent and unrelated in time, making it difficult to simulate the actual business data sequence.
[0005] There are also related studies in the prior art that propose to use deep learning to establish business models, such as learning the distribution of business data IP packets through Generative Adversarial Networks (GAN), thereby generating traffic data close to reality. However, the GAN network can only learn the average distribution of a large amount of data, and cannot control the type of generated data, making it difficult to generate data that obeys a specific distribution on demand. In addition, the data generated by this method is not related in time, and cannot simulate the continuity and periodicity in real business data, and cannot be directly applied to the training and verification of algorithms such as traffic prediction.
[0006] In summary, the existing business modeling methods cannot generate a large amount of realistic network traffic data. Therefore, it is necessary to study efficient and accurate network traffic data generation models to generate network traffic data that is close to the actual distribution and conforms to its time characteristics, so as to better support the development of functions such as traffic management and resource optimization, as well as the creation of digital twin networks. Summary of the invention
[0007] In view of the difficulty in obtaining traffic data in the existing network, the present invention proposes a communication network traffic data generation method based on mechanism knowledge and generative models. The present invention decouples data by mechanism analysis, constructs a data generation model using a generative adversarial neural network with a long short-term memory (LSTM) structure, and captures the characteristics of traffic data in terms of time and statistics. It can generate a large number of network traffic data sequences that are close to the real distribution at a specified time point on demand.
[0008] In order to achieve the above-mentioned invention object, the technical solution adopted by the present invention is:
[0009] The communication network traffic data generation method based on mechanism knowledge and generative model includes the following steps:
[0010] Step 1, perform STL time series decomposition on the original flow data sequence to obtain trend items, seasonal items and random items;
[0011] Step 2: construct a data generation model based on CcGAN, take the seasonal term as the learning object of the data generation model, and train the data generation model;
[0012] Step 3: Perform online data augmentation through the data generation model, splice the augmented data according to the positions marked by the regression labels to obtain a seasonal item data sequence of the required length, and then superimpose it with the trend item and random item obtained in step 1 to synthesize a complete network traffic data sequence.
[0013] Furthermore, the specific method of step 1 is:
[0014] Step 101: the original traffic data sequence X=[x 1 ,x 2 ,...,x N ] for each point x n The local data is weighted regressed. The window size of weighted regression is the length of the entire original traffic data sequence. The trend item sequence Y = [y 1 ,y 2 ,...,y N ]The value of each element in:
[0015]
[0016] Among them, y n For position x i The smoothed estimate, w i is a weight function used to measure the position x n With position x i The effect of the distance on the estimate is defined using a Gaussian kernel:
[0017]
[0018] Where h is the smoothing parameter;
[0019] Step 102: remove the trend term from the original traffic data sequence X to obtain the detrended data Z = [z 1 ,z 2 ,...,z N ], where each element is:
[0020] z n =x n -y n
[0021] Perform Fourier spectrum analysis on the detrended data:
[0022]
[0023] Find the V main frequency components [f 1 ,..,f v ,...,f V ], each main frequency component represents a periodic feature of the data series, and these main frequency components are used to determine the window size of seasonal smoothing Indicates rounding up, thereby extracting V seasonal components representing different periodic characteristics:
[0024]
[0025] Wherein, Loess represents the Loess smoothing estimation function;
[0026] For seasonal ingredients Summing, we get:
[0027]
[0028] Finally, we get the seasonal term sequence S = [s 1 ,...,s n ,...,s N ];
[0029] Step 103, remove the trend term and seasonal term components from the original flow data sequence, and the remaining part is the random term sequence R = [r 1 ,...,r n ,...,r N ].
[0030] Furthermore, the specific method of step 2 is:
[0031] Step 201, constructing a data generation model based on CcGAN, wherein both the generator and the discriminator adopt an LSTM network structure, the generator receives conditional labels and realizes the combination with the conditional labels through layer normalization, and in the discriminator, the labels are embedded into the input of the discriminator through label projection to distinguish the input data corresponding to different conditional labels;
[0032] Step 202, taking the seasonal item sequence S as the learning object of the data generation model, dividing the sequence S into M data segments of length L, M=N / L, a single data segment as the input of the model, L is a hyperparameter; the continuity condition of the model is the label corresponding to each data segment, and the label is determined by the time period of the data, that is, the combination of indexes corresponding to each seasonal item cycle;
[0033] Step 203, setting the loss functions of the generator and the discriminator. The loss function of the discriminator includes two types: a hard neighborhood discriminator loss function and a soft neighborhood discriminator loss function. Among them, the hard neighborhood discriminator loss function is:
[0034]
[0035] Among them, C 1 ,C 2 are constants respectively; represents the i-th real data sample, represents the i-th data sample generated by the generator; q j represents the conditional label of the jth data sample; N r is the number of true samples that meet the conditions in the neighborhood of the specified conditional label, N gis the number of generated samples; ∈ r and ∈ g represents the Gaussian noise added to the label; D is the discriminator; is Gaussian distribution; E is the expected function; W 1 ,W 2 Respectively represent the weight coefficients of the samples:
[0036]
[0037] Among them, k is the neighborhood radius; Refers to satisfying |qq j -∈ r |≤kq j The number of Refers to satisfying |qq j -∈ g |≤kq j The number of A is a knowledge function supported by subscripts, which is used to indicate whether an element belongs to a set. If element x belongs to set A, then 1 A (x)=1, otherwise 1 A (x) = 0;
[0038] The soft neighborhood discriminator loss function is:
[0039]
[0040] Among them, C 3 ,C 4 are constants, W 3 ,W 4 Represents the weight coefficient of the sample:
[0041]
[0042] Among them, the ω function is defined as:
[0043]
[0044] μ is a hyperparameter, μ>0;
[0045] The generator loss function is:
[0046]
[0047] Among them, G represents the generator;
[0048] Step 204, the generator and the discriminator are trained in an adversarial manner, the goal of the generator is to maximize the error probability of the discriminator, and the goal of the discriminator is to minimize the classification error; the specific training process is as follows:
[0049] 1) Split the seasonal item sequence S into sequence segments of length L in chronological order, and combine their corresponding continuous labels as the input of the data generation model to form the training data set of the network;
[0050] 2) Using continuous labels as control conditions, randomly generated noise is input into the generator to generate data segments corresponding to the labels;
[0051] 3) The seasonal item data fragment generated by the generator Actual seasonal data snippet And the control condition q i Input into the discriminator together;
[0052] 4) Select the loss function of the discriminator as Or a weighted function of the two, calculate the gradient of the discriminator loss function, and update the parameters of the discriminator through back propagation;
[0053] 5) Calculate the generator loss function and update the generator parameters through back propagation;
[0054] 6) Repeat steps 2) to 5) until the loss functions of the generator and discriminator stabilize and reach a balance.
[0055] Furthermore, the specific method of step 3 is:
[0056] Step 301, use the trained generator G to perform online data augmentation, given a continuous label, input random noise, and obtain the data segment corresponding to the label;
[0057] Step 302: for the data fragments at M different positions generated by the generator, use the symbol Represents the mth data segment generated by the generator, and its corresponding label is q m , splice M data fragments into a data sequence according to the position of the label
[0058] Step 303: convert the data sequence S r The required flow data is synthesized by superimposing the trend item sequence and random item sequence X r Each element The calculation method is:
[0059]
[0060] The beneficial effects of the present invention are:
[0061] 1. The present invention decouples data by means of mechanism analysis, and constructs a data generation model using a generative adversarial neural network with a long short-term memory (LSTM) structure. It captures the temporal and statistical characteristics of traffic data at the same time, and can generate a large number of network traffic data sequences that are close to the real distribution at a specified time point on demand.
[0062] 2. The present invention retains the temporal characteristics of the sequence while learning the data distribution characteristics, and can obtain a large amount of data close to the actual business characteristics in a short time, greatly reducing the cost of obtaining business data.
[0063] 3. The present invention uses a trained generator to quickly generate realistic traffic data sequences. The application scenarios can be network traffic at a base station or business traffic at a single user. It can solve the problem of difficulty in acquiring data in the existing network and provide high-precision data support for communication network traffic management, business load prediction, and the construction of digital twin networks. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 It is a schematic diagram of the principle of the flow data generation method based on mechanism knowledge and generative model.
[0065] Figure 2 This is a schematic diagram of the data generation model based on CcGAN. DETAILED DESCRIPTION
[0066] In order to facilitate the understanding of the technical solution of the present invention by those skilled in the art, and at the same time, to make the technical purpose, technical solution and beneficial effects of the present invention clearer, the technical solution of the present invention is further and more detailedly explained in the form of specific cases below.
[0067] A traffic data generation method based on mechanism knowledge and generative models. The method first extracts different seasonal terms through Fourier spectrum analysis to obtain the implicit periodic characteristics of the data, takes different periodic label sequences in the seasonal terms as conditions, and uses a continuous conditional generative adversarial network (CcGAN) to generate data at different time positions. The generated data fragments are then synthesized into a target length in sequence, and then the trend term and random term are superimposed to obtain a traffic data sequence of the required length.
[0068] This method is based on mechanism analysis and combines the CcGAN generative model to propose a generation mechanism for network / business traffic data. Figure 1 As shown, the method comprises the following steps:
[0069] Step 1: Data decoupling based on mechanism knowledge.
[0070] Some communication network service data sets are known. Service data usually exhibits complex time dynamic characteristics, so it is necessary to decouple the data. Here, we adopt (Seasonal and Trend Decomposition using Loess, STL) time series decomposition. As a decomposition method based on local weighted regression (Loess), STL decomposition can handle complex and changing seasonal characteristics, and obtain the main characteristic component seasonal term as the input of the generative model to simplify the subsequent data generation process.
[0071] The traffic data is decoupled through STL decomposition to obtain trend terms, seasonal terms and random terms. First, the entire original sequence is subjected to weighted regression to obtain the trend term, which reflects the long-term trend of the data and generally changes slowly. Next, the sequence after the trend term is removed is subjected to Fourier spectrum analysis. The spectrum information of the business data contains the laws of business operation in time and complex periodic changes. After Fourier transform, the main frequency components are extracted to obtain the periodic characteristics of the business, and multiple seasonal terms are extracted according to each main frequency component. Finally, the trend term and seasonal term are removed from the original sequence to obtain the random term, which is the sum of irregular noise or random fluctuations and the remaining periodic terms with smaller components.
[0072] Step 2: Generate a data model based on CcGAN.
[0073] Through the aforementioned data decoupling, trend items, seasonal items and random items can be obtained. After removing trend items and random items, the seasonal item contains a large number of characteristics of the time dimension and eliminates the influence of random fluctuations and noise. It is the main component of traffic data. Therefore, the present invention selects the seasonal item as the learning object of the data generation model.
[0074] In order to better learn the temporal distribution characteristics of traffic data, this method uses CcGAN with LSTM structure to build a generative model. CcGAN network is a generative adversarial network with continuous digital labels. The labels of traditional CGAN are discrete categories, while CcGAN combines an improved label input mechanism, whose labels can be continuous numbers, namely regression labels, so that the continuous condition of time scale can be better input into the generator and discriminator, so that the generated business data has more temporal consistency and authenticity. This method uses a combination of multiple main frequency components as the control condition of the CcGAN model, and uses the seasonal items obtained by the above data decoupling as the training data set. The control condition marks the position of the input data block in the entire time series, maps it to a high-dimensional vector space through an embedding layer, and inputs it into the generator and discriminator respectively, to assist the generator in learning the distribution of seasonal item data sequence blocks at different positions. The generator is responsible for simulating the distribution of network traffic seasonal item data at different time positions, while the discriminator is used to evaluate the authenticity of the generated data. Specific discriminator losses are introduced, including Hard Vicinal Discriminator Loss (HVDL) based on hard neighborhood samples and Soft Vicinal Discriminator Loss (SVDL) based on soft neighborhood samples, to improve the model's ability to generate high-quality data under continuous conditions. Through adversarial learning between the discriminator and the generator, the trained generator can generate data sequence blocks at specified positions.
[0075] Step 3: data splicing and synthesis.
[0076] By splicing the data generated by the CcGAN model according to the position marked by its regression label, we can get the seasonal item data sequence of the required length, and then superimpose it with the trend item and random item obtained by data decoupling to synthesize the complete network traffic data sequence. This not only simulates the distribution of real data, but also retains the temporal characteristics of the data sequence.
[0077] Since CcGAN can synthesize a large amount of data as required, this method can provide the required number of realistic data sequences in a short time, which can be either traffic data for base stations or different types of service data for single users, thereby supporting network simulation optimization, improving the accuracy of service load prediction, and providing reliable data support for the diversified application scenarios of future mobile communications.
[0078] Here is a more concrete example:
[0079] A method for generating traffic data based on mechanism knowledge and generative models is specifically implemented as follows:
[0080] Step 1. Data decoupling based on mechanism knowledge.
[0081] It is known that a raw traffic data sequence X=[x 1 ,x 2 ,...,x N ], and perform STL time series decomposition on it to obtain trend terms, seasonal terms and random terms in turn.
[0082] 1.1 Trend Item Extraction
[0083] The Loess method is to calculate the n The local data is weighted and regressed to smooth the time series data to estimate the trend or seasonal component. The difference is that in the estimation of the trend term, the window size of the weighted regression is the entire original sequence length, while the weighted regression window in the seasonal term calculation is the inverse of the main frequency component. 1 ,y 2 ,...,y N ] represents the long-term trend, and its extraction process is as follows:
[0084]
[0085] Among them, y n For position x i The smoothed estimate, w is the weight function used to measure the position x n With position x i The effect of the distance on the estimate is usually defined using a Gaussian kernel:
[0086]
[0087] Among them, h is the smoothing parameter, which determines the range of the local area and affects the degree of smoothing.
[0088] 1.2 Seasonal item extraction
[0089] First, the original sequence is detrended, and the trend term is removed from the original traffic data sequence X to obtain the detrended data Z = [z 1 ,z 2 ,...,z N ]:
[0090] z n =x n -y n
[0091] Next, perform Fourier spectrum analysis on the detrended data:
[0092]
[0093] By observing the spectrum F(f), we can find several main frequency components with higher amplitudes [f 1 ,..,f v ,...,f V ], V is the number of main frequency components, each main frequency represents a periodic feature of the data series, and multiple main frequency components indicate that the flow data series presents multiple patterns of periodicity in time. These frequencies are used to determine the window size of seasonal smoothing Thus, multiple seasonal components representing different periodic characteristics are extracted. The application window size is w v The Loess smoothing method is used to estimate the seasonal term
[0094]
[0095] Then the sum of all seasonal terms gives the data sequence S = [s 1 ,...,s n ,...,s N ], as the object of subsequent generative model learning, s n The calculation method is as follows:
[0096]
[0097] 1.3 Random Term Calculation
[0098] Calculate the residual term, that is, the random term R = [r 1 ,...,r n ,...,r N ], the remainder after removing the trend and seasonal components from the original traffic data series is the random term:
[0099] r n =x n -y n -s n ,n=1,...,N
[0100] Step 2. Data generation model based on CcGAN
[0101] 2.1 Network Structure of CcGAN
[0102] Both the generator and the discriminator adopt the LSTM network structure to effectively extract the features of the time series input data. The generator receives the conditional label and combines it with the conditional label through layer normalization. In the discriminator, the label is embedded into the input of the discriminator through label projection so that the discriminator can distinguish the input data corresponding to different conditional labels. The overall structure of the CcGAN network is as follows: Figure 2shown.
[0103] The seasonal term sequence S obtained in the first step of data decoupling is used as the learning object of the generative model. The sequence S is divided into M data segments of length L, where M = N / L. A single data segment is used as the input of the model, and L is a hyperparameter. The continuity condition of the model is the label corresponding to each data segment, which is determined by the time period in which the data is located, that is, the combination of indexes corresponding to each seasonal term cycle. For example, there are 3 main frequency components of the seasonal term, and their reciprocals are the cycle sizes, which are expressed as [a, b, c] respectively. There are 7 a's in a b cycle and 4 b's in a c cycle. The entire data sequence contains 3 c cycles in total, so the label can be expressed as [a i ,b j ,c u ],i∈[0,6],j∈[0,3],u∈[0,2]. Assuming that the length of the data segment L = the length of the minimum period, the label value of the first segment of data is [a 0 ,b 0 ,c 0 ], the second data label value is [a 1 ,b 0 ,c 0 ], the eighth data segment label is [a 0 ,b 1 ,c 0 ]. This multi-dimensional continuous label is used to indicate each input data fragment, and the normalization layer and projection layer are used to inform the network of this temporal feature, so as to better guide the generator to generate data fragments of the corresponding type (i.e., position), and improve the temporal consistency of the generated data with the real data.
[0104] 2.2 Loss Function of CcGAN
[0105] The generator and the discriminator are trained through an adversarial process. The goal of the generator is to generate more realistic data as much as possible, while the goal of the discriminator is to distinguish between real data and generated data as much as possible. Therefore, the goal of the generator is to maximize the error probability of the discriminator, while the goal of the discriminator is to minimize the classification error. The discriminator loss of the CcGAN network includes two types: hard neighbor discriminator loss (HVDL) and soft neighbor discriminator loss (SVDL) to deal with the scarcity of generated data under regression labels. HVDL or SVDL can be selected according to the data generation effect in actual applications, or a weighted combination of HVDL and SVDL can be used as the loss function of the discriminator according to performance indicators. At the same time, the loss function of the generator has also been newly adjusted.
[0106] 2.2.1 Loss Function of Discriminator
[0107] (1)HVDL
[0108] The hard neighborhood is defined by a neighborhood with fixed boundaries and only samples within a certain range are selected.
[0109]
[0110] Among them, C 1 ,C 2 are constants respectively; represents the i-th real data sample, represents the i-th data sample generated by the generator; q j represents the conditional label of the jth data sample; N r is the number of true samples that meet the conditions in the neighborhood of the specified conditional label, N g is the number of generated samples; ∈ r and ∈ g represents the Gaussian noise added to the label; W 1 ,W 2 Respectively represent the weight coefficients of the samples:
[0111]
[0112] in, Refers to satisfying |qq j -∈ r |≤kq j The number of Refers to satisfying |qq j -∈ g |≤kq j k is the neighborhood radius, which is a threshold defined in the label space and is used to determine whether a sample belongs to the neighborhood range of a specific label. A is a knowledge function supported by subscripts, which is used to indicate whether an element belongs to a set. If element x belongs to set A, then 1 A (x)=1, otherwise 1 A (x)=0.
[0113] (2)SVDL
[0114] Compared with hard decision, soft decision defines the domain by randomly perturbing the label and sampling near the label in a weighted manner. The loss function differs from hard decision only in constant and weight coefficient.
[0115]
[0116] Among them, C 3 ,C 4 are constants, W 3 ,W 4 Represents the weight coefficient of the sample:
[0117]
[0118] Among them, the ω function is defined as:
[0119]
[0120] μ is a hyperparameter and needs to satisfy μ>0.
[0121] 2.2.2 Generator loss function:
[0122]
[0123] This loss function adds Gaussian noise to the labels based on the standard generator loss function, allowing the generator to tolerate small label perturbations, that is, to generate realistic data even on continuous labels not seen in the training set, thus solving the limitations of traditional CGAN on continuous labels.
[0124] 2.3 Training steps of CcGAN
[0125] 1) Split the seasonal item data sequence S into sequence segments of length L in chronological order, and combine their corresponding continuous labels as the input of the CcGAN network to form the training data set of the network;
[0126] 2) Using continuous labels as control conditions, randomly generated noise is input into the generator to generate data segments corresponding to the labels;
[0127] 3) The seasonal item data fragment generated by the generator Actual seasonal data snippet And the control condition q i Input into the discriminator together;
[0128] 4) Calculate the loss function of the discriminator and choose according to the actual situation / The weighted function of the two, calculates its gradient, and updates the parameters of the discriminator through back propagation, so that the discriminator can better distinguish between real data and generated data;
[0129] 5) According to Calculate the generator loss function and update the generator parameters through back propagation so that the data generated by the generator can deceive the discriminator as much as possible;
[0130] 6) Repeat steps 2-5 until the loss functions of the generator and discriminator tend to stabilize and reach a balance;
[0131] 7) The trained generator G can be used for online data augmentation. Given a continuous label, the data segment corresponding to the label can be obtained by inputting random noise.
[0132] Step 3. Data splicing and synthesis
[0133] Assume that the generator generates a total of M data fragments at different positions, symbolized by Represents the mth data segment generated by the generator, and its corresponding label is q m , splice M data fragments into a data sequence according to the position of the label
[0134] The generated seasonal item data is superimposed with the trend item and random item to synthesize the required flow data The calculation method is:
[0135]
[0136] In summary, the present invention combines the mechanism characteristics of business data, decouples the traffic data through time series decomposition and Fourier spectrum analysis, and obtains the main component seasonal term that reflects the dynamic changes in time. By utilizing the learning ability of the CcGAN network with LSTM structure for time series data, introducing an improved label input mechanism and hard neighborhood and soft neighborhood discriminator losses, the distribution characteristics and time characteristics of seasonal term data blocks can be accurately captured. The present invention can solve the problem of difficult data acquisition in existing networks, and provide high-precision data support for communication network traffic management, business load forecasting, and the construction of digital twin networks.
[0137] Although the above illustrative specific embodiments of the present invention are described to facilitate the understanding of the present invention by those skilled in the art, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the claims, these changes are obvious, and all inventions and creations using the concept of the present invention are protected.
Claims
1. A method for generating communication network traffic data based on mechanism knowledge and generative model, characterized in that: The following steps are involved: Step 1, perform STL time series decomposition on the original flow data sequence to obtain trend items, seasonal items and random items; Step 2: construct a data generation model based on CcGAN, take the seasonal term as the learning object of the data generation model, and train the data generation model; Step 3: Perform online data augmentation through the data generation model, splice the augmented data according to the positions marked by the regression labels to obtain a seasonal item data sequence of the required length, and then superimpose it with the trend item and random item obtained in step 1 to synthesize a complete network traffic data sequence.
2. The method for generating communication network traffic data based on mechanism knowledge and generative model according to claim 1 is characterized in that: The specific method of step 1 is: Step 101: the original traffic data sequence X = [x1, x2, ..., x N ] for each point x n The local data is weighted regressed. The window size of the weighted regression is the length of the entire original traffic data sequence. The trend item sequence Y = [y1, y2, ..., y N ]The value of each element in: Among them, y n For position x i The smoothed estimate, w i is a weight function used to measure the position x n With position x i The effect of the distance on the estimate is defined using a Gaussian kernel: Where h is the smoothing parameter; Step 102: remove the trend term from the original traffic data sequence X to obtain the detrended data Z = [z1, z2, ..., z N ], where each element is: z n =x n -y n Perform Fourier spectrum analysis on the detrended data: Find the V main frequency components [f1,..,f v ,...,f V ], each main frequency component represents a periodic feature of the data series, and these main frequency components are used to determine the window size of seasonal smoothing Indicates rounding up, thereby extracting V seasonal components representing different periodic characteristics: Wherein, Loess represents the Loess smoothing estimation function; For seasonal ingredients Summing, we get: Finally, we get the seasonal term sequence S = [s1,...,s n ,...,s N ]; Step 103, remove the trend term and seasonal term components from the original flow data sequence, and the remaining part is the random term sequence R = [r1, ..., r n ,...,r N ].
3. The method for generating communication network traffic data based on mechanism knowledge and generative model according to claim 2 is characterized in that: The specific method of step 2 is: Step 201, constructing a data generation model based on CcGAN, wherein both the generator and the discriminator adopt an LSTM network structure, the generator receives conditional labels and realizes the combination with the conditional labels through layer normalization, and in the discriminator, the labels are embedded into the input of the discriminator through label projection to distinguish the input data corresponding to different conditional labels; Step 202, taking the seasonal item sequence S as the learning object of the data generation model, dividing the sequence S into M data segments of length L, M=N / L, a single data segment as the input of the model, L is a hyperparameter; the continuity condition of the model is the label corresponding to each data segment, and the label is determined by the time period of the data, that is, the combination of indexes corresponding to each seasonal item cycle; Step 203, setting the loss functions of the generator and the discriminator. The loss function of the discriminator includes two types: a hard neighborhood discriminator loss function and a soft neighborhood discriminator loss function. Among them, the hard neighborhood discriminator loss function is: Among them, C1 and C2 are constants respectively; represents the i-th real data sample, represents the i-th data sample generated by the generator; q j represents the conditional label of the jth data sample; N r is the number of true samples that meet the conditions in the neighborhood of the specified conditional label, N g is the number of generated samples; ∈ r and ∈ g represents the Gaussian noise added to the label; D is the discriminator; is a Gaussian distribution; E is the expected function; W1 and W2 represent the weight coefficients of the samples respectively: Among them, k is the neighborhood radius; Refers to satisfying |qq j -∈ r |≤kq j The number of Refers to satisfying |qq j -∈ g |≤kq j The number of A is a knowledge function supported by subscripts, which is used to indicate whether an element belongs to a set. If element x belongs to set A, then 1 A (x)=1, otherwise 1 A (x) = 0; The soft neighborhood discriminator loss function is: Among them, C3, C4 are constants, W3, W4 represent the weight coefficients of the samples: Among them, the ω function is defined as: μ is a hyperparameter, μ>0; The generator loss function is: Among them, G represents the generator; Step 204, the generator and the discriminator are trained in an adversarial manner, the goal of the generator is to maximize the error probability of the discriminator, and the goal of the discriminator is to minimize the classification error; the specific training process is as follows: 1) Split the seasonal item sequence S into sequence segments of length L in chronological order, and combine their corresponding continuous labels as the input of the data generation model to form the training data set of the network; 2) Using continuous labels as control conditions, randomly generated noise is input into the generator to generate data segments corresponding to the labels; 3) The seasonal item data fragment generated by the generator Actual seasonal data snippet And the control condition q i Input into the discriminator together; 4) Select the loss function of the discriminator as Or a weighted function of the two, calculate the gradient of the discriminator loss function, and update the discriminator parameters through back propagation; 5) Calculate the generator loss function and update the generator parameters through back propagation; 6) Repeat steps 2) to 5) until the loss functions of the generator and discriminator stabilize and reach a balance.
4. The method for generating communication network traffic data based on mechanism knowledge and generative model according to claim 3 is characterized in that: The specific method of step 3 is: Step 301, use the trained generator G to perform online data augmentation, given a continuous label, input random noise, and obtain the data segment corresponding to the label; Step 302: for the data fragments at M different positions generated by the generator, use the symbol Represents the mth data segment generated by the generator, and its corresponding label is q m , splice M data fragments into a data sequence according to the position of the label Step 303: convert the data sequence S r The required flow data is synthesized by superimposing the trend item sequence and random item sequence X r Each element The calculation method is: