Non-stationary network flow prediction method and system based on trend variance normalization
By processing non-stationary network traffic data based on trend variance normalization, using multi-head attention model and perceptron network prediction feature parameters, the distribution offset problem caused by non-stationary is solved, and the accuracy and real-timeness of network traffic prediction are improved.
Patent Information
- Application Number
- CN202510504275.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-25
AI Technical Summary
The prior art is difficult to effectively process non-stationary network traffic data, resulting in poor performance in real-time and accuracy of network traffic prediction models, especially in complex network environments, which are difficult to adapt to the distribution offset caused by non-stationary.
The non-stationary network traffic data is processed through multi-scale slices and local normalization, combined with the multi-head attention model and perceptron network prediction feature parameters, the Fourier transform selects the main period for slices, and performs anti-normalization and weighting operations through the backbone model to build a non-stationary network traffic prediction model.
It improves the accuracy and robustness of network traffic prediction, reduces the modeling difficulties caused by non-stationarity, and enhances the generalization and real-time prediction capabilities of the model.
Smart Images

Figure CN120378318A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of communication technologies, and relates to a non-stationary network traffic prediction based on trend variance normalization, which can be used for network optimization, routing design, and load balancing design. Background Art
[0002] Due to the huge scale, heterogeneous and complex architecture, and rich dynamic characteristics of the Internet, there are increasingly rich service types in the network. In order to more effectively perform network optimization, routing design, and load balancing design, so as to ensure and improve the service quality of the existing network, modeling and predicting network traffic is a key technology in research.
[0003] Network traffic prediction is essentially time series prediction. Traditional methods usually assume that the time series is stationary when designed, that is, the distribution of the time series remains stable over time, and there needs to be a constant conditional dependence between future predictions and historical observations. However, real-world network traffic data is usually affected by various complex factors such as sampling accuracy, the number of users, and user behavior. Therefore, most network traffic is non-stationary, and the data distribution changes significantly. It is precisely because of the essential characteristics of traffic data that violate the premise of assuming independent and identically distributed data in machine learning, and at the same time are contrary to the basic assumptions of existing deep models and statistical models. Therefore, these methods have poor generalization ability when facing future unknown network traffic, especially perform poorly in real-time network traffic prediction and cannot adapt to this "non-stationary" environment.
[0004] To solve the non-stationary problem brought by network traffic sequences, Zhiding Li et al. proposed a lightweight adaptive normalization model, which attempts to perform adaptive normalization on the input sequence on a single fine-grained slice by calculating the mean and variance within the slice and using traditional normalization methods. Compared with the global normalization level, this slice method can better alleviate the distribution shift problem within the time series. However, this method has a single choice for the slice length and does not adaptively select it in combination with data characteristics, and the simple normalization mechanism makes it unable to effectively eliminate the impact brought by non-stationarity.
[0005] The patent document with the application number CN202410143511.2 discloses a network traffic prediction method based on an improved LSTM model with VMD, which includes a modal decomposition VMD module and a data prediction module GA-LSTM. This model decomposes traffic data into modes with different frequencies through VMD, inputs the decomposed sequences into the improved LSTM prediction module to extract feature components and make predictions, performs secondary decomposition on the decomposed residuals, and repeats this operation to extract components with different frequencies of the network traffic signal, ensuring that the data is approximately stationary within a local range, thereby weakening the non-stationarity of the overall signal and improving the effect of the subsequent prediction model. However, due to the highly dependent on parameter selection for the decomposition results of this method, and VMD involves multiple fast Fourier transforms and inverse Fourier transforms, the model has poor interpretability and high computational complexity. For applications in real-time prediction or large-scale network environments such as cloud data centers and CDNs, the computational efficiency of VMD may become a bottleneck, resulting in a decline in real-time performance and response speed. Summary of the Invention
[0006] The object of the present invention is to propose a network traffic prediction method and system based on trend variance normalization in view of the above deficiencies of the prior art, so as to eliminate the distribution shift caused by the non-stationarity of network traffic, reduce the distribution difference of data, make the data more stationary, and improve the accuracy and generalization ability of the network traffic prediction model.
[0007] To achieve the above object, the technical solution of the present invention includes the following:
[0008] 1. A network traffic prediction method based on trend variance normalization, characterized by including:
[0009] (1) Collect a non-stationary network traffic dataset X = {x 1 , x 2 ,..., x n ,..., x N}, and divide it into a training set X1 and a test set X2, where x n is a non-stationary network traffic sequence and N is the total number of sequences;
[0010] (2) Construct a non-stationary network traffic prediction model including a preprocessing module, a parameter predictor module, and a network traffic prediction module:
[0011] The preprocessing module includes multi-scale slicing and normalization operations for processing the non-stationary network traffic sequences in X into stationary network traffic sequences at different scales;
[0012] The parameter predictor module includes a slope predictor, an intercept predictor, and a variance predictor, all composed of a multi-head attention model and a perceptron network MLP, for predicting the characteristic parameters of non-stationary network traffic sequences;
[0013] The network traffic prediction module includes any backbone model, denormalization, and weighted sum operations for obtaining the final non-stationary network traffic prediction sequence;
[0014] (3) Input the training set X1 into the non-stationary network traffic prediction model and train it through a two-stage training architecture to obtain a trained prediction model;
[0015] (4) Input the test set X2 into the trained non-stationary network traffic prediction model, and its output result is the predicted non-stationary network traffic sequence.
[0016] Furthermore, the preprocessing module processes the non-stationary network traffic sequences in X into stationary network traffic sequences at different scales, and its implementation includes the following:
[0017] 2a) Perform a fast Fourier transform on the input non-stationary network traffic sequence to convert x n from the time domain to the frequency domain, and select the first K periods corresponding to the larger amplitudes in the frequency domain as the main periods of the network traffic
[0018] 2b) Use the main period as the slice length to perform non-overlapping slicing on the network traffic sequence x n to obtain K slice data For sequences that are not long enough for slicing, complete them by copying a section from the end, where is the i-th slice data obtained with T i as the slice length, and is the j-th sub-slice in x i ;
[0019] 2c) Perform linear fitting calculations on each sub-slice in x i to obtain the corresponding slope and intercept, calculate the variance within the slice based on these two parameters, and perform local normalization operations within the sub-slice to obtain a stationary network traffic sequence
[0020] Furthermore, the network traffic prediction module obtains the final non-stationary network traffic prediction sequence, and its implementation includes the following:
[0021] 2d) Use the stationary network traffic sequence as the input, obtain the stationary network traffic prediction sequence by the backbone model, and perform denormalization on it within the local slice using the characteristic parameters output by the parameter predictor module;
[0022] 2e) Calculate the amplitudes of the K main periods using the softmax function to obtain the weight information corresponding to K scales;
[0023] 2f) Weight the denormalized non-stationary network traffic prediction sequence with its corresponding weight information to obtain the non-stationary network traffic prediction sequence finally output by this network traffic prediction module
[0024] 2. A non-stationary network traffic prediction system based on trend variance normalization, characterized by including:
[0025] An acquisition module, configured to acquire a non-stationary network traffic data set X and divide it into a training set X1 and a test set X2;
[0026] A model construction module, configured to construct a non-stationary network traffic prediction model including a preprocessing module, a parameter predictor module, and a network traffic prediction module:
[0027] The preprocessing module is configured to process the non-stationary network traffic sequence in the data set X into a stationary network traffic sequence at different scales;
[0028] The parameter predictor module is configured to predict the characteristic parameters of the non-stationary network traffic sequence;
[0029] The network traffic prediction module is configured to perform denormalization, weighted sum operations on the stationary network traffic sequence using the characteristic parameters to obtain a non-stationary network traffic sequence;
[0030] A training module, configured to train the constructed non-stationary network traffic prediction model using the training set X1;
[0031] A test module, configured to test the test set X2 using the trained non-stationary network traffic prediction model to finally obtain the predicted value of the non-stationary network traffic sequence.
[0032] Compared with the prior art, the present invention has the following advantages:
[0033] 1. The preprocessing module in the non-stationary network traffic prediction model constructed by the present invention uses a simple linear equation to fit local slices, calculates the corresponding variance using the slope and intercept of this linear equation, and realizes the operation of local stationarization around the linear trend. Compared with the existing adaptive normalization model, the present invention can more effectively reduce the modeling difficulty caused by non-stationary factors in the face of the non-stationarity of network traffic, effectively control the fluctuation of data around the trend, and improve the prediction accuracy and robustness of the model.
[0034] 2. In the preprocessing module of the present invention, the slice length is adaptively selected according to the characteristics of network traffic data. The Fourier transform is used to select the main period of the traffic sequence in the frequency domain as the slice length, so that the traffic patterns within each slice are similar, and the data distribution difference between slices is smaller than when other slice lengths are selected. This is more conducive to the subsequent normalization operation to eliminate non-stationary information and make the data more stationary. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is the implementation flowchart of the non-stationary network traffic prediction method provided in Embodiment 1 of the present invention;
[0036] Figure 2 is the structural diagram of the parameter predictor module in Embodiment 1 of the present invention;
[0037] Figure 3 is the schematic structural diagram of the non-stationary network traffic prediction model constructed in Embodiment 1 of the present invention:
[0038] Figure 4 is the structural block diagram of the non-stationary network traffic prediction system provided in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0040] Embodiment 1, Non-stationary Network Traffic Prediction Method Based on Trend Variance Normalization
[0041] Refer to Figure 1 , the implementation steps of this example include the following:
[0042] Step 1: Obtain the non-stationary network traffic data set X, perform data preprocessing on it, and divide it into a training set X1 and a test set X2.
[0043] (1.1) Collect the non-stationary network traffic data set X of the target network from the network traffic monitor device;
[0044] (1.2) Set the size L of each sample and the corresponding label size H, that is, predict the network traffic sequence of the next H moments according to the network traffic sequence of L historical moments where is the sample value corresponding to the t1-th moment in the L historical moments of x for x n and is y n the label value corresponding to the t2-th moment among the H moments of
[0045] (1.3) Divide the non-stationary network traffic dataset X into multiple samples and labels according to the set L and H, and use the first 80% of the samples and their corresponding labels as the training set X1, and the remaining 20% of the samples and their corresponding labels as the test set X2.
[0046] Step 2: Construct a preprocessing module that performs the following operations.
[0047] (2.1) For the non-stationary network traffic sequence x input to the model n use the fast Fourier transform FFT to convert it from the time domain to the frequency domain, and select the first K larger amplitude values {A1, A2,..., A K} corresponding frequencies as the main frequencies {f1, f2,..., f K}, and use the formula to calculate the K main periods {T1, T2,..., T K} of the sequence;
[0048] (2.2) Let the K main periods be the slice lengths, and perform non-overlapping slicing on the non-stationary network traffic sequence x n at K different scales to obtain K slice data For sequences that are not long enough for slicing, complete them by copying a section from the end, where is the i-th slice data obtained with T i as the slice length, is the j-th sub-slice in x i , and x i is cut into sub-slices, and the sequence length of each sub-slice is T i ;
[0049] (2.3) Assume that each data point i in each sub-slice of the slice data x follows a simple linear equation Calculate the slope corresponding to this linear equation by first taking the difference and then summing: Then combine the slopes of each sub-slice to obtain the slope parameter corresponding to this slice data x i expressed as: where is the data corresponding to t = T i , is the data corresponding to t = 1;
[0050] (2.4) For the sub - slices at \(t = 1,2,\cdots,T\) i Obtain an equation according to the single - linear equation Sum both sides of this equation:
[0051] (2.5) According to the equation after summation and the obtained slope Obtain the intercept corresponding to this linear equation Combine the intercepts of each sub - slice to obtain the intercept parameter representation of the slice data \(x\) i Corresponding intercept parameter representation
[0052] (2.6) According to the obtained slope Intercept parameter Calculate the sub - slice Corresponding variance parameter
[0053] And combine the variances of each sub - slice to obtain the corresponding variance parameter representation of the slice data \(x\) i Corresponding variance parameter representation
[0054] (2.7) Use the variance parameter of the sub - slice To Perform normalization to obtain the normalization equation: After normalizing all sub - slices and splicing them together, obtain the stabilized network traffic sequence Where \(\varepsilon\) is a very small constant that can stabilize the numerical value and avoid division - by - zero errors.
[0055] (2.8) Establish a new input feature representation: (2.8) Establish a new input feature representation:
[0056] Combine the three feature parameters \(k\) i Of the slice data \(x\) calculated by the pre - processing module i , \(b\) i , \(\sigma\) i In sequence to obtain a new input feature representation, that is:
[0057] Concatenate \(k\) i With \(b\) i , \(\sigma\) i Parameters in dimension to obtain the first group of new input features \([k\) i , \(b\) i and \([k\) i , \(\sigma\) i ;
[0058] Concatenate \(b\) i With \(k\) i , \(\sigma\) iThe parameters are concatenated in dimension to obtain a second set of new input features [b i , k i and [b i , σ i ;
[0059] Concatenate σ i with k i and b i parameters in dimension respectively to obtain a third set of new input features [σ i , k i and [σ i , b i ;
[0060] According to the above new input features, define their respective parameter combinations for different prediction parameters:
[0061] If predicting the slope parameter, define the slope-intercept parameter combination as: X in_1 = [k i , b i , and the slope variance parameter combination as X in_1 ' = [k i , σ i ;
[0062] If predicting the intercept parameter, define the intercept-slope parameter combination as X in_2 = [b i , k i , and the intercept variance parameter combination as X in_2 ' = [b i , σ i ;
[0063] If predicting the variance parameter, define the variance-slope parameter combination as X in_3 = [σ i , k i , and the variance-intercept parameter combination as X in_3 ' = [σ i , b i .
[0064] Step 3: Construct a parameter predictor module.
[0065] Refer to Figure 2 , the implementation of this step includes the following:
[0066] 3.1) Establish a slope predictor including a slope multi-head attention model and a slope perceptron network MLP
[0067] 3.1.1) Establish a slope multi-head attention model:
[0068] The number of heads of the slope multi-head attention model is 2, and the attention model of each head includes a slope linear mapping layer, a slope attention calculation layer, a slope concatenation layer, and a slope output linear layer;
[0069] This linear mapping layer uses different linear transformation matrices W Q , W K , W V to map the slope intercept parameter combination X in_1 and the slope variance parameter combination X in_1 ' into the query Q1, key K1, value V1 matrices of the first head and the query Q2, key K2, value V2 matrices of the second head respectively, which are expressed as follows:
[0070] Q1 = X in_1 *W Q , Q2 = X in_1 '*W Q
[0071] K1 = X in_1 *W K , K2 = X in_1 '*W K
[0072] V1 = X in_1 *W V , V2 = X in_1 '*W V ;
[0073] According to the query Q1, key K1, value V1 matrices of the first head and the query Q2, key K2, value V2 matrices of the second head, the attention scores of the two heads are calculated using the attention mechanism in this slope attention calculation layer:
[0074]
[0075] where head 1k and head 2k represent the attention scores of the first and second attention heads respectively, is the dimension of query Q1, is the dimension of query Q2, and softmax(*) is the normalization exponential function;
[0076] In this slope concatenation layer, the attention scores head 1k and head 2k of the two heads are concatenated, and then the output of the multi-head attention model is obtained through the slope output linear layer:
[0077] output1 = Concat[head 1k , head 2k *Wo
[0078] Among them, output1 is the feature representation of the slope parameter output by the slope multi-head attention model, and W o is the weight parameter of the slope output linear layer;
[0079] 3.1.2) Establish a slope perceptron network MLP:
[0080] The perceptron network MLP has 2 neuron layers, and the neurons between adjacent layers are fully connected. Each neuron uses a non-linear activation function to introduce non-linearity, and the tanh function is selected as the activation function of this network;
[0081] 3.1.3) Connect the slope multi-head attention model and the slope perceptron network MLP to form a slope predictor, that is, use the output output1 of the slope multi-head attention model as the input of the slope perceptron network to obtain the output of this slope predictor:
[0082]
[0083] Among them, is the slope parameter predicted by the slope predictor module, W1 and b1 are the weight and bias of the first neuron layer of the MLP respectively, W2 and b2 are the weight and bias of the second neuron layer of the MLP respectively, and φ(*) is the activation function.
[0084] 3.2) Construct an intercept predictor including an intercept multi-head attention model and an intercept perceptron network MLP.
[0085] 3.2.1) Establish an intercept multi-head attention model:
[0086] The intercept multi-head attention model has 2 heads, and each head's attention model includes an intercept linear mapping layer, an intercept attention calculation layer, an intercept splicing layer, and an intercept output linear layer;
[0087] The intercept linear mapping layer uses different linear transformation matrices W Q , W K , W V to map the intercept slope parameter combination X in_2 , the intercept variance parameter combination X in_2 ' into the query Q1', key K1', value V1' matrices of the first head and the query Q2', key K2', value V2' matrices of the second head respectively, and their expressions are as follows:
[0088] Q1' = X in_2 * W Q , Q2' = X in_2 ' * W Q
[0089] K1' = X in_2 *W K , K2' = X in_2 '*W K
[0090] V1' = X in_2 *W V , V2' = X in_2 '*W V ;
[0091] According to the query Q1', key K1', value V1' matrix of the first head and the query Q2', key K2', value V2' matrix of the second head, in this intercept attention calculation layer, use the attention mechanism to calculate the attention scores of the two heads:
[0092]
[0093] where head 1b and head 2b represent the attention scores of the first and second attention heads respectively, is the dimension of the query matrix Q1′ of the first head, and d k4 is the dimension of the query matrix Q2′ of the second head;
[0094] This intercept concatenation layer is used to concatenate the attention scores head 1b and head 2b ;
[0095] This intercept output linear layer is used to perform a linear transformation on the concatenated attention scores to obtain the output of the intercept multi-head attention model:
[0096] output2 = Concat[head 1b , head 2b *W o
[0097] where output2 is the feature representation of the intercept parameter output by the intercept multi-head attention model, and W o is the weight parameter of the intercept output linear layer;
[0098] 3.2.2) Establish an intercept perceptron network MLP:
[0099] The number of neuron layers of the intercept perceptron network MLP is 2, the neurons between adjacent layers are in a fully connected manner, each neuron uses a non-linear activation function to introduce non-linearity, and the tanh function is selected as the activation function of this network;
[0100] 3.2.3) Connect the intercept multi - head attention model with the intercept perceptron network MLP to form an intercept predictor, that is, take the output output2 of the intercept multi - head attention model as the input of the perceptron network, and obtain the output of this intercept predictor:
[0101]
[0102] Among them, is the intercept parameter predicted by the intercept predictor, W1 and b1 are the weights and biases of the first neuron layer of the MLP respectively, W2 and b2 are the weights and biases of the second neuron layer of the MLP respectively, and φ(*) is the activation function.
[0103] 3.3) Construct a variance predictor including a variance multi - head attention model and a variance perceptron network MLP.
[0104] 3.3.1) Establish a variance multi - head attention model:
[0105] The number of heads of the variance multi - head attention model is 2, and each head's attention model includes a variance linear mapping layer, a variance attention calculation layer, a variance concatenation layer, and a variance output linear layer;
[0106] This variance linear mapping layer uses different linear transformation matrices W Q , W K , W V to map the variance slope parameter combination X in_3 and the variance intercept parameter combination X in_3 ' into the query Q1”, key K1”, value V1” matrices of the first head and the query Q2”, key K2”, value V2” matrices of the second head respectively, which are expressed as follows:
[0107] Q1” = X in_3 *W Q , Q2” = X in_3 '*W Q
[0108] K1” = X in_3 *W K , K2” = X in_3 '*W K
[0109] V1” = X in_3 *W V , V2” = X in_3 '*W V ;
[0110] According to the query Q1”, key K1”, value V1” matrix of the first head and the query Q2”, key K2”, value V2” matrix of the second head, the attention scores of the two heads are calculated using the attention mechanism in this variance attention calculation layer:
[0111]
[0112] where head 1σ and head 2σ represent the attention scores of the first and second attention heads respectively, is the dimension of the query matrix Q1″ of the first head, and d k6 is the dimension of the query matrix Q2″ of the second head;
[0113] This variance concatenation layer concatenates the attention scores head 1σ and head 2σ ;
[0114] This variance output linear layer is used to linearly transform the concatenated attention scores to obtain the output of the variance multi-head attention model:
[0115] output3 = Concat[head 1σ , head 2σ * W o
[0116] where output3 is the feature representation of the variance parameter output by the variance multi-head attention model, and W o is the weight parameter of the variance output linear layer.
[0117] 3.3.2) Establish a variance perceptron network MLP:
[0118] The number of neuron layers of the variance perceptron network MLP is 2, the neurons between adjacent layers are in a fully connected manner, and each neuron uses a non-linear activation function to introduce non-linearity, and the relu function is selected as the activation function;
[0119] 3.3.3) Connect the variance multi-head attention model and the variance perceptron network MLP to form a variance predictor, and use the output output3 of the variance multi-head attention model as the input of the variance perceptron network to obtain the output of this variance predictor:
[0120]
[0121] where, is the variance parameter predicted by the variance predictor, W1 and b1 are the weights and biases of the first neuron layer of the MLP respectively, W2 and b2 are the weights and biases of the second neuron layer of the MLP respectively, and φ(*) is the activation function.
[0122] 3.4) Arrange the slope predictor, intercept predictor, and variance predictor in parallel to form a parameter predictor module.
[0123] Step 4: Construct a network traffic prediction module that performs the following operations.
[0124] 4.1) Use the detrended network traffic sequence as the input, and the detrended network traffic prediction sequence predicted by the backbone model where H is the length of the network traffic prediction sequence and C is the number of dimensions. The backbone models that can be selected include Transformer, Informer, DLinear, etc.;
[0125] 4.2) Use the same slicing method as the preprocessing module to perform non-overlapping slices of length T i to obtain sliced data where is the j-th sub-slice in;
[0126] 4.3) According to the slope intercept and variance parameters of the sub-slice perform local inverse normalization inside the sub-slice : and splice each inverse-normalized sub-slice to obtain a non-stationary network traffic prediction sequence at the T i scale
[0127] 4.4) Calculate the weight information corresponding to K scales using the amplitudes of K main periods: w i = softmax(A i ), where A i is the amplitude corresponding to the main period T i of the non-stationary network traffic sequence in the frequency domain, and w i is the weight information corresponding to the slice length of T i , and softmax(*) is the normalization exponential function;
[0128] 4.5) The inverse-normalized non-stationary network traffic prediction sequence and its corresponding weight information w iPerform a weighted sum to obtain the non-stationary network traffic sequence for the next H time instants, which is the final output of the network traffic prediction module.
[0129]
[0130] Step 5: Construct a non-stationary network prediction traffic model.
[0131] Connect the output end of constructing the new input features in the preprocessing module to the input end of the parameter predictor module, connect the normalized output end in the preprocessing module to the input end of the backbone model in the network traffic prediction module, and connect the output end of the parameter predictor module to the inverse normalization input end in the network traffic prediction module to form a complete non-stationary network prediction traffic model, as Figure 3 shown.
[0132] Step 6: Input the training set X1 into the non-stationary network traffic prediction model and perform iterative training using a two-stage training architecture to obtain a trained traffic prediction model.
[0133] This step is divided into two stages:
[0134] (6.1) The first training stage:
[0135] (6.1.1) Initialize the network parameters of the parameter predictor module and the network traffic prediction module in the non-stationary network traffic prediction model as θ, Set the model learning rate as α and the convergence threshold as γ;
[0136] (6.1.2) Use the mean square loss as the loss function of this traffic prediction model. Obtain the characteristic parameter values corresponding to the non-stationary network traffic prediction sequence output by the parameter predictor module, and calculate its loss value using this characteristic parameter and the characteristic parameter of the true non-stationary network traffic sequence y n :
[0137] (6.1.3) Judge whether the parameter predictor module reaches the convergence state according to the loss value:
[0138] If the loss value of this module is less than the threshold γ, then this module converges, and the optimal network parameters are
[0139] Otherwise, return to step (6.1.2) to continue training.
[0140] (6.2) The second training stage:
[0141] (6.2.1) Use the optimal parameters of the parameter predictor module as the input, and obtain the non-stationary network traffic prediction sequence output by the network traffic prediction module Using the real non-stationary network traffic sequence y n and the prediction sequence to calculate its loss value
[0142]
[0143] (6.2.2) Judging whether the network prediction module reaches the convergence state according to the loss value:
[0144] If the loss value of the module is less than the threshold γ, the optimal network parameter of the module is θ * , and the training is completed;
[0145] Otherwise, return to step (6.2.1) to continue training.
[0146] Step 7: Input the test set X2 into the trained non-stationary network traffic prediction model, and the output result is the non-stationary network traffic sequence predicted by the model.
[0147] Embodiment 2, a non-stationary network traffic prediction system based on trend variance normalization.
[0148] Referring to Figure 4 , this example includes: a collection module 1, a preprocessing module 2, a parameter predictor module 3, a network traffic prediction module 4, a model construction module 5, a training module 6, and a testing module 7. Among them:
[0149] The collection module 1 obtains the non-stationary network traffic dataset X of the target network in the network traffic monitor device, and divides it into a training set X1 and a test set X2;
[0150] The preprocessing module 2 is used to process the non-stationary network traffic sequence in the dataset X into a stationary network traffic sequence at different scales;
[0151] The parameter predictor module 3 is used to predict the characteristic parameters of the non-stationary network traffic sequence;
[0152] The network traffic prediction module 4 is used to perform denormalization and weighted sum operations on the stationary network traffic sequence by using the characteristic parameters to obtain the non-stationary network traffic sequence;
[0153] The model construction module 5 is used to construct a non-stationary network traffic prediction model including a preprocessing module, a parameter predictor module, and a network traffic prediction module:
[0154] The training module 6 is used to iteratively train the non-stationary network traffic prediction model by using the training set X1 to obtain the trained non-stationary network traffic prediction model;
[0155] A test module 7 is used to test the test set X2 by using the trained non-stationary network traffic prediction model, and finally obtain the predicted values of the non-stationary network traffic sequence.
[0156] It should be noted that the step numbers in the specification and claims of the present invention are only for clearly describing the embodiments of the present invention for easy understanding, and their sequence numbers are not limited.
Claims
1. A non-stationary network traffic prediction method based on trend variance normalization, characterized in that, Including: (1) Collect the non-stationary network traffic dataset X and divide it into a training set X1 and a test set X2; (2) Construct a non-stationary network traffic prediction model including a preprocessing module, a parameter predictor module, and a network traffic prediction module: The preprocessing module includes multi-scale slicing and normalization operations for processing the non-stationary network traffic sequences in X into stationary network traffic sequences at different scales; The parameter predictor module includes a slope predictor, an intercept predictor, and a variance predictor, each composed of a multi-head attention model and a perceptron network MLP, for predicting the characteristic parameters of non-stationary network traffic sequences; The network traffic prediction module includes any backbone model, denormalization, and weighted sum operations for obtaining the final non-stationary network traffic prediction sequence; (3) Input the training set X1 into the non-stationary network traffic prediction model and train it through a two-stage training architecture to obtain a trained prediction model; (4) Input the test set X2 into the trained non-stationary network traffic prediction model, and its output result is the predicted non-stationary network traffic sequence.
2. The method according to claim 1, wherein The preprocessing module in step (2), which is used to process the non-stationary network traffic sequences in X into stationary network traffic sequences at different scales, is implemented as follows: 2a) Perform a fast Fourier transform on the input non-stationary network traffic sequence to convert x n from the time domain to the frequency domain, and select the first K periods corresponding to the larger amplitudes in the frequency domain as the main periods of the network traffic 2b) Slice the network traffic sequence x with the main period as the slice length n to obtain K sliced data through non-overlapping slicing For sequences that are not long enough for slicing, copy a segment from the end to complete them, where is the i-th sliced data obtained with T i as the slice length, is the j-th sub-slice in x i ; 2c) For x i Perform a linear fitting calculation on each sub-slice in to obtain the corresponding slope and intercept. Calculate the variance within the slice based on these two parameters, and perform a local normalization operation within the sub-slice to obtain a smoothed network traffic sequence 3. The method according to claim 2, wherein The main period in step 2a) Its calculation formula is as follows: Among them, L is the length of the non-stationary network traffic, and f i is the frequency corresponding to the i-th larger amplitude after the Fourier transform of the sequence, and T i is the i-th period among the K main periods obtained by calculating through L and f i .
4. The method according to claim 2, wherein In step 2c), for each sub-slice in x i perform a linear fitting calculation to obtain the corresponding slope, intercept, and variance within the slice, and perform a local normalization operation within the sub-slice. The implementation includes the following: 2c1) Make the sub - slices of the sliced data x i Each data point in the sub - slices follows a simple linear equation The slope corresponding to this linear equation is obtained by first taking differences and then summing where is the data corresponding to t = T i and is the data corresponding to t = 1; 2c2) For the sub - slices at \(t = 1,2,\cdots,T\) i Following this linear equation gives an equation Sum both sides of the equation: Based on the equation after summation and the obtained slope Obtain the intercept corresponding to this linear equation 2c3) Use the formula and the slope and intercept parameters of the linear equation to obtain the variance parameter corresponding to this sub-slice 2c4) Use the normalization equation Perform local normalization on the inner part of the sub-slice to obtain the stabilized network traffic sequence where ε is a very small constant that can stabilize the value and avoid division-by-zero errors.
5. The method according to claim 1, wherein The multi-head attention model and the perceptron network MLP in the parameter predictor module in step (2), their structures and functions are respectively as follows: The multi-head attention model maps the input into query Q, key K, and value V matrices through different linear transformation matrices W Q , W K , W V and calculates the attention scores of two heads using the self-attention mechanism: After concatenating the attention results of the two heads, the output is obtained, where head i represents the output of the i-th attention head, and d k is the dimension of the query and the key; The perceptron network MLP is composed of multiple neuron layers. The neurons between adjacent layers are fully connected. Each neuron uses a non-linear activation function to introduce non-linearity so that the network can learn complex patterns.
6. The method according to claim 1, wherein The network traffic prediction module in step (2) obtains the final non-stationary network traffic prediction sequence, and its implementation is as follows: 2d) Use the smoothed network traffic sequence as the input, obtain the smoothed network traffic prediction sequence by the backbone model, and perform inverse normalization on it within the local slice using the feature parameters output by the parameter predictor module; 2e) Use the softmax function to calculate the amplitudes of K main periods to obtain the weight information corresponding to K scales; 2f) Weight the denormalized non-stationary network traffic prediction sequence with its corresponding weight information to obtain the non-stationary network traffic prediction sequence finally output by this network traffic prediction module 7. The method according to claim 6, wherein In step 2d), the characteristic parameters output by the parameter predictor module are used to perform denormalization on the stationary network traffic prediction sequence output by the backbone model within the local slice, and its implementation is as follows: 2d1) Take the smoothed network traffic sequence as input, and the predicted smoothed network traffic prediction sequence obtained by using the backbone model where H is the length of the network traffic prediction sequence and C is the number of dimensions; 2d2) Use the slicing method of the preprocessing module with a slice length of T i For perform non-overlapping slicing to obtain sliced data where is the j-th sub-slice in; (2d3) Use the parameter predictor module to separately predict and obtain the slope, intercept, and variance parameters corresponding to each sub-slice in Utilize the formula to perform local denormalization within the sub-slice, obtaining the non-stationary network traffic prediction sequence at scale T i 8. The method according to claim 6, wherein: In step 2e), the softmax function is used to calculate the amplitudes of K main periods to obtain the weight information corresponding to K scales, and its formula is as follows: w i = softmax(A i ) In the formula, is the main period T of the non-stationary network traffic sequence in the frequency domain i The corresponding amplitude, w i is the weight information corresponding to the slice length with T i as the slice length, and softmax(*) is the normalized exponential function; In step 2f), the denormalized non-stationary network traffic prediction sequence is weighted with its corresponding weight information, and its formula is as follows: Wherein, is the non-stationary network traffic prediction sequence after anti-normalization, is the result after weighting the non-stationary network traffic prediction sequences of K scales.
9. The method according to claim 1, characterized in that, In step (3), the training set X1 is input into the non-stationary network traffic prediction model and trained through a two-stage training architecture, and its implementation is as follows: 3a) First-stage training: 3a1) Initialize the network parameters of the parameter predictor module and the network traffic prediction module in the non-stationary network traffic prediction model to θ, Set the model learning rate to α and the convergence threshold to γ; 3a2) Take the mean square loss as the loss function of the traffic prediction model, and obtain the non-stationary network traffic prediction sequence output by the parameter predictor module The corresponding characteristic parameter values, and use the characteristic parameters and the true non-stationary network traffic sequence y n Calculate its loss value using the characteristic parameters of 3a3) Judge whether the parameter predictor module reaches the convergence state according to the loss value: If the loss value of the module is less than the threshold γ, the module converges, and the optimal network parameters are obtained as Otherwise, return to 3a2) and continue training. 3b) Second-stage training: 3b1) Optimal parameters using the parameter predictor module As input, a non-stationary network traffic prediction sequence is output by the network traffic prediction module Using the real non-stationary network traffic sequence y n And the prediction sequence Calculate its loss value 3b2) Judge whether the network prediction module reaches the convergence state according to the loss value: If the loss value of the module is less than the threshold γ, the module converges to obtain the optimal network parameter θ * , and the training is completed; Otherwise, return to 3b1) and continue training.
10. A non-stationary network traffic prediction system based on trend variance normalization, characterized in that, Including: A collection module for collecting the non-stationary network traffic dataset X and dividing it into a training set X1 and a test set X2; A model construction module for constructing a non-stationary network traffic prediction model including a preprocessing module, a parameter predictor module, and a network traffic prediction module: The preprocessing module is used to process the non-stationary network traffic sequence in the dataset X into stationary network traffic sequences at different scales; The parameter predictor module is used to predict the characteristic parameters of the non-stationary network traffic sequence; The network traffic prediction module is used to perform denormalization and weighted sum operations on the stationary network traffic sequence using the characteristic parameters to obtain a non-stationary network traffic sequence; A training module for training the constructed non-stationary network traffic prediction model using the training set X1; A testing module for testing the test set X2 using the trained non-stationary network traffic prediction model to finally obtain the predicted value of the non-stationary network traffic sequence.
Citation Information
Patent Citations
Network traffic prediction method of improved LSTM model based on VMD
CN118041867A