An Irregular Periodic Flow Prediction Method and Device

By extracting multi-scale and multi-period traffic characteristics and using Transformer model for prediction, the problem of low prediction performance of irregular periodic traffic data in the prior art is solved, and more efficient traffic prediction is achieved.

CN118646664BActive Publication Date: 2025-06-03FIBERHOME TELECOMMUNICATION TECHNOLOGIES CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410862173.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-28
Publication Date
2025-06-03
Estimated Expiration
2044-06-28

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with irregular periodic traffic data, resulting in a degradation in predictive performance in the face of such data.

Method used

By extracting one-dimensional multi-scale flow characteristics and transforming them into multiple periodic two-dimensional flow matrixes, the multi-period time characteristics are extracted, and the multi-feature fusion model of Transformer is used for prediction.

Benefits of technology

This method can accurately capture the multi-scale and periodic characteristics in network traffic data, significantly enhancing the prediction performance of irregular periodic traffic data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118646664B_ABST
    Figure CN118646664B_ABST
Patent Text Reader

Abstract

An irregular periodic traffic prediction method and device, which relate to the fields of computer networks and artificial intelligence technologies. The method includes: extracting one-dimensional multi-scale traffic features according to different time scales; within each scale, transforming the multi-scale traffic features into multiple periodic two-dimensional traffic matrices, and extracting multi-period time features; the multi-period time features include multi-period time features within each scale and multi-period time features across scales; splicing the multi-scale traffic features and the multi-period time features to obtain fused features; inputting the fused features into a multi-feature fusion model based on Transformer to output predicted traffic. The present invention can accurately capture multi-scale and periodic features in network traffic data, thereby enhancing the prediction performance of irregular periodic traffic data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of computer networks and artificial intelligence, and particularly relates to a method and device for predicting irregular periodic traffic. Background Art

[0002] Network Traffic Prediction (NTP), as a typical traffic monitoring and management means, plays an indispensable role in congestion control, resource allocation, and anomaly detection. Traditional traffic prediction methods learn the periodic patterns in data to fit the traffic change trend within a certain period in the future, so as to more accurately predict and plan the resource allocation and performance optimization of the network or system. In recent years, the development of deep learning has greatly promoted the progress of network traffic sequence analysis, and researchers have proposed a large number of deep time series models to model complex network traffic sequences. The current mainstream deep network methods mainly include methods based on Long Short-Term Memory (LSTM) networks, methods based on deep convolutional networks, and methods based on Transformer networks.

[0003] Specifically, Long Short-Term Memory (LSTM) is an improved Recurrent Neural Network (RNN), which is suitable for processing and predicting network traffic sequences with long time intervals and time delays. However, when there are irregular outliers or sudden events, the memory of LSTM will be damaged or disrupted by the information bias introduced by the outliers, so this type of method is not robust to irregular periodic data.

[0004] Convolutional Neural Networks (CNN) can capture local features in time series data, enabling the network to make more effective use of spatial information. Therefore, many methods integrate one-dimensional convolution into LSTM to enhance the local spatial feature extraction ability of the model. However, the combination of convolution operation and LSTM cannot well capture the long-range dependencies in the input sequence, so there are still deficiencies when this type of method predicts irregular periodic data.

[0005] Transformer is a deep learning architecture that has achieved great success in natural language processing and other sequence modeling tasks in recent years. As a model structure based on the multi-head self-attention mechanism, it completely gets rid of the sequential processing limitations of traditional sequence models (such as LSTM and GRU), can process the entire sequence in parallel without being affected by the sequence length, so as to better capture the long-range dependencies in the input sequence. Therefore, Transformer is also widely applied to traffic prediction tasks. However, since traffic data often consists of small-sample datasets, Transformer cannot well extract local periodic features in the case of small samples, so the performance of this type of method is poor in fitting data details when predicting irregular periodic data.

[0006] With the wide application of new mode networks such as the Internet of Things (IoT), Internet of Vehicles (IoV), and 6G, the scale, heterogeneity, and complexity of network traffic are experiencing a rapid growth stage. This surge not only makes network traffic data become huge in scale, but also exacerbates the irregular changes in the traffic data cycle, thus increasing the difficulty of traffic prediction. The above existing prediction methods often struggle to effectively cope with this variability, resulting in a decline in prediction performance when facing irregular periodic traffic data. Summary of the Invention

[0007] This application provides an irregular periodic traffic prediction method and device, which can solve the problem of the decline in prediction performance in the prior art when facing irregular periodic traffic data.

[0008] In a first aspect, an embodiment of this application provides an irregular periodic traffic prediction method, and the method includes:

[0009] Extract one-dimensional multi-scale traffic features according to different time scales;

[0010] Within each scale, transform the multi-scale traffic features into multiple periodic two-dimensional traffic matrices, and extract multi-period time features; the multi-period time features include multi-period time features within each scale and multi-period time features across scales;

[0011] Concatenate the multi-scale traffic features and the multi-period time features to obtain a fused feature;

[0012] Based on the multi-feature fusion model of Transformer, input the fused feature and output the predicted traffic.

[0013] In combination with the first aspect, in an implementation, the extracting one-dimensional multi-scale traffic features according to different time scales includes:

[0014] Set time scales according to different time spans;

[0015] Based on the original network traffic sequence, extract the traffic features of each time scale in parallel to obtain multi-scale traffic features; the extraction method includes a truncation method or a moving window averaging method, and for the traffic features of each time scale, the traffic data corresponding to the time scale is selected in the order from the back to the front.

[0016] Combined with the first aspect, in one implementation, before extracting one-dimensional multi-scale traffic features according to different time scales, it further includes performing a stationary processing on the original network traffic sequence:

[0017] Adjust the original network traffic sequence according to the calculated mean and standard deviation of the original network traffic sequence to reduce or eliminate the non-stationarity in the original network traffic sequence data;

[0018] The predicted traffic output is the traffic data after de-stationarization, which is consistent with the scale and distribution of the data in the original network traffic sequence.

[0019] Combined with the first aspect, in one implementation, after performing a stationary processing on the original network traffic sequence, through positional encoding, token embedding, and time embedding, the data is transformed from a low-dimensional space to a high-dimensional space; the time embedding is a fixed embedding for the time dimension, or an embedding is directly generated using time-related feature encoding.

[0020] Combined with the first aspect, in one implementation, the transformation of multi-scale traffic features into multiple periodic two-dimensional traffic matrices includes:

[0021] Use the Fast Fourier Transform (FFT) to extract the periodic features of the network traffic sequence at each time scale. By analyzing the frequency composition of the traffic sequence, extract k important periods with the largest amplitudes, and intercept the one-dimensional network traffic sequence into k periodic two-dimensional traffic matrices, where k is an integer greater than 1.

[0022] Combined with the first aspect, in one implementation, extracting multi-period time features includes:

[0023] Pass the two-dimensional traffic matrix through a two-dimensional convolutional Inception network to obtain k adjusted one-dimensional periodic time features;

[0024] Obtain the multi-period time features within each scale through an attention mechanism based on FFT amplitude information; and obtain the cross-scale multi-period time features through a cross-scale multi-period attention mechanism.

[0025] In combination with the first aspect, in one implementation, the Transformer-based multi-feature fusion model inputs the fusion features and outputs the predicted traffic, including:

[0026] Introduce the fusion features into the positional encoding, use the Transformer-based multi-feature encoder to learn the complex patterns and irregular dynamic changes of the fusion features, and finally use the fully connected layer to map the high-dimensional fusion features to the predicted traffic.

[0027] In combination with the first aspect, in one implementation, introducing the fusion features into the positional encoding includes:

[0028] Obtain the position information of each data point in the network traffic sequence of the concatenated fusion features,

[0029] Generate the positional encoding by calculating the values of the sine and cosine functions related to each position information and directly add it to the fusion features.

[0030] In combination with the first aspect, in one implementation, using the Transformer-based multi-feature encoder to learn the complex patterns and irregular dynamic changes of the fusion features includes:

[0031] The multi-feature encoder adopts the multi-head attention mechanism to capture the dependencies in the fusion features;

[0032] After the residual connection and normalization steps, further extract and refine the fusion features through the feed-forward neural network.

[0033] In the second aspect, the embodiments of the present application provide a prediction device based on the above-mentioned irregular periodic traffic prediction method, and the prediction device includes:

[0034] A feature extraction module, which includes a multi-scale time block and a multi-period time block. The multi-scale time block is used to extract one-dimensional multi-scale traffic features according to different time scales; the multi-period time block is used to transform the multi-scale traffic features into multiple periodic two-dimensional traffic matrices within each scale and extract multi-period time features; the multi-period time features include multi-period time features within each scale and multi-period time features across scales;

[0035] A splicing module, which is used to splice the multi-scale traffic features and multi-period time features to obtain fusion features;

[0036] A Transformer-based multi-feature fusion model, which is used to input the fusion features and output the predicted traffic.

[0037] The beneficial effects brought by the technical solutions provided by the embodiments of the present application include:

[0038] Extract multi-scale traffic characteristics in different time ranges, and in a two-dimensional space, extract multi-period time characteristics within each scale and multi-period time characteristics across scales. After splicing, the fused characteristics are obtained, and the predicted traffic is output by a multi-feature fusion model based on Transformer. This application can accurately capture multi-scale and periodic characteristics in network traffic data, thereby enhancing the prediction performance of irregular periodic traffic data. Brief Description of the Drawings

[0039] Figure 1 It is a schematic flowchart of the method for predicting irregular periodic traffic in an embodiment of this application;

[0040] Figure 2 It is a logical schematic diagram of the device for predicting irregular periodic traffic in an embodiment of this application;

[0041] Figure 3 It is a schematic diagram for extracting multi-scale traffic characteristics and multi-period time characteristics in an embodiment of this application;

[0042] Figure 4 It is a logical schematic diagram of the multi-feature fusion model based on Transformer in an embodiment of this application. Detailed Embodiment

[0043] In order to enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0044] To make the purpose, technical solution, and advantages of this application clearer, the embodiments of this application will be further described in detail below in conjunction with the drawings.

[0045] In a first aspect, an embodiment of this application provides a method for predicting irregular periodic traffic.

[0046] In one embodiment, with reference to Figure 1 , Figure 1 It is a schematic flowchart of an embodiment of the method for predicting irregular periodic traffic in this application. As Figure 1 shown, the above prediction method includes:

[0047] S1: Extract one-dimensional multi-scale traffic characteristics according to different time scales.

[0048] S2: Within each scale, transform the multi-scale traffic features into multiple periodic two-dimensional traffic matrices, and extract the time features of multiple periods; the time features of multiple periods include the time features of multiple periods within each scale and the time features of multiple periods across scales.

[0049] S3: Concatenate the multi-scale traffic features and the time features of multiple periods to obtain fused features.

[0050] S4: Input the fused features into the multi-feature fusion model based on Transformer and output the predicted traffic.

[0051] In this embodiment, by extracting traffic data of different time scales, one-dimensional multi-scale traffic features are obtained, and the time features of multiple periods within each scale and the time features of multiple periods across scales in the two-dimensional space are obtained. After concatenation, fused features are obtained, and the predicted traffic is output by the multi-feature fusion model based on Transformer. The problem of the decline in prediction performance when facing irregular periodic traffic data is solved.

[0052] The above-mentioned irregular periodic traffic prediction method can be implemented by a prediction device. As Figure 2 shown, a logical schematic diagram of an embodiment of an irregular periodic traffic prediction device is provided. In this embodiment, the prediction device is implemented by a multi-scale and multi-period model, that is, the Mmformer (Multi-Scale and Multi-Period Transformer Model) model, which includes a multi-scale and multi-period time feature extraction algorithm and a multi-feature fusion algorithm based on Transformer.

[0053] Further, in one embodiment, in the above S1, extracting one-dimensional multi-scale traffic features according to different time scales includes:

[0054] Set time scales according to different time spans; parallelly extract traffic features of each time scale based on the original network traffic sequence to obtain multi-scale traffic features, where the original network traffic sequence is a one-dimensional time series.

[0055] Specifically, multi-scale traffic feature extraction can be realized through multi-scale time blocks. As Figure 2As shown in part (b), the multi-scale time block divides the original network traffic sequence into multiple different time spans, and each time span serves as a time scale. In this embodiment, it is dynamically divided into three different time spans: hours, days, and weeks. The one-dimensional traffic features of different time scales are extracted in parallel, and multiple computing units are used to execute tasks simultaneously. In this embodiment, three identical network traffic sequences are input for parallel processing, and each scale runs independently to extract the corresponding scale features. This method not only improves the processing speed but also ensures the richness of capturing the network traffic sequence features at different time scales.

[0056] Considering a given original network traffic sequence of length T, which is a time series, this one-dimensional network traffic sequence can be represented as X 1D ∈R T , where R represents the set of real numbers, indicating that each data point in the network traffic sequence is a real number, T represents the length of the network traffic sequence, that is, the number of time steps, and this length can indicate how many data points are included in the network traffic sequence. Using the multi-scale decomposition method, this original network traffic sequence of length T is decomposed into subsequences of different time scales, corresponding to time spans such as hours, days, and weeks. This decomposition process can be described by the mathematical expression (1):

[0057] X multi-scale ={X s |X s =f s (X 1D ), s∈{s 1 , s 2 , s 3}}(1)

[0058] Among them, X s represents the subset of the network traffic sequence at scale s, and s 1 , s 2 , s 3 respectively represent three different selected time scales (hours, days, weeks). f s (·) is responsible for mapping X 1D to X s , which can be achieved through specific operations such as truncation, moving window averaging, or other aggregation methods. In this embodiment, the truncation method is adopted, which can capture the specific dynamic changes and patterns at each scale more comprehensively. Assuming there are 2 days of data, and the three scales of 2 days, 1 day, and 12 hours will be used for processing, then the three scales will take three different scale data from the back to the front in the 2-day data. Among them, the 1-day scale takes the whole day's data of the second day, and the 12-hour scale takes the data from 12:00 to 24:00 of the second day.

[0059] Further, in one embodiment, as Figure 2 shown in part (a) of Figure 2 , before the original traffic sequence enters the multi-scale time block, it also includes a process of embedding through an embedding layer. The Mmformer model combines different types of embedding techniques such as positional encoding (PositionalEmbedding), token embedding (Token Embedding), and temporal embedding (Temporal Embedding) to achieve the conversion of data from a low-dimensional space to a high-dimensional space and enrich the context information of the data. In this embodiment, the high-dimensional space refers to the dimensionality of the data space. For example, the space dimensionality of network traffic data is 50, and after embedding, the data space dimensionality becomes 200, and at the same time, the meaning of the space expression is also different, which is called the conversion from a low-dimensional space to a high-dimensional space. The aforementioned one-dimensional network traffic sequence refers to the dimensionality of the data format. When one-dimensional data becomes two-dimensional data, such as the length 50 becoming 5×10, only the data format changes, and its space dimensionality does not change and remains in the original data space.

[0060] Specifically, the positional encoding uses a combination of sine and cosine functions to generate the embedding of position information for the Mmformer model to capture the positional relationship of each element in the sequence. In this way, even in the processing of network traffic sequences, the Mmformer model can maintain sensitivity to the element order, which is crucial for understanding and predicting tasks based on network traffic sequences.

[0061] The token embedding uses a one-dimensional convolutional network (Conv1d) to map the input vector at each time point into a high-dimensional space. This process not only increases the representation ability of the Mmformer model but also allows the Mmformer model to capture the local dependencies between adjacent time points. Through the cyclic padding strategy, the Mmformer model can maintain continuity when processing sequence boundaries, thereby effectively using all available context information.

[0062] The temporal embedding (i.e., temporal encoding) can be a fixed embedding (Fixed Embedding) or a time feature-based embedding (Time Feature Embedding) according to the configuration options. The fixed embedding uses a method similar to the positional embedding but is for the time dimension, such as the periodic patterns of time units like hours, days, weeks, etc. The time feature-based embedding directly uses time-related features (such as dates, times, etc.) to generate embeddings, further enhancing the Mmformer model's ability to capture time patterns.

[0063] The combination of the above three embedding methods provides a powerful data representation for the Mmformer model, capturing not only the information at each time point, but also the relationships between time points and the global temporal patterns of the entire network traffic sequence. This rich data representation provides a solid foundation for capturing complex spatio-temporal features, understanding the order and temporal dependencies in the data, and significantly enhancing the representational ability of the Mmformer model. Through such data embedding processing, the Mmformer model can effectively process and analyze network traffic sequence data, improving the accuracy and robustness of predictions.

[0064] Further, in one embodiment, in the above S2, multi-period time blocks can be used to transform the multi-scale traffic features into multiple periodic two-dimensional traffic matrices within each scale, and extract multi-period temporal features, including:

[0065] S21: As Figure 2 shown in part (c) of the figure, use the Fast Fourier Transform (FFT) to extract the periodic features of the network traffic sequence at each time scale, and identify and select important periods by analyzing the frequency composition of the network traffic sequence.

[0066] Specifically, as Figure 3 shown in part (a) of the figure, a network traffic sequence is shown. It can be seen that the network traffic data has periodicity and irregularity. As Figure 3 shown in part (b) of the figure, k important periods with the largest amplitudes are extracted by FFT, where the amplitude reflects the importance of these frequency components in the network traffic sequence. Using FFT to transform the data of the network traffic sequence into the frequency domain representation enables the multi-period time block to identify the frequency components in the sequence.

[0067] After the FFT transformation, the data of the network traffic sequence is transformed from the time domain to the frequency domain. In the frequency domain, each frequency component is a complex number, and these complex numbers contain the amplitude and phase information of the frequency component. To analyze the frequency composition, usually the amplitude information of these frequency components is concerned. The amplitude is the absolute value of the frequency component, indicating the intensity of the signal at the corresponding frequency. By calculating the mean value of the absolute values of the above frequency domain representation, the frequency composition of the network traffic sequence can be analyzed. The amplitude is the absolute value of the frequency component, indicating the intensity of the signal at a specific frequency. Calculating the mean value of the absolute values can understand the average intensity distribution of the signal in the entire frequency domain. The mean value of the absolute values provides an overall concept of the frequency intensity distribution, which can help understand the average intensity of the signal at different frequencies.

[0068] In this embodiment, the calculation of the mean of the absolute values takes into account all batches and channels, thereby generating a list that reflects the average amplitude of each frequency component. To focus on non-DC (non-zero frequency) components, the amplitude corresponding to the DC component in the list is set to 0 because the DC component usually represents the average level of the sequence rather than the periodic changes we are interested in. Next, the k frequency components with the largest average amplitude are selected from the frequency list, and these components represent the most significant periodic patterns in the data. To ensure that the selected frequency components are all valid periods, multiple period values are selected as alternatives, so that the k selected amplitudes and periods have a sufficient number of non-zero frequency components, where k is an integer greater than 1. Then, the actual period length is calculated by dividing the time step of the network traffic data after FFT conversion by the index of each frequency component, and the index of each frequency component corresponds to a specific frequency. Finally, k important periods with the largest amplitudes and their corresponding frequency domain characteristics are obtained, providing important period information for the extraction of multi-period characteristics.

[0069] S22: After selecting the k important periods with the largest amplitudes, the one-dimensional network traffic sequence is intercepted into k periodic two-dimensional traffic matrices according to the k periods.

[0070] As Figure 3 shown in part (c) of, for example, assume that one of the period lengths is P and the length of the sample sequence (network traffic sequence) is T. Then, T / P + 1 vectors can be intercepted, and one of the vectors has a length less than P. By using the padding method to fill it with zeros to adapt to the period length of the two-dimensional traffic matrix. Subsequently, the dimension of the network traffic sequence is adjusted by using the Reshape method, and the padded network traffic sequence is reorganized into a two-dimensional form for further processing. Finally, the size of the two-dimensional traffic matrix is P×(T / P + 1).

[0071] After transforming the periodic two-dimensional traffic matrix, a two-dimensional convolutional Inception network is used to process these network traffic sequences. The Inception network is constructed by stacking two Inception_Block_V1 modules and a GELU activation function. The Inception_Block_V1 module is the classic Inception network architecture, which can effectively extract features at different scales by capturing information using convolutional kernels of different sizes in parallel. The first Inception_Block_V1 module: is used to accept the periodic two-dimensional traffic sequence as input, where the number of input channels is equal to the hidden layer dimension defined in the Inception_Block_V1 module configuration, and the number of output channels is equal to the feedforward layer dimension. The Inception_Block_V1 module is composed of multiple two-dimensional convolutional kernels of different sizes, and the size of each convolutional kernel ranges from 1×1 to (2 * number of convolutional kernels - 1)×(2 * number of convolutional kernels - 1), and so on. The padding of each convolutional layer is set to keep the output size unchanged. This design allows the network to extract features in different receptive fields, thereby capturing information at different scales.

[0072] Using the GELU activation function can enhance the non-linear processing ability, and a [missing part] is connected after the first Inception_Block_V1 module. The input of the second Inception_Block_V1 module is the result processed by the GELU activation function, and at the same time, the number of output channels is restored to the hidden layer dimension. The second Inception_Block_V1 module adopts the same design as the first Inception_Block_V1 module, using multiple two-dimensional convolutional kernels (Conv2d) of different sizes to extract rich feature representations in parallel. The entire Inception network is composed of multiple convolutional blocks, and each convolutional block contains various numbers of convolutional kernels inside, that is, different sizes and numbers of convolutional kernels are used in each convolutional block to process the input data in parallel, so as to be able to extract features at different scales.

[0073] In addition, the Inception_Block_V1 module uses the Kaiming initialization method to initialize the weights of the convolutional layer during initialization to improve the convergence speed and stability of the Inception_Block_V1 module. After these features are processed by the network, a ReshapeBack operation is used to convert the two-dimensional features back into one-dimensional time features. To ensure the consistency of the sequence length and remove the redundant information introduced by padding, the model truncates the extended network traffic sequence, retaining the part with the same length as the original input network traffic sequence, thereby obtaining k adjusted one-dimensional periodic time features. This design enables the Mmformer model to effectively process periodic two-dimensional network traffic sequences, thereby improving the ability to understand the data of network traffic sequences.

[0074] As Figure 2 and Figure 3 shown, a multi-period attention mechanism (FFTAttention) based on FFT amplitude information is introduced in the multi-period time block to strengthen the extraction of multi-period time features within each scale. Specifically, through an attention mechanism based on FFT amplitude information, the multi-period time block is enhanced with the ability to extract features sensitive to frequency features. In this embodiment, this type of module receives the input data x and the corresponding frequency weights weights, where the shape of x is [B, T, N, k], representing the batch size, the length of the network traffic sequence, the feature dimension, and the number of frequency components. The shape of the weights weights is [B, k], reflecting the importance of each frequency component for feature extraction. By applying the softmax function to weights, it is ensured that the sum of the weights of all frequency components is 1, meaning a normalized importance assignment is performed across the entire frequency range. Subsequently, the dimension of weights is expanded to match the shape of the input data x, and the weights are applied to the input data through element-wise multiplication, achieving the weighting of different frequency components. Finally, by summing the weighted data over the frequency dimension, the multi-period time block can comprehensively consider the information of all frequency components, thereby generating a comprehensive feature representation. By introducing this attention mechanism based on FFT amplitude information, the Mmformer model enables the multi-period time block to focus on the more important components in the multi-period features within each scale.

[0075] After that, cross-scale multi-period time features are obtained through a cross-scale multi-period attention mechanism. Specifically, the cross-scale multi-period attention mechanism is used to learn the correlation between multi-period features extracted within different scales, further enhancing the ability to selectively focus on important time features. The cross-scale multi-period attention mechanism first calculates the average value of the outputs of different scales in the time dimension and flattens it in the feature dimension to obtain a compact representation of the output of each scale. These representations are then used to calculate the inner product with a predefined query vector, and after a linear transformation, to generate the attention weights for each scale. Similarly, these weights are normalized by the softmax function. In this way, the cross-scale multi-period attention mechanism assigns different weights to the multi-period outputs of different scales, allowing the Mmformer model to dynamically adjust its focus of attention and preferentially select multi-period features at scales considered to be more important or more informative.

[0076] Furthermore, in the above S3, in the one-dimensional space, the multi-scale traffic features and the multi-period time features are concatenated to obtain a fused feature.

[0077] Furthermore, in the above S4, based on the multi-feature fusion model of Transformer, the fused feature is input, and the predicted traffic is output, specifically including: introducing the fused feature into the positional encoding, using the multi-feature encoder based on Transformer to learn the complex patterns and irregular dynamic changes of the fused feature, and finally using the fully connected layer to map the high-dimensional fused feature to the predicted traffic.

[0078] In one embodiment, as Figure 4 shown, the fused feature is introduced into the positional encoding to provide the specific position information of each data point in the network traffic sequence of the fused feature and ensure that the encoding of each position is unique. This enables the multi-feature fusion model of Transformer to more effectively understand and utilize the sequential dependence of network traffic features. In this embodiment, the multi-feature fusion model of Transformer adopts the strategy of fixed positional encoding. Considering the maximum length of the network traffic sequence and the dimension of the model's hidden layer, the positional encoding is generated by calculating the values of sine and cosine functions related to each position information. The periods of these functions change with the increase of the dimension. In this way, the position information is encoded as a vector with the same dimension as the multi-scale traffic features and the multi-period time features and directly added to the fused feature, thus seamlessly integrating the position information into the model's representation.

[0079] As Figure 4As shown, the above uses a Transformer-based multi-feature encoder to learn complex patterns and irregular dynamic changes in fused features. The multi-feature encoder adopts the multi-head attention mechanism, effectively capturing and understanding the dependencies in the network traffic sequence of fused features, enabling the Mmformer model to simultaneously focus on different parts of the network traffic with multi-scale and multi-period fused features, enhancing the adaptability of the Mmformer model to features of different scales and periods.

[0080] Specifically, the dependencies in the network traffic sequence of fused features refer to the fact that the features at a certain position in the sequence may have a strong correlation with the features at other positions. In network traffic data, the traffic peak at a certain moment may affect the traffic in the following several moments. Through the multi-head attention mechanism, the dependencies between different positions can be noticed in different heads, and each head can focus on different parts of the network traffic sequence of fused features. This mechanism not only improves the Mmformer model's ability to understand complex patterns but also better captures the dynamic changes between different moments. In addition, through residual connections and normalization steps to stabilize and optimize the model performance, the built-in structured hierarchy of the multi-feature encoder further extracts and refines the fused features through a feed-forward neural network, deepening the understanding of the underlying patterns and dynamics of the data.

[0081] Finally, a fully connected layer (linear layer Linear) is used to map the high-dimensional features refined by the multi-feature encoder to the predicted traffic output. The fully connected layer maintains the mapping relationship between the features and the output and ensures that the dimension of the output matches the requirements of the prediction task. At the same time, the conversion efficiency from features to the final predicted value is optimized, improving the accuracy of network traffic prediction.

[0082] Furthermore, in one embodiment, in the above irregular periodic traffic prediction method, before extracting one-dimensional multi-scale traffic features at different time scales, it further includes a step of stabilizing the network traffic sequence, which is implemented before the embedding layer. As Figure 2 shown, at the initial stage of data processing, the series stationary technology is adopted to adjust the data. The purpose is to reduce or eliminate the non-stationarity in the network traffic sequence data. By calculating the mean and standard deviation of the network traffic sequence to adjust the network traffic sequence, the input of the Mmformer model can be optimized to eliminate the trend and seasonal components in the data, making the network traffic sequence more stable and facilitating subsequent processing. Calculate the mean and de-mean through formulas (2) and (3):

[0083]

[0084] x′ t = x t - μ(3)

[0085] Wherein, x t is the observed value at time point t, T is the total length of the network traffic sequence, μ is the mean value of the network traffic sequence, and x′ t is the observed value after mean removal. The standard deviation and standardization are calculated through formulas (4) and (5):

[0086]

[0087]

[0088] Wherein, σ is the standard deviation of the network traffic sequence after mean removal, and x″ t is the standardized form of the network traffic sequence. In this embodiment, due to the addition of the stationary processing, the subsequent steps extract the traffic features of each time scale in parallel according to the original network traffic sequence after the stationary processing.

[0089] Furthermore, after the Mmformer model completes the prediction, the series de-stationarization technology will be adopted to convert the output prediction result (i.e., the predicted traffic) back to the scale and distribution of the data in the original network traffic sequence before the stationary processing. This step is the inverse operation of the stationary processing step. By reapplying the mean value and standard deviation of the network traffic sequence, it ensures that the prediction result can reflect the true characteristics and distribution of the original data, which is crucial for the accuracy and usability in practical applications.

[0090] Specifically, inverse standardization is performed through formula (6), and the mean value is restored through formula (7).

[0091]

[0092]

[0093] Wherein, is the predicted stationary sequence value, is the inverse standardized network traffic sequence value, is the restored mean value.

[0094] Second, as Figure 2 shown, this application provides an embodiment of an irregular periodic traffic prediction device. This device can be implemented by using an Mmformer model, and specifically includes a feature extraction module, a splicing module, and a multi-feature fusion model based on Transformer.

[0095] A feature extraction module, which includes a multi-scale time block and a multi-period time block. The multi-scale time block is used to extract one-dimensional multi-scale traffic features according to different time scales; the multi-period time block is used to transform the multi-scale traffic features into multiple periodic two-dimensional traffic matrices within each scale, and extract multi-period time features; the multi-period time features include multi-period time features within each scale and multi-period time features across scales.

[0096] A splicing module, which is used to splice the multi-scale traffic features and the multi-period time features to obtain fused features.

[0097] A multi-feature fusion model based on Transformer, which is used to input the fused features and output the predicted traffic. Specifically, the multi-feature fusion model based on Transformer also includes a multi-feature encoder, and the multi-feature encoder uses a multi-head attention mechanism to capture the dependencies in the fused features.

[0098] In addition, the multi-feature fusion model of Transformer stabilizes and optimizes the model performance through residual connections and normalization steps. The built-in structured hierarchy of the multi-feature encoder further extracts and refines the fused features through a feed-forward neural network, deepening the understanding of the underlying patterns and dynamics of the data.

[0099] Finally, the multi-feature fusion model of Transformer uses a fully connected layer to map the high-dimensional features refined by the multi-feature encoder to the predicted traffic output.

[0100] Among them, the function implementation of each module in the above traffic prediction device corresponds to each step in the above traffic prediction method embodiment, and its function and implementation process will not be elaborated here one by one.

[0101] To verify the effectiveness and universality of the Mmformer model, this application provides a verification example, which is verified on two real network traffic data sets respectively, including the United Kingdom Academic Network (UKAN) and a core network data set of a certain European city (European).

[0102] The United Kingdom Academic Network traffic data set comes from the Internet traffic data (in bits) of an Internet service provider, and is the aggregated traffic in the backbone of the United Kingdom Academic Network.

[0103] The core network data set of a certain European city comes from the Internet traffic data (in bits) of a private ISP with centers in 11 European cities.

[0104] Before the training phase, the present invention uses a sliding window method to generate samples from time series data, converting the continuous network traffic sequence prediction task into a supervised learning form suitable for the Mmformer model. Subsequently, a fixed dataset partitioning strategy is adopted, that is, the entire dataset is divided into a training set and a test set through a deterministic segmentation method, where the first 80% of the data is allocated for use as the training set, while the remaining 20% of the data is reserved as the test set to ensure that the prediction performance of the Mmformer model on different parts of the dataset can be comprehensively evaluated.

[0105] To ensure the reproducibility of the experiment and the comparability of the results, the experiment was carefully planned and multiple rounds of tests were carried out. Specifically, the training cycle of all comparison models in the experiment was set to 100 epochs to ensure that the model has enough learning cycles to learn complex patterns and irregularities, and to ensure that each model can perform at a stable level. In addition, the learning rate of all models was uniformly set to 0.0001 to balance the learning efficiency and convergence speed of the model, ensuring that the model can effectively avoid getting stuck in local optima prematurely.

[0106] Considering the computational efficiency and stability requirements during model training, the batch size was set to 32. In the selection of the model optimizer, the Adam optimizer was adopted. This optimizer is widely praised for its efficient convergence characteristics and stable performance, which helps to improve the overall efficiency and effect of model training. In particular, the multi-scale and multi-period strategy of the Mmformer model can adapt to different prediction scenarios, and the scale size is dynamically divided according to the specific requirements of the actual prediction task. For example, time scales of 12 hours, 1 day, and 2 days are used to predict 12-hour data.

[0107] To enhance the model's ability to recognize and process periodic features, the Multi-Period Time Blocks were set to 2 layers. In capturing periodic features, the number of periods k was set to 5. The core parameters of the model include the model dimension (d_model) set to 10, which directly affects the model's expressive ability; the number of attention heads (n_heads) set to 8, enabling the model to process features in different subspaces in parallel, improving efficiency and performance; the number of encoder and decoder layers (e_layers and d_layers) were set to 2 and 1 respectively. Such a hierarchical setting helps the model learn complex features in the data while maintaining computational manageability. The feed-forward network dimension (d_ff) was set to 10. The appropriate setting of this parameter can improve the model's ability to process information. In addition, a dropout rate of 0.1 was adopted to prevent overfitting, and the GELU activation function was selected to provide non-linear processing ability, further enhancing the model's performance.

[0108] This embodiment is based on Python 3.11.4 and implements the proposed Mmformer model using Pytorch-GPU 2.0.1+cu118. The Mmformer model is trained on a computer equipped with a GPU, and the specific configuration is as follows: the CPU is an Intel(R) Xeon(R) Gold 6148 CPU @ 2.40GHz, with 40 physical cores (20 cores per socket, 2 sockets in total), supporting 80 threads; the GPU is 4 pieces of Nvidia GeForce RTX 3090; the memory is 502GB; the operating system is 64-bit Ubuntu 18.04.6 LTS; the CUDA version is 11.8.

[0109] Network traffic prediction belongs to the category of network traffic sequence prediction tasks. The proposed Mmformer model is benchmarked against recognized and state-of-the-art models in the field of network traffic sequence prediction, including LSTM (1997), GRU (2014), TCN (2018), Transformer (2017), ETSformer (2022), Pyraformer (2022), Non-stationary Transformer (2022), TimesNet (2023), and iTransformer (2023), a total of nine comparison models. The implementation of all benchmark models is based on the configurations specified in the corresponding original papers or provided by the official code repositories. To ensure the fairness of model comparison, in this embodiment, the input embeddings and output mappings of all models are standardized. In addition, for the variant models of Transformer, the same number of decoder layers is used in this embodiment.

[0110] Through such benchmark comparisons, this embodiment comprehensively evaluates the performance of the Mmformer model in this application in network traffic prediction and compares it with the current state-of-the-art technologies. The comparative analysis is not only based on the prediction accuracy of the models, but also takes into account the computational efficiency and scalability of the models to ensure the practical application value of the proposed Mmformer model. In addition, through unified input-output processing, the consistency and fairness of the comparison between different models are ensured, providing a solid foundation for a deep understanding of the advantages and limitations of different models. Finally, this comprehensive comparative analysis helps to promote the development of network traffic sequence prediction technology, especially in the context of the growing demand for network traffic data prediction.

[0111] To evaluate the prediction accuracy of the Mmformer model, four evaluation metrics were adopted to quantify the mean absolute percentage error (MAPE), root mean square error (RMSE), and mean absolute error (MAE) of the actual network traffic. Specifically, MAPE, RMSE, and MAE are all used to measure the prediction error, and the smaller the value, the better the prediction effect.

[0112] Table 1 shows the comparison results between the Mmformer model and other models. Here, Data represents two datasets, and Len represents the traffic window length. The former represents the historical traffic window length, and the latter represents the predicted traffic window length. The expression "576-144" for Len in Table 1 means using 2 days of data to predict 12 hours of data. Since the dataset is collected every 5 minutes, there are 144 data points in 12 hours, 288 data points in 1 day, and 576 data points in 2 days.

[0113] Table 1

[0114]

[0115]

[0116] Mmformer achieved excellent performance in both experimental datasets. All four evaluation metrics were better than the other 9 models. The best results are marked in bold. The lower the values of the four metrics MAPE, MSE, RMSE, and MAE, the more accurate the prediction results.

[0117] To study the effects of each part of the Mmformer model architecture, in Table 2, ablation studies were conducted on the core modules respectively. Experiments were carried out on the above three core structures: multi-scale time block, multi-period time block, and multi-feature encoder. In the experiment, "Mmformer" is the Mmformer model in this application. In the experiment "Without-Multi-Scale", the multi-scale time block is no longer used. In the experiment "Without-Multi-Period", the multi-period time block is no longer used. In the experiment "Without-MF-Encoder", the multi-feature encoder with multi-head attention mechanism is no longer used. For the fairness and sufficiency of the experiment, the same parameters and data were used for the experiment. To make the training effect more sufficient, 200 epochs were adopted.

[0118] Table 2

[0119]

[0120] The results show that the performance of the Mmformer model significantly depends on its core components. The absence of the Multi-Period Time Blocks leads to the largest performance loss, where the MAPE metric of the UKAN dataset with Len of 864 - 288 changes from 4.372% to 11.0737%, highlighting the importance of multi-period features in improving the prediction accuracy of network traffic sequences. The removal of the Multi-Feature Fusion Module (MF-Encoder) significantly weakens the model's ability to capture the dependencies at each time point in the sequence, where the MAPE metric of the UKAN dataset with Len of 864 - 288 changes from 4.372% to 8.9751%. Although the "Multi-Scale" module has a relatively small effect on performance improvement, its contribution in enabling the model to adapt to different time-scale features cannot be ignored.

[0121] It should be noted that the serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.

[0122] The terms "including" and "having" and any variations thereof in the specification and claims of the present application and the above drawings are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices. The descriptions with terms such as "first", "second", and "third" are used to distinguish different objects, etc., and do not represent a sequential order, nor do they limit that "first", "second", and "third" are of different types.

[0123] In the description of the embodiments of the present application, words such as "exemplary", "for example", or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary", "for example", or "for instance" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary", "for example", or "for instance" is intended to present relevant concepts in a specific manner. Additionally, in the description of the embodiments of the present application, "a plurality of" means two or more than two.

[0124] In some processes described in the embodiments of the present application, there are multiple operations or steps that appear in a specific order. However, it should be understood that these operations or steps may not be executed in the order in which they appear in the embodiments of the present application or may be executed in parallel. The serial numbers of the operations are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. Additionally, these processes may include more or fewer operations, and these operations or steps may be executed in order or in parallel, and these operations or steps may be combined.

[0125] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above and includes several instructions for causing a terminal device to execute the methods described in various embodiments of the present application.

[0126] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A method for predicting irregular periodic flow, characterized in that: The method comprises: Extract one-dimensional multi-scale traffic features at different time scales; In each scale, the multi-scale traffic features are transformed into multiple periodic two-dimensional traffic matrices to extract multi-periodic time features; the multi-periodic time features include multi-periodic time features in each scale and multi-periodic time features across scales; the multi-periodic time features in each scale are obtained through an attention mechanism based on FFT amplitude information; the correlation between multi-periodic features extracted in different scales is learned through a cross-scale multi-periodic attention mechanism; The multi-scale flow features and the multi-period time features are concatenated to obtain fusion features; The Transformer-based multi-feature fusion model inputs the fused features and outputs the predicted traffic, including: introducing the fused features into position encoding, using the Transformer-based multi-feature encoder to learn the complex patterns and irregular dynamic changes of the fused features, and finally using a fully connected layer to map the high-dimensional fused features to the predicted traffic.

2. The irregular periodic flow prediction method according to claim 1, characterized in that: The one-dimensional multi-scale flow characteristics are extracted according to different time scales, including: Set time scales according to different time spans; Based on the original network traffic sequence, the traffic features of each time scale are extracted in parallel to obtain multi-scale traffic features; the extraction method includes a truncation method or a sliding window averaging method, and the traffic features of each time scale are selected in the order from back to front in time, and the traffic data of the corresponding time scale is selected.

3. The irregular periodic flow prediction method according to claim 2, characterized in that: Before extracting one-dimensional multi-scale traffic features at different time scales, the original network traffic sequence is stabilized: The original d network traffic sequence is adjusted according to the calculated mean and standard deviation of the original network traffic sequence to reduce or eliminate the non-stationarity in the original network traffic sequence data; The output predicted traffic is the traffic data after destabilization, which is consistent with the scale and distribution of the data in the original network traffic sequence.

4. The irregular periodic flow prediction method according to claim 3, characterized in that: After the original network traffic sequence is stabilized, the data is converted from low-dimensional space to high-dimensional space through position encoding, token embedding and time embedding; the time embedding is a fixed embedding for the time dimension, or directly generates embedding using time-related feature encoding.

5. The irregular periodic flow prediction method according to claim 1, characterized in that: The method of converting the multi-scale traffic characteristics into a plurality of periodic two-dimensional traffic matrices includes: Fast Fourier transform (FFT) is used to extract the periodic characteristics of network traffic sequences on various time scales. By analyzing the frequency composition of the traffic sequence, k important periods with the largest amplitudes are extracted, and the one-dimensional network traffic sequence is truncated into k periodic two-dimensional traffic matrices, where k is an integer greater than 1.

6. The irregular periodic flow prediction method according to claim 1, characterized in that: Introducing fusion features into position encoding, including: Obtain the location information of each data point in the network traffic sequence after splicing and fusion of features, The position code is generated by calculating the values ​​of the sine and cosine functions associated with each position information and directly added to the fused feature.

7. The irregular periodic flow prediction method according to claim 1, characterized in that: Use a Transformer-based multi-feature encoder to learn complex patterns and irregular dynamic changes in fused features, including: The multi-feature encoder uses a multi-head attention mechanism to capture the dependencies in the fused features; After residual connection and normalization steps, the fused features are further extracted and refined through a feed-forward neural network.

8. A prediction device based on the irregular periodic flow prediction method according to any one of claims 1 to 7, characterized in that: The prediction device comprises: The feature extraction module includes a multi-scale time block and a multi-cycle time block. The multi-scale time block is used to extract one-dimensional multi-scale traffic features according to different time scales; the multi-cycle time block is used to convert the multi-scale traffic features into multiple periodic two-dimensional traffic matrices within each scale to extract multi-cycle time features; the multi-cycle time features include multi-cycle time features within each scale and multi-cycle time features across scales; A splicing module, which is used to splice the multi-scale flow features and multi-period time features to obtain fusion features; The Transformer-based multi-feature fusion model is used to input the fusion features and output the predicted traffic.

Citation Information

Patent Citations

  • Long-term network traffic prediction method based on deep learning

    CN113316163A

  • Base station flow prediction method based on space-time convolution and multiple time scales

    CN115134816A

  • Base station flow prediction system and method based on time domain, frequency domain and space domain combined characteristics

    CN117640414A