Cluster network flow prediction method based on multi-scale time feature fusion

Through Fourier transform and discrete wavelet transform combined with multi-scale Transformer architecture, the insufficient capture of multi-scale features and long-term dependencies in cluster network traffic is solved, and high-precision and efficient traffic prediction are achieved.

CN120358156APending Publication Date: 2025-07-22XI AN JIAOTONG UNIV
View PDF 0 Cites 9 Cited by

Patent Information

Application Number
CN202510717153.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture the multi-scale time characteristics and long-term dependencies in cluster network traffic, resulting in insufficient accuracy and real-time performance in non-stationary and nonlinear traffic prediction.

Method used

Fourier transform and discrete wavelet transform combined with multi-scale Transformer architecture, through a multi-head self-attention mechanism and a double-layer encoder, global periodic features and local multi-resolution details are captured, and multi-scale time dependencies are fused.

Benefits of technology

It significantly improves the accuracy and real-timeness of cluster network traffic prediction, reduces the computational complexity, and enhances the robustness of burst traffic and abnormal fluctuations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358156A_ABST
    Figure CN120358156A_ABST
Patent Text Reader

Abstract

The invention provides a cluster network flow prediction method based on multi-scale time feature fusion, and belongs to the technical field of computer network flow prediction. The method comprises the following steps: determining a multi-index prediction sequence based on traffic load characteristics of cluster IP instances, and constructing a high-quality time sequence data set; fourier transform and discrete wavelet transform are used for time-frequency feature analysis, and noise filtering and data dimension reduction are completed; projecting sequences of different time granularities to a unified model dimension, performing one-dimensional channel convolution merging, inputting the merged sequences into a time encoder and a cross-channel encoder, and capturing cross-scale long-term time dependence and a coupling relationship between variables; in the loss function design, time domain and frequency domain loss are fused, double-domain error calculation is carried out on a prediction result and a label through Fourier transform, and the robustness of a model to non-stationary fluctuation is enhanced; and through linear layer decoding and reverse normalization processing, the abstract feature is converted into an actual flow prediction value. According to the invention, the precision and reliability of cluster network flow prediction are significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of network traffic prediction, and particularly relates to a method for predicting cluster network traffic based on multi-scale time feature fusion. Background Art

[0002] With the rapid development of cloud computing, big data, and edge computing technologies, the cluster network has become the core infrastructure to support large-scale distributed systems. In data centers, container orchestration platforms (such as Kubernetes), and microservices architectures, the cluster network carries a large amount of inter-node communication traffic, and its dynamic changes directly affect resource scheduling efficiency, quality of service (QoS), and system stability. For example, accurate traffic prediction can trigger resource expansion and contraction strategies in advance to avoid service interruptions caused by sudden traffic; by predicting peak traffic periods, the load balancing algorithm can be optimized to reduce the probability of network congestion; in the scenario of a green data center, traffic prediction can also assist the power management system to achieve energy consumption optimization.

[0003] Cluster network traffic exhibits significant multi-scale dynamic characteristics. In the time dimension, it includes multi-layer time features such as second-level burst traffic (such as microservice call peaks), minute-level periodicity (such as user access tidal effects), and hourly / daily trends (such as differences in traffic patterns on weekdays / weekends). In the space dimension, the inter-node communication traffic is affected by business dependency relationships, load balancing strategies, and fault recovery mechanisms, forming complex spatio-temporal coupling relationships. At the same time, traffic data also exhibits characteristics such as non-stationarity (such as abnormal fluctuations caused by attack traffic), non-linearity (such as exponential trends in traffic growth), and noise interference (such as acquisition device errors), further increasing the difficulty of prediction. In this context, cluster network traffic prediction technology is not only the core link of network management but also the key prerequisite for realizing intelligent resource scheduling and improving system reliability. Accurately capturing the multi-scale time features of traffic and modeling its long-term dependency relationships have become the core challenges in current research.

[0004] In the network traffic prediction scenario, the commonly used prediction models are roughly divided into three types, namely statistical models, machine learning models, and neural network models. In the early stage, the changes and distributions of network traffic were not complex, and researchers usually adopted network traffic prediction methods based on statistical models, mainly including exponential smoothing (ES), autoregressive (AR), autoregressive integrated moving average (ARIMA), Holt-Winters model (HW), etc. The methods based on statistical models characterize the relationship between input and output through polynomial functions. The advantages of these methods are simplicity and ease of use, but they also have some disadvantages, such as the need to manually select model parameters, lack of robustness to outliers and noise, and insufficient adaptability to non-linear and non-stationary time series, etc.

[0005] With the in-depth study of network traffic, researchers have started to study prediction models that can fit non-linear characteristics in view of the non-linear characteristics of network traffic, and have tried to use traditional machine learning methods such as Support Vector Regression (SVR), Random Forest (RF), and Gaussian Process Regression (GPR) to build models. Although the above non-linear prediction methods based on traditional machine learning have achieved good prediction results compared with the prediction methods based on statistical models, these methods are based on certain fixed assumptions and need to preconceive accurate and effective features in advance, which often requires rich artificial experience in this field. The model design is difficult, the generalization ability is relatively poor, and the fitting effect of SVM and SVR models is greatly affected by the kernel function, and the selection of the kernel function can only rely on artificial experience. At present, the methods based on traditional machine learning have been widely replaced by the network traffic prediction methods based on neural networks.

[0006] With the rapid rise of neural networks, people have found that neural network-based models have achieved good results in processing non-linear data. Neural networks have the advantages of non-linear fitting, high robustness, distributed parallel processing, and strong self-adaptability, which has also prompted a large number of researchers to introduce neural network methods into the field of network traffic prediction. According to the main architecture of the model, these methods can be divided into five categories, namely methods based on RNN, CNN, GNN, MLP, and Transformer. In addition to the basic model architecture, researchers usually jointly model different Backbones to predict network traffic, or perform feature enhancement and mutual compensation in different feature aspects to improve the prediction effect of the model. However, these methods still have the contradiction between local feature modeling and global dependence. Their feature fusion methods are single, usually adopting the strategy of reconstructing after hierarchical prediction, and do not achieve the deep fusion of multi-scale features inside the model, making it difficult to meet the network traffic prediction needs of large-scale clusters.

[0007] According to the applicant's retrieval and novelty search, the following several patents related to the present invention in the technical field of network traffic prediction were retrieved, which are respectively:

[0008] 1. CN119583370A, a network traffic prediction method based on graph convolutional neural network.

[0009] 2. CN119728457A, network traffic prediction method, device, electronic device and storage medium.

[0010] Patent 1 provides a network traffic prediction method based on a graph convolutional neural network. The method first selects the network traffic data of a specific time period in a geographical area with spatial correlation with the target area as the target time series data based on the periodic characteristics of the network traffic data to predict the network traffic in the next two hours. Next, the network traffic data is represented as a byte traffic graph, with bytes as nodes and the correlation of byte pairs as edges to form an adjacency matrix, with the original byte values as node features. Afterwards, the byte traffic graph is subjected to multi-level feature extraction using the average pooling layer and the maximum pooling layer to obtain coarse-grained and fine-grained change trends and splice them into a sequence feature graph. Based on the graph convolutional neural network combined with the long short-term memory network and the gated recurrent unit, a graph convolution operation is performed on the sequence feature graph to extract spatiotemporal features, and finally the prediction result is output after linear transformation.

[0011] Patent 2 provides a prediction method based on a graph attention mechanism. The method first determines multiple indicator sequences based on the changes in multiple traffic indicators of multiple network element devices in a preset time period, and each sequence corresponds to a network element device and a traffic indicator. These sequences are input into the spatial feature extraction layer of the network traffic prediction model. This layer splices and prunes the sequences based on the knowledge of the causal relationship between the network element devices and the traffic indicators, and extracts spatial features using the graph attention mechanism to output the spatial feature vector. Next, the spatial feature vector is input into the temporal feature extraction layer to obtain the spatiotemporal feature vector. Finally, the spatiotemporal feature vector is input into the linear mapping layer to output the predicted values of multiple network indicators of multiple network element devices within the prediction period. Among them, the causal relationship knowledge is determined by obtaining the start and end time of the log events of each traffic indicator of the network element device, and constructing a directed weighted graph of causal relationships between network element devices and traffic indicators.

[0012] Patent 1 converts traffic data into a graph structure through joint spatiotemporal modeling, and explicitly models the spatial associations between nodes (such as geographic location, communication dependencies) through GCN, breaking through the limitations of traditional methods that only focus on a single dimension of time or space. However, a graph structure with bytes as nodes may magnify the details of the underlying data and ignore the macro characteristics of traffic, such as the overall throughput of the region. Patent 2 explicitly models causal relationships, uses log event data to build a causal relationship graph for device indicators, explicitly characterizes the dependency between network element devices and traffic indicators, and improves the interpretability of feature correlation. However, the combination of the graph attention mechanism and the time series model increases the parameter scale. In large-scale network scenarios, the time cost of causal graph construction and feature calculation increases significantly, and real-time prediction performance may be limited. Summary of the invention

[0013] To overcome the above-mentioned shortcomings of the prior art, the purpose of the present invention is to provide a cluster network traffic prediction method based on multi-scale time feature fusion, which captures the global periodic features and local multi-resolution details of the traffic through Fourier transform and discrete wavelet transform respectively, constructs feature vectors containing different time scales, and designs a multi-scale perception Transformer architecture. By improving the position encoding and attention weight allocation, the model's ability to model the dependency relationships of different time scales is enhanced, thereby enhancing the accuracy and real-time performance of cluster network traffic prediction.

[0014] To achieve the above purpose, the technical solution adopted by the present invention is:

[0015] A cluster network traffic prediction method based on multi-scale time feature fusion, comprising the following steps:

[0016] S1. Collect and format the original data of the network traffic time series in the actually deployed gateway cluster. Specifically:

[0017] Based on the traffic load capacity and characteristic information of multiple IP instances in the cluster, determine multiple network metric sequences to be predicted, including the number of incoming and outgoing packets, incoming and outgoing bandwidth, total number of connections, number of connections per second, and number of queries per second;

[0018] By analyzing the periodic characteristics of the network traffic and the actual prediction scenario, determine the time range and time granularity of the data set. Since the change of computer traffic depends on people's activity time, the collection period can be selected as one week, and the time granularity can be selected as one minute;

[0019] In the actually deployed cluster, collect the network traffic time series for multi-variable prediction according to the required metrics and format it into a csv file.

[0020] S2. Perform preprocessing and instance normalization on the time series data of the cluster network traffic. Specifically:

[0021] First, perform data grouping. The original data can be grouped and divided according to the unique identifier field SetID or the preset field SetID in the data set to ensure that each group of data is processed independently. Divide the original data set into non-overlapping sub-data sets to obtain the grouped data set G = G1, G2,..., G K . Where G K is the Kth sub-data set, the SetIDs of each sub-data set are different, and K is the number of categories of SetID. This operation realizes data isolation of different network clusters or business units and avoids interference of cross-group data confusion on time series features. All subsequent processing steps are independently executed for a single sub-data set to ensure the homogeneity of the data within the group.

[0022] Secondly, timestamp sorting is performed, and the cleaned data is arranged in ascending order according to the timestamp field to ensure the continuity and temporal logic of the time series.

[0023] Furthermore, for the grouped sub-datasets, there may be multiple records with the same cluster c and timestamp t, which need to be deduplicated and aggregated. Ensure that each (cid, t) corresponds to a unique record and eliminate redundant information caused by collection noise. Define the cluster identification field as cid, and the outbound bandwidth e b Inbound bandwidth b As the core indicator, for the data subset with the same (cid, t) Use mean aggregation to remove duplicates:

[0024]

[0025] After deduplication, the grouped data Perform secondary grouping by cluster c to obtain a cluster-level data set Then sort them in ascending order by the timestamp field to form a strictly ordered time series. Where t1 <t2<…<t N .

[0026] Then, abnormal data filtering is performed, and clusters with too small flow data are filtered through the preset threshold τ, and the outbound k 、inbound k Respectively represent the maximum outbound bandwidth and inbound bandwidth of the kth cluster. When a time series data sequence satisfies max(outbound k ,inbound k )<τ, it is defined as abnormal data and filtered out. After filtering, a valid cluster set is obtained, which reduces the interference of invalid data on model training.

[0027] To unify the analysis cycle, set the target time period ([T start ,T end ]), remove the data points outside this interval. For the time series of cluster c Keep the condition that satisfies t∈[T start ,T end ] records, forming a truncated sequence

[0028] The sequence after the above processing There are timestamps missing, which are repaired by combining linear interpolation and forward filling. Suppose the target timestamp set is (Equidistant sampling), for the missing time point t, linear interpolation is used to fill it. For t prev and t next are respectively defined as the nearest valid timestamps before and after the missing time point t, then they are respectively the actual values of the network metrics at the moments of t prev and t next . For the missing data x t ', the linear interpolation method is used for automatic filling, and it is repaired by linear fitting:

[0029]

[0030] To eliminate the dimensional difference of different bandwidth metrics, reversible instance normalization (InstanceNormalization) is used for processing. The mean μ i and the standard deviation σ i are independently calculated for each data instance i, and normalization is performed:

[0031]

[0032] where x i is the original data instance, T is the time step, γ and β are learnable affine parameters, ∈ is a small value to prevent division by zero, for example, ∈ = 10 -8 , and the statistical parameters μ i , σ i , γ and β in the normalization process are saved to support subsequent inverse normalization operations.

[0033] The reversible instance normalization operation scales the data to a space with zero mean and unit variance while preserving the relative relationship of internal features within a single time point, and realizes the reversible inverse normalization of the prediction result by saving the mean and standard deviation at each step, ensuring the lossless nature of the preprocessing process.

[0034] S3. Based on the Fourier transform and the discrete wavelet transform, the time-domain and frequency-domain features of the time-series data are analyzed, high-frequency noise is filtered out, data compression and dimensionality reduction are performed, and hierarchical analysis of the hidden periodic components, mutation features, and multi-scale fluctuations in the time-series data is realized. The global frequency-domain decomposition of the Fourier transform for stationary signals is the basis for the multi-scale analysis of the wavelet transform, and the latter solves the "frequency aliasing" problem of the former in the processing of non-stationary signals through a localized time-frequency window. In addition, the frequency-domain amplitude / phase of the Fourier transform and the multi-scale components of the discrete wavelet transform form feature orthogonality. The former reveals the frequency composition of the signal, and the latter depicts the dynamic change of the frequency components on the time axis, jointly constructing a time-frequency joint representation space. Specifically:

[0035] Let the preprocessed time-series data be x[n] ∈ RN , where \(n = 0, 1, \ldots, N - 1\), and \(N\) represents the sequence length of the preprocessed time series data. First, the time-domain signal is mapped to the frequency domain through the discrete Fourier transform:

[0036]

[0037] Among them, the amplitude \(|X[k]|\) of \(X[k]\) characterizes the signal energy corresponding to the frequency \(f\) k \(= k / N\), and the phase \(\angle X[k]\) reflects the phase shift of this frequency component. By calculating the power spectral density:

[0038]

[0039] Frequency points with power significantly higher than the noise level are screened to extract the dominant periodic components. For example, the daily cycle corresponds to \(f = 1 / 24\), and the weekly cycle corresponds to \(f = 1 / 168\). After retaining the main frequency components, the low-frequency signal \(x\) FT [n] is reconstructed through the inverse DFT to achieve the global separation of the trend term and the periodic term in the time series data, providing a stationary basic signal for the subsequent wavelet transform.

[0040] For the non-stationary features in \(x[n]\) that are not captured by the Fourier transform, the DWT is used for multi-scale decomposition. Through the binary scale \(a = 2\) j and the integer translation \(b = 2\) j \(k\), the signal is decomposed layer by layer into the approximation component \(c\) j low-frequency trend and the detail component \(d\) j high-frequency fluctuation:

[0041] \(c\) j \(= H * c\) j-1 , \(d\) j \(= G * c\) j-1

[0042] where \(H\) and \(G\) are the low-pass and high-pass filters respectively, and \(*\) represents the convolution operation. The detail component \(d\) j in the \(j\)-th layer corresponds to the frequency interval \([\pi / 2\) j , \(\pi / 2\) j-1 , and the approximation component \(c\) j corresponds to \([0, \pi / 2\) j . Through \(J\) layers of decomposition, a feature sequence \(\{c\) J , \(d\) J , \(d\) J-1 , \ldots, d1\}\) with increasing frequency resolution is obtained, corresponding to multi-scale features from ultra-low frequency to high frequency respectively.

[0043] The joint analysis of the Fourier transform and the DWT forms a "global - local" feature complementarity. The dominant frequency components extracted by the Fourier transform clarify the periodic nature of the time series data. By reconstructing the signal \(x\) FT[n] Provide a low-frequency trend benchmark; the detail component d of DWT j Capture high-frequency fluctuations at different time scales by calculating the energy proportion of each layer Quantify the contributions of minute-level and hour-level fluctuations to the overall traffic, and combine extreme point detection to locate the time positions of burst traffic; splice the dominant frequency amplitude and phase of the Fourier transform with the mean, variance, energy and other statistics of the approximation / detail components of each layer of DWT to form a composite feature vector f = [f FT , f d1 , …, f dJ , f cJ , providing multi-dimensional input for subsequent prediction models.

[0044] S4. Project the multi-scale time embedding representation into the model dimension and merge the multi-scale time series through one-dimensional channel convolution. Let the input data be X in ∈ R C×L , where C is the traffic index dimension and L is the time series length. First, expand the boundary of the preprocessed network traffic time series by repeating and padding the tail data to ensure the integrity of the chunking operation. Divide the expanded sequence into P L non-overlapping time chunks with a fixed chunk length P N and a step size S (the length of the non-overlapping region), and each chunk contains P L time points.

[0045] To retain variable independence and capture local dynamics, perform embedding independently for each variable channel: expand the input sequence X in into a three-dimensional tensor of , and implement chunk embedding through a 1D convolutional layer with a kernel size of P L and an output dimension of :

[0046]

[0047] where represents the single-scale embedding result at scale s, and explicitly models the variable dynamics (such as short-term traffic fluctuations) within the local time window through convolution operations, avoiding information aliasing between variables caused by traditional fully connected embedding.

[0048] To capture dependencies at different time resolutions, use num different chunk lengths and corresponding step sizes {S 1 , S 2 , …, S num} for parallel embedding. Ensure that the number of chunks P N after embedding at each scale is the same by adjusting the step size, that is:

[0049]

[0050] For each scale \(k\), perform block convolution embedding independently to obtain where \(d\) k \(= d\) model / num is the single-scale feature dimension, and \(d\) model is the total target dimension.

[0051] Next, concatenate the embedding results of each scale along the model dimension to form a composite feature vector containing multi-temporal resolution information:

[0052]

[0053] In this way, local features of different scales, such as block lengths capture minute-level fluctuations, capture hourly-level trends, align local features of different scales in the same feature space, and provide a basis for cross-scale interaction for the subsequent self-attention mechanism.

[0054] S5. Input the sequence after multi-scale temporal embedding into the temporal encoder and cross-channel encoder, and efficiently capture multi-scale and cross-scale long-term temporal dependencies through the multi-head self-attention mechanism. Among them:

[0055] The temporal encoder aims to capture long-range dependencies at different temporal resolutions. First, map the features after multi-scale embedding to a subspace through a linear transformation, and obtain after concatenation. Then, use the multi-head self-attention mechanism to extract cross-block temporal dependencies. Specifically, \(X\) can be converted to temp through a trainable linear layer where \(d = d\) model / h. Then, calculate the attention scores to quantify the degree of dependence between different time blocks. Then, the attention output passes through layer normalization and a feed-forward network to generate

[0056] The cross-channel encoder focuses on extracting variable correlation features. The output of the temporal encoder is reshaped and then undergoes 1D convolution to obtain To address the complexity and overfitting problems of high-dimensional variables, the cross-channel encoder adopts a dimensionality reduction attention mechanism. Specifically, first keep the original dimension of the Query matrix \(Q\) chan , and reduce the Key and Value matrices to \(r\times d\) through 1D convolution to obtain \(K\) chan , \(V\) chan \(\in\mathbb{R}\) B×r×d . Then, calculate the lightweight attention scores Reduce the complexity to O(C·r). The attention output passes through layer normalization and a feed-forward network to generate Realize the modeling of the coupling relationship between variables.

[0057] The entire encoding process follows a progressive logic from local to global in the time dimension and from independent to interactive in the channel dimension. The time encoder captures local dynamics first through multi-scale chunking and self-attention, and then integrates them into global time dependencies; the cross-channel encoder, while preserving the independence of variables, explores the causal associations between variables. The two work together to balance computational efficiency and representational ability, effectively solving the modeling problem of time multi-resolution dynamics and variable coupling relationships in network traffic data.

[0058] S6. Use the Fourier transform to convert the label sequence from the time domain to the frequency domain, calculate the time-domain loss and the frequency-domain loss, and fuse them to calculate the final loss. Specifically:

[0059] First, perform feature extraction according to the multi-scale time embedding method in the above process to obtain the feature representation X emb . The feature representation X emb is transformed to the frequency domain through the fast Fourier transform (FFT) to obtain the frequency-domain feature F(Y), and this process can be expressed as where Y t is the value of the time series at time step t, k is the frequency index, and T is the length of the time series. Then, in the frequency domain, independent predictions are made for each frequency component F k This process is implemented by the above method based on the double-layer encoder.

[0060] Next, to ensure that the prediction results also have good performance in the frequency domain, the frequency-domain loss function L freq is used here. This loss function optimizes the model by calculating the difference between the predicted frequency-domain feature and the true frequency-domain feature. The frequency-domain loss function can be expressed as: where is the predicted frequency-domain feature and F(Y) is the true frequency-domain feature.

[0061] To make full use of the information in the time domain and the frequency domain, finally, the final loss is calculated by fusing. Combine the time-domain loss L tmp and the frequency-domain loss L freq to form the final loss function L α , where α is a weight parameter used to control the relative contributions of the time-domain and frequency-domain losses.

[0062] L α = α·L freq +(1 - α)·L tmp

[0063] S7 uses a linear layer to decode time series data, mapping the model dimension to the prediction length; and performs denormalization on the predicted network traffic data.

[0064] In the above steps, the designed double-layer encoder captures the long-term temporal dependence and channel dependence. The input of the decoder is the spatio-temporal feature vector output by the encoder. Where B is the batch size, C is the dimension of the traffic metric, and d model is the model feature dimension. W dec and b dec are the linear layer parameters, which are the weight matrix and the bias vector respectively, and are optimized during training through backpropagation to make the predicted value as close as possible to the true value. The core objective of decoding is to map the high-dimensional features to the prediction length T (such as the time steps in the next two hours), which is specifically achieved through a trainable linear layer:

[0065]

[0066] This operation maps the d model -dimensional features of each metric dimension C to the predicted logarithmic values for T time steps, and the output tensor shape is Through linear transformation, while maintaining the independence of variables, the model uses the cross-scale dependence relationships (such as the fused features of long-term trends and short-term fluctuations) extracted by the encoder to generate multi-step prediction results.

[0067] Finally, since the traffic data was reversibly instance-normalized in the preprocessing stage, the logarithmic values of the prediction output need to be denormalized to restore the original dimension. Let the mean of each time step saved during preprocessing be and the variance be For the prediction time step t′ ∈ [L + 1, L + T], the linear interpolation method is used to estimate its normalization parameters μ t ′ and σ t ′, and the denormalization formula is:

[0068]

[0069] is the final predicted network traffic value, with the unit consistent with the original data before preprocessing.

[0070] On the other hand, the present invention also provides a computer-readable storage medium for storing a computer program, which when executed by a processor causes the computer to execute the method as described above.

[0071] On the other hand, the present invention also provides a computer program product, including computer program instructions, which when run on a computer cause the computer to execute the method as described above.

[0072] Compared with the prior art, the present invention has at least the following beneficial effects:

[0073] Through multi-dimensional time-frequency feature fusion and multi-scale modeling, the present invention effectively solves the problem that traditional methods are insufficient in capturing periodic and mutational features in non-stationary traffic prediction. Based on time-frequency analysis using Fourier transform and discrete wavelet transform, hierarchical analysis of the global periodic components and local multi-resolution details in network traffic is achieved, significantly improving the characterization ability of complex traffic patterns compared with single transformation methods. Through reversible instance normalization and linear interpolation preprocessing, the data quality is guaranteed, providing high-precision input for subsequent feature analysis.

[0074] Aiming at the spatio-temporal coupling characteristics of multi-variable cluster network traffic, the present invention designs a multi-scale time embedding and dual-encoder structure. By merging multi-scale sequences through one-dimensional channel convolution and combining the multi-head self-attention mechanism, long-term time dependencies across scales (such as the association between hourly trends and minute-level fluctuations) and cross-channel variable collaboration (such as the coupling relationship between inbound / outbound bandwidth) are efficiently captured, breaking through the modeling limitations of traditional models for single time series or independent variables, significantly reducing feature redundancy in high-dimensional traffic scenarios, and improving the generalization ability of the prediction model.

[0075] In the design of the loss function, the present invention innovatively fuses time-domain and frequency-domain losses. By using Fourier transform to convert the prediction results and label sequences into the frequency domain for error calculation, the prediction accuracy of the model in the time-frequency dual domain is ensured. This strategy effectively suppresses the influence of noise interference and non-stationary fluctuations, and has stronger robustness especially for scenarios containing bursty traffic or abnormal fluctuations. At the same time, the linear decoding layer combined with inverse normalization processing realizes the accurate recovery of physical quantities of the prediction results on the premise of ensuring computational efficiency, providing a reliable decision-making basis for practical applications such as network resource scheduling and congestion control.

[0076] In summary, through the technical architecture of "deep fusion of time-frequency features - multi-scale modeling - dual-domain loss constraint", the present invention significantly improves the prediction accuracy, robustness, and computational efficiency in cluster network traffic prediction, effectively solving the core difficulties of the prior art in multi-variable non-stationary time series prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] Figure 1 It is a flowchart of the cluster network traffic prediction method based on multi-scale time feature fusion in an embodiment of the present invention.

[0078] Figure 2 It is a schematic structural diagram of the cluster traffic prediction model in an embodiment of the present invention.

[0079] Figure 3This is a schematic diagram of the data change trend based on the preprocessed cluster traffic in the embodiments of the present invention.

[0080] Figure 4 This is a schematic diagram of the cluster traffic prediction trend and the comparison between the predicted value and the real value in the cluster network traffic prediction scenario in the embodiments of the present invention. Detailed implementation manners

[0081] In order to enable those skilled in the art of this technology to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention and are exemplary, aiming to explain the present invention and should not be construed as a limitation to the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0082] As Figure 1 shown, a cluster network traffic prediction method based on multi-scale time feature fusion provided by the embodiments of the present invention includes the following steps:

[0083] S1, based on the traffic load capacity and feature information of multiple IP instances in the cluster, perform raw data collection and formatting.

[0084] The hardware environment is deployed in the gateway cluster of a certain cloud computing center, including several gateway servers, each configured with 10 10Gbps ports and centrally managed by an SDN controller. The data collection module uses Python scripts to regularly obtain traffic data from the relevant interfaces of each server and store it in the HDFS distributed file system. The collection parameters are mainly prediction indicators: 7 network indicators are selected to form a multi-variable sequence, including: outbound packet volume (e pkts ), inbound packet volume (i pkts ), outbound bandwidth (e b , Mbps), inbound bandwidth (i b , Mbps), total number of connections (conn total ), number of connections per second number of queries per second The collection period is 7 days (corresponding to the human activity cycle), covering the traffic fluctuations within a complete week, with a time granularity of 1 minute, forming an equally spaced time series. The data formatting method is to generate a CSV file by aggregating every minute, and the fields include: setid (business unit identifier), cid (cluster identifier), timestamp (timestamp), and each prediction indicator, etc.

[0085] S2. Perform data cleaning on the original data. After preprocessing such as grouping and deduplication, sorting by timestamp, filtering out abnormal data, and linear interpolation, perform reversible instance normalization on the data.

[0086] First, perform data grouping. The original data can be grouped according to the unique identifier field SetID in the dataset or a preset field SetID to ensure independent processing of each group of data. Divide the original dataset into non-overlapping sub-datasets to obtain the grouped dataset G = G1, G2, …, G K . Among them, G K is the Kth sub-dataset, and the SetIDs of each sub-dataset are different. K is the number of categories of SetID. This operation realizes data isolation for different network clusters or business units, avoiding interference with time series features caused by cross-group data confusion. All subsequent processing steps are executed independently for a single sub-dataset to ensure the homogeneity of the data within the group.

[0087] Secondly, perform timestamp sorting. Sort the cleaned data in ascending order according to the timestamp field to ensure the continuity and time series logic of the time series.

[0088] Furthermore, for the grouped sub-datasets, there may be multiple records for the same cluster c and timestamp t, and deduplication and aggregation are required. The aggregated data ensures that each (cid, t) corresponds to a unique record, eliminating redundant information caused by acquisition noise. Define the cluster identification field as cid, and the outbound bandwidth e b and the inbound bandwidth i b as the core metrics. For the data subset with the same (cid, t) use mean aggregation to remove duplicates:

[0089]

[0090] For the deduplicated grouped data perform secondary grouping by cluster c to obtain the cluster-level dataset Subsequently, sort in ascending order according to the timestamp field to form a strictly ordered time series where t1 < t2 < … < t N .

[0091] Subsequently, perform abnormal data filtering. Filter out clusters with too small traffic data through a preset threshold τ. Set outbound k , inbound k to represent the maximum values of the outbound bandwidth and inbound bandwidth of the kth cluster respectively. Then, when a time series data sequence satisfies max(outbound k , inboundk ) At time <t>, it is defined as abnormal data and filtered. After filtering, an effective cluster set is obtained, reducing the interference of invalid data on model training.

[0092] To unify the analysis period, a target time period ([T start , T end ) is set, and data points outside this interval are removed. For the time series of cluster c Keep the records that satisfy t ∈ [T start , T end , forming a truncated sequence

[0093] The sequence after the above processing has missing timestamps and is repaired using a strategy that combines linear interpolation and forward filling. Let the set of target timestamps be (equidistant sampling). For the missing time point t, linear interpolation is used to fill it. Let t prev , t next be defined as the nearest valid timestamps before and after the missing time point t respectively, then are the actual values of the network metrics at times t prev , t next respectively. The missing data x t ′ is automatically filled using linear interpolation and repaired through linear fitting:

[0094]

[0095] To eliminate the dimensional differences of different bandwidth metrics, reversible instance normalization (InstanceNormalization) is used for processing. The mean μ i and standard deviation σ i are calculated independently for each data instance i, and normalization is performed:

[0096]

[0097] where, x i is the original data instance, T is the time step, γ, β are learnable affine parameters, ∈ is a small value to prevent division by zero, for example, ∈ = 10 -8 , and the statistical parameters μ i , σ i , γ and β during the normalization process are saved to support subsequent denormalization operations.

[0098] The reversible instance normalization operation scales the data into a space with zero mean and unit variance while preserving the relative relationships of internal features at a single time point, and realizes the reversible denormalization of the prediction result by saving the mean and standard deviation at each step, ensuring the lossless nature of the preprocessing process.

[0099] S3. Based on the Fourier transform and the discrete wavelet transform, analyze the time-domain and frequency-domain features of the time series data, filter out high-frequency noise, and perform data compression and dimensionality reduction.

[0100] Through joint analysis, the joint analysis of the Fourier transform and the DWT forms a "global-local" feature complementarity. The dominant frequency components extracted by the Fourier transform clarify the periodic nature of the time series data, and the reconstructed signal x FT [n] provides a low-frequency trend benchmark; the detail component d of the DWT j captures high-frequency fluctuations at different time scales, and quantifies the contributions of minute-level and hour-level fluctuations to the overall traffic by calculating the energy proportion E of each layer j at each layer, and locates the time positions of burst traffic by combining extreme point detection; splice the dominant frequency amplitude and phase of the Fourier transform with the mean, variance, energy and other statistics of the approximation / detail components of each layer of the DWT to form a composite feature vector f = [f FT , f d1 , …, f dJ , f cJ that provides multi-dimensional inputs for the subsequent prediction model.

[0101] S4. Project the multi-scale time embedding representation into the model dimension and merge the multi-scale time series through one-dimensional channel convolution.

[0102] Suppose the input data of the embodiment at this time is X in ∈R C×L , where C is the dimension of the traffic index and L is the sequence length. First, expand the sequence by repeating and padding the tail data to ensure the integrity of the block operation. Divide the expanded sequence into P L non-overlapping time blocks with a fixed block length P N and a step size S (the length of the non-overlapping region), and each block contains P L time points.

[0103] To preserve variable independence and capture local dynamics, perform embedding independently for each variable channel: expand X in into a three-dimensional tensor of , and realize block embedding through a one-dimensional convolutional layer with a kernel size of P L and an output dimension of :

[0104]

[0105] wherein represents the single-scale embedding result at scale s, explicitly modeling the variable dynamics within the local time window (such as short-term traffic fluctuations) through convolution operations, and avoiding the information aliasing between variables caused by traditional fully-connected embeddings.

[0106] To capture dependencies at different time resolutions, num different block lengths and corresponding step sizes {S 1 , S 2 , …, S num} are used for parallel embedding. By adjusting the step sizes, the number of blocks P N after embedding at each scale is ensured to be consistent, that is:[[]]

[0107]

[0108] For each scale k, blockwise convolutional embedding is independently performed to obtain where d k = d model / num is the single-scale feature dimension, and d model is the total target dimension.

[0109] Next, the embedding results at each scale are concatenated along the model dimension to form a composite feature vector containing multi-time-resolution information:[[]]

[0110]

[0111] In this way, local features at different scales, such as block lengths capturing minute-level fluctuations,[[]] capturing hourly-level trends, are aligned in the same feature space, providing a basis for cross-scale interaction for the subsequent self-attention mechanism.

[0112] For S5, the sequence after multi-scale time embedding is input to the time encoder and the cross-channel encoder, and the multi-scale and cross-scale long-term time dependencies are efficiently captured through the multi-head self-attention mechanism.

[0113] The time encoder aims to capture long-range dependencies at different time resolutions. First, the features after multi-scale embedding are mapped to a subspace through a linear transformation, and after concatenation, is obtained. Next, the multi-head self-attention mechanism is used to extract cross-block time dependencies. Specifically, first, X temp is converted to through a trainable linear layer, where d = d model / h. Then, the attention scores are calculated to quantify the degree of dependence between different time blocks. Then, the attention output passes through layer normalization and a feed-forward network to generate

[0114] The cross-channel encoder focuses on extracting variable correlation features. The time encoder outputs After reshaping and 1D convolution, it obtains To address the complexity and overfitting problems of high-dimensional variables, the cross-channel encoder adopts a dimensionality reduction attention mechanism. Specifically, first, for the Query matrix Q chan Keep the original dimension, and reduce the Key and Value matrices to r×d through 1D convolution to obtain K chan , V chan ∈R B×r×d . Then, calculate the lightweight attention scores Reduce the complexity to O(C·r). The attention output is passed through layer normalization and a feed-forward network to generate Realize the modeling of the coupling relationship between variables.

[0115] S6. Use the Fourier transform to convert the label sequence from the time domain to the frequency domain, calculate the time-domain loss and the frequency-domain loss, and fuse them to calculate the final loss.

[0116] First, perform feature extraction according to the multi-scale time embedding method of the above embodiment to obtain the feature representation X emb . The feature representation X emb Is converted to the frequency domain through the fast Fourier transform (FFT) to obtain the frequency-domain feature F(Y), and this process can be expressed as Where Y t Is the value of the time series at time step t, k is the frequency index, and T is the length of the time series. Then in the frequency domain, for each frequency component F k Perform independent prediction, and this process is implemented by the above method based on the double-layer encoder.

[0117] Next, to ensure that the prediction results also have good performance in the frequency domain, the frequency-domain loss function L freq Is used here. This loss function optimizes the model by calculating the difference between the predicted frequency-domain feature and the true frequency-domain feature. The frequency-domain loss function can be expressed as: Where Is the predicted frequency-domain feature, and F(Y) is the true frequency-domain feature.

[0118] To make full use of the information in the time domain and the frequency domain, the final loss is calculated by fusion here, and the time-domain loss L tmp And the frequency-domain loss L freq Are combined to form the final loss function L α , where α is a weight parameter used to control the relative contribution of the time-domain and frequency-domain losses.

[0119] L α = α·L freq +(1 - α)·L tmp

[0120] S7 uses a linear layer to decode time - series data, maps the model dimension to the prediction length, and performs inverse normalization on the predicted network traffic data.

[0121] In this embodiment, the input of the decoder is the spatio - temporal feature vector output by the encoder where B is the batch size, C is the dimension of the traffic metric, and d model is the model feature dimension. The core goal of decoding is to map the high - dimensional features to the prediction length T (such as the time steps in the next two hours), which is specifically achieved through a trainable linear layer:

[0122]

[0123] Finally, since the traffic data is reversibly instance - normalized in the pre - processing stage, the logarithm of the predicted output needs to be restored to the original dimension through inverse normalization. Let the mean of each time step saved during pre - processing be the variance be For the prediction time step t′ ∈ [L + 1, L + T], a linear interpolation method is used to estimate its normalization parameters μ t ′ and σ t′ , and the inverse normalization formula is:

[0124]

[0125] is the finally predicted network traffic value, and the unit is the same as the original data before pre - processing.

[0126] The structure of the cluster traffic prediction model in this embodiment is as Figure 2 shown. When the model is trained, the input of the model is the pre - processed cluster network traffic time - series data, which is collected from the actually deployed gateway cluster and covers multiple network metric sequences of multiple IP instances in the cluster. The output of the model is the predicted network traffic value, which is compared with the actual real traffic value to calculate the loss function. Through the backpropagation algorithm, the model updates its parameters according to the value of the loss function, thus continuously optimizing the prediction performance.

[0127] In subsequent prediction scenarios, the input to the model is the time-series data of the new cluster network traffic. These data need to go through the same preprocessing steps as in the training phase. The output of the model is the predicted values of the network traffic for future time periods, and the data dimension is the same as the original input data. These predicted values can provide a decision-making basis for practical applications such as network resource scheduling and congestion control. Since the model has learned the patterns and rules in the data during the training phase, it can accurately predict the future network traffic based on the input data during the inference phase.

[0128] Figure 3 This is a schematic diagram of the change trend of the preprocessed data based on cluster traffic after data preprocessing in this embodiment. The horizontal axis of this figure is time, and the vertical axis is the value of each traffic metric at each moment, including the inbound packet volume (PpsIn), outbound packet volume (PpsOut), inbound bandwidth (BpsIn, Mbps), outbound bandwidth (BpsOut, Mbps), total number of connections (conns), and queries per second (QPS). It can also be extended to other network traffic metrics. By observing its change trend through the time-series diagram, it is not difficult to see that it has a strong periodicity and an obvious daily trend (large traffic during the day and small traffic at night).

[0129] Figure 3 This shows a schematic diagram of the change trend of the preprocessed data based on cluster traffic after data preprocessing in this embodiment. The horizontal axis represents the time dimension and records the timestamps corresponding to each sampling point; the vertical axis covers the values of multiple key traffic metrics, specifically including the inbound packet volume (PpsIn), outbound packet volume (PpsOut), inbound bandwidth (BpsIn, in Mbps), outbound bandwidth (BpsOut, in Mbps), total number of connections (conns), and queries per second (QPS). In addition to these common metrics listed in the figure, this method also has good scalability and can easily be compatible with other types of network traffic metrics to meet the diverse needs in different scenarios.

[0130] By carefully observing this time-series diagram, it can be clearly found that various traffic metrics show significant periodic fluctuation patterns. The daily periodic change trend is particularly obvious: during the day, various business activities are frequent, the number of network users increases, and the demand for data interaction is strong, resulting in high levels of metrics such as inbound and outbound packet volumes, bandwidth occupancy, and query request counts; while as night falls, the network usage frequency decreases, the business activity weakens, and the corresponding traffic metrics also drop significantly. This obvious day-night difference provides important prior feature patterns for the subsequent training of the prediction model, helps the model better understand the internal logic of traffic changes, and thus enables more accurate traffic trend prediction.

[0131] To verify the effectiveness and feasibility of the present invention, this embodiment also conducts tests on a real cluster traffic dataset. Figure 4 It shows the prediction performance of this embodiment in the actual network traffic time series. In the figure, the black curve represents the real network traffic value, while the blue curve corresponds to the traffic value predicted by the model. Specifically, the grey range represents the maximum and minimum values of the traffic change predicted by the model. The blue data on the last day intuitively depicts the trend of the cluster traffic change predicted by the model. It can be seen from the figure that this model has good prediction performance, with a small deviation between the real value and the predicted value, and the prediction result is closely close to the actual observed value. In addition, the model also makes accurate predictions on the trend of network traffic changes, and its trend is basically completely consistent with the actual change trend.

[0132] For this embodiment, the performance of the method can be verified on a public time series dataset. There are 6 commonly used datasets in the current time series field, namely Electricity, ETT, Traffic, Weather, covering the fields of electricity consumption, temperature, traffic, and weather. These datasets are stored in.csv format files. Among them, the ETT dataset has four tables, representing different sampling periods and data sources respectively. Therefore, there are a total of 7 public datasets for this experiment.

[0133] The performance of the model is verified on the above datasets. The specified input sequence length is 64, and the results of time series prediction are shown in Table 1. The evaluation metrics are mainly the MAE (Mean Absolute Error) and MSE (Mean Squared Error) of the prediction results. In complex time series tasks, both MAE and MSE can be reported to comprehensively evaluate the model performance. MAE is the average of the absolute differences between the predicted value and the real value, and the penalty for all errors is linear, with low sensitivity to outliers (such as spikes or outliers). In a time series, if there are periodic fluctuations or sudden outliers, MAE can more stably reflect the overall error level. MSE is the average of the squared errors, and the penalty for larger errors increases quadratically, so it is very sensitive to outliers. In long-term trend prediction, the high penalty of MSE for large errors helps the model capture key fluctuations.

[0134] Table 1

[0135] Dataset ETTh1 ETTh2 ETTm1 ETTm2 Weather Electricity Traffic MSE 0.411 0.316 0.348 0.256 0.222 0.156 0.387 MAE 0.423 0.384 0.375 0.315 0.262 0.159 0.262

[0136] In summary, based on the above solutions provided by the embodiments of the cluster network traffic prediction method based on multi-scale time feature fusion, through multi-dimensional time-frequency feature fusion and multi-scale modeling, the problem that traditional methods are insufficient in capturing periodic and mutational features in non-stationary traffic prediction is effectively solved. Based on the time-frequency analysis of Fourier transform and discrete wavelet transform, the hierarchical analysis of the global periodic components and local multi-resolution details in network traffic is realized. Compared with single transformation methods, the representation ability of complex traffic patterns is significantly improved. In the cluster network traffic prediction, significant improvements in prediction accuracy, robustness and computational efficiency are achieved, effectively solving the core difficulties in multi-variable non-stationary time series prediction in the prior art. The present invention can be applied to network traffic prediction fields such as data centers, cellular networks, and the Internet, and is particularly suitable for cluster network traffic prediction scenarios with multi-variable, non-stationary and multi-scale characteristics.

[0137] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for predicting cluster network traffic based on multi-scale time feature fusion, characterized in that, It includes the following steps: S1. In the actually deployed gateway cluster, based on the traffic load capacity and characteristic information of multiple IP instances in the cluster, collect and format the original data of the network traffic time series; S2. Perform data cleaning on the original data. After the preprocessing process of grouping and removing duplicates, timestamp sorting, abnormal data filtering, and linear interpolation, perform reversible instance normalization on the data; S3. Based on the Fourier transform and discrete wavelet transform, analyze the time-domain and frequency-domain characteristics of the time-series data, filter out high-frequency noise, and perform data compression and dimensionality reduction; S4. Project the multi-scale time embedding representation into the model dimension, and merge the multi-scale time series through one-dimensional channel convolution; S5. Input the sequence after multi-scale time embedding into the time encoder and cross-channel encoder, and efficiently capture the long-term time dependencies of multi-scale and cross-scale through the multi-head self-attention mechanism; S6. Use the Fourier transform to convert the label sequence from the time domain to the frequency domain, calculate the time-domain loss and frequency-domain loss, and fuse and calculate the final loss; S7. Use a linear layer to decode the time-series data and map the model dimension to the prediction length; Perform denormalization on the predicted network traffic data.

2. The method for predicting cluster network traffic based on multi-scale time feature fusion according to claim 1, wherein In S1, the collection and formatting of the original data set in the actually deployed gateway cluster includes: Based on the traffic load capacity and characteristic information of multiple IP instances in the cluster, determine multiple network metric sequences to be predicted, including the number of incoming and outgoing packets, incoming and outgoing bandwidth, total number of connections, number of connections per second, and number of queries per second; Determine the time range and time granularity of the data set through the periodic characteristics of network traffic and the actual prediction scenario; In the actually deployed cluster, collect the network traffic time series for multivariate prediction according to the required metrics and format it into a csv file.

3. The method for predicting cluster network traffic based on multi-scale time feature fusion according to claim 1, characterized in that In S2, the preprocessing and instance normalization of the data include: Group and partition the original data according to the preset field SetID to ensure that each group of data is processed independently. For the original data set D raw , the grouped data set G = G1, G2, …, G K ; G K is the Kth sub-data set, and the SetIDs of each sub-data set are different. K is the number of categories of SetID; Sort the cleaned data in ascending order according to the timestamp field to ensure the continuity and chronological logic of the time series; Filter clusters with too small traffic data through a preset threshold τ, and set outbound k , inbound k respectively represent the maximum outbound bandwidth and inbound bandwidth of the k-th cluster. Then, when a certain time-series data satisfies max(outbound k , inbound k ) < τ, it is defined as abnormal data and filtered; Define t prev and t next as the nearest valid timestamps before and after the missing time point t respectively. Then they are respectively the actual values of the network metrics at time t prev and t next . Automatically fill the missing data x t ' using the method of linear interpolation: Calculate the mean μ independently for each data instance i i and the standard deviation σ i , and perform standardization: where x i is the original data instance, T is the time step, γ and β are learnable affine parameters, ∈ is a small value to prevent division by zero, and μ i , σ i , γ, and β are saved to support subsequent denormalization operations.

4. A method for predicting cluster network traffic based on multi-scale time feature fusion according to claim 1, wherein In S3, based on the Fourier transform and discrete wavelet transform, analyze the time-domain and frequency-domain characteristics of the time-series data to achieve a hierarchical analysis of the implicit periodic components, mutation characteristics, and multi-scale fluctuations in the time-series data. Specifically, it includes: Let the preprocessed time-series data be \(x[n]\in\mathbb{R}\) N , where \(n = 0, 1,\cdots, N - 1\) and \(N\) represents the sequence length of the preprocessed time-series data. First, map the time-domain signal to the frequency domain through the discrete Fourier transform: By calculating the power spectral density Screen the frequency points where the power is significantly higher than the noise level, extract the dominant periodic components, and achieve the global separation of the trend term and the periodic term in the time series data; For the non-stationary features in x[n] that are not captured by the Fourier transform, DWT is used for multi-scale decomposition to decompose the signal layer by layer into an approximation component c j low-frequency trends and detail components d j high-frequency fluctuations: c j = H * c j-1 , d j = G * c j-1 where H and G are low-pass and high-pass filters respectively, * represents the convolution operation, and the detail component d at the j-th layer j corresponds to the frequency interval [π / 2 j , π / 2 j-1 , and the approximation component c j corresponds to [0, π / 2 j . Through J-layer decomposition, a feature sequence {c J , d J , d J-1 , …, d1} with increasing frequency resolution is obtained, corresponding to multi-scale features from ultra-low frequency to high frequency respectively.

5. The method for predicting cluster network traffic based on multi-scale time feature fusion according to claim 1, wherein In S4, project the multi-scale time embedding representation into the model dimension, and merge the multi-scale time series through one-dimensional channel convolution. Specifically, it includes: First, the preprocessed network traffic time series is extended at the boundaries by repeatedly padding the tail data, and the extended sequence is divided into P L non-overlapping time blocks of fixed block length P N and step size S, where each block contains P L time points; Embed each variable channel independently, where the input sequence X in is extended to a three-dimensional tensor, where C represents the traffic index dimension, L represents the time series length, and block embedding is implemented through a 1D convolutional layer with a kernel size of P L and an output dimension of : Adopt num different block lengths and corresponding step sizes {S 1 , S 2 , …, S num} for parallel embedding, and ensure that the number of blocks P N after embedding at each scale is consistent by adjusting the step size, that is: For each scale k, perform block convolution embedding independently to obtain where d k = d model / num is the single-scale feature dimension, d model is the total target dimension, and P N is the number of time blocks; then, concatenate the embedding results of each scale along the model dimension to form a composite feature vector containing multi-time resolution information: After obtaining the local features of different scales, align them in the same feature space to provide a basis for cross-scale interaction for the subsequent self-attention mechanism.

6. The method for predicting cluster network traffic based on multi-scale time feature fusion according to claim 1, wherein, In S5, input the sequence after multi-scale time embedding into the time encoder and cross-channel encoder, and efficiently capture the long-term time dependencies of multi-scale and cross-scale through the multi-head self-attention mechanism. Specifically, it includes: First, the features after multi-scale embedding are mapped to a subspace through a linear transformation, and after concatenation, we get Then, it is converted into through a trainable linear layer, where d = d model / h; then the attention scores are calculated to quantify the dependence between different time blocks; the attention output passes through layer normalization and a feed-forward network to generate The output of the time encoder is obtained after reshaping and 1D convolution The cross-channel encoder adopts a dimensionality reduction attention mechanism; first, for the Query matrix Q chan Keep the original dimension, and reduce the Key and Value matrices to r×d through 1D convolution to obtain K chan , V chan ∈R B×r×d ; then, calculate the lightweight attention score Reduce the complexity to O(C·r), and the attention output goes through layer normalization and a feed-forward network to generate Realize the modeling of the coupling relationship between variables.

7. The method for predicting cluster network traffic based on multi-scale time feature fusion according to claim 1, wherein In S6, use the Fourier transform to convert the label sequence from the time domain to the frequency domain, calculate the time-domain loss and frequency-domain loss, and fuse and calculate the final loss. Specifically, it includes: First, feature extraction is performed according to the multi-scale temporal embedding method in S5 to obtain the feature representation X emb ; the feature representation X emb is transformed into the frequency domain through the fast Fourier transform to obtain the frequency-domain feature F(Y); then, in the frequency domain, independent predictions are made for each frequency component F k , and this process is implemented by the method based on the double-layer encoder in S5; then, the frequency-domain loss function L freq is used to optimize the model by calculating the difference between the predicted frequency-domain feature and the true frequency-domain feature; the frequency-domain loss function is expressed as: where is the predicted frequency-domain feature, and F(Y) is the true frequency-domain feature; Finally, perform a fusion calculation on the final loss, combining the time-domain loss $L$ tmp and the frequency-domain loss $L$ freq to form the final loss function $L$ α : L α = α·L freq +(1 - α)·L tmp Where α is a weight parameter used to control the relative contribution of the time-domain and frequency-domain losses.

8. The method for predicting cluster network traffic based on multi-scale time feature fusion according to claim 1, wherein In the S7, a linear layer is used for decoding time series data, mapping the model dimension to the prediction length, and performing inverse normalization processing on the predicted network traffic data, specifically including: The input of the decoder is the spatio-temporal feature vector output by the encoder where B is the batch size, C is the dimension of the traffic metric, and d model is the dimension of the model feature; W dec and b dec are the parameters of the linear layer, which are the weight matrix and the bias vector respectively, and are optimized during training through backpropagation to make the predicted value as close as possible to the true value; the decoding objective is to map the high-dimensional features to the predicted length T, which is achieved through a trainable linear layer: This operation maps the d-dimensional features of each metric dimension C into the predicted logarithmic values for T time steps, and the output tensor shape is model Through linear transformation, while maintaining the independence of variables, the model generates multi-step prediction results by utilizing the cross-scale dependencies extracted by the encoder; ​ Finally, the logarithm value of the predicted output is restored to the original dimension through denormalization; let the mean value of each time step saved during preprocessing be and the variance be For the predicted time step \(t'\in[L + 1,L+T]\), the linear interpolation method is used to estimate its normalization parameters \(\mu\) t ' and \(\sigma\) t ', and the denormalization formula is: is the network traffic value of the final prediction, and the unit is the same as that of the original data before preprocessing.

9. A computer-readable storage medium, characterized in that, A computer program storage for implementing the method according to any one of claims 1 to 8 when the computer program is executed by a processor.

10. A computer program product comprising instructions, characterized in that, When the instruction runs on a computer, the computer is caused to execute the method according to any one of claims 1-8.

Citation Information

Cited By

  • Battery state prediction method and device based on battery time sequence model

    CN120064995A

  • A battery state prediction method and device based on a battery timing model

    CN120064995B

  • Data processing method and device, computer equipment and storage medium

    CN120897218A

  • Data processing method and device, computer device and storage medium

    CN120897218B

  • College water resource water supply pipe network monitoring system and method based on Internet of Things and medium

    CN120975393A