Network traffic prediction method based on periodic decomposition and similarity matching

By employing periodic decomposition and similarity matching methods, and utilizing Fast Fourier Transform and Dynamic Time Stretch Distance for network traffic prediction, this approach addresses the issues of high computational cost and poor interpretability in existing technologies. It achieves efficient and accurate network traffic prediction, making it suitable for system resource allocation and malicious attack detection.

CN120934783APending Publication Date: 2025-11-11UNIV OF ELECTRONICS SCI & TECH OF CHINA

Patent Information

Application Number
CN202510823240.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing network traffic prediction techniques involve huge computational demands and are susceptible to noise interference when processing time series of network traffic characteristic indicators. They also lack interpretability and are difficult to accurately predict short-term data changes.

Method used

By employing periodic decomposition and similarity matching, the system extracts time periods using Fast Fourier Transform, performs clustering based on dynamic time stretching distance, and uses the energy amplitude ratio method for prediction, thereby improving the interpretability and computational efficiency of network traffic prediction.

Benefits of technology

It achieves efficient and accurate network traffic prediction, can dynamically process large-scale data, has strong noise resistance, improves the interpretability and computational efficiency of prediction results, and is suitable for system resource allocation and malicious attack detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120934783A_ABST
    Figure CN120934783A_ABST
Patent Text Reader

Abstract

The invention discloses a network traffic prediction method based on periodic decomposition and similarity matching, and relates to the technical field of time sequence prediction, and the method comprises the steps: S1, collecting network traffic historical data; s2, extracting a time period from the network traffic historical data sequence through fast Fourier transform, and segmenting the network traffic historical data sequence into a plurality of periodic subsequences and a to-be-predicted sequence; s3, clustering all the periodic subsequences based on the dynamic time stretching distance to obtain a median subsequence capable of representing each class of features; s4, matching a median sequence with the highest similarity with the to-be-predicted sequence as a reference median sequence; and S5, based on the reference median subsequence, predicting the residual sequence value of the to-be-predicted sequence by using an energy amplitude ratio method. According to the method, subsequence matching is carried out by utilizing the periodicity and similarity of the network traffic, the data change in a short period can be accurately predicted, a complicated parameter selection process is not needed, and thus the prediction efficiency of the network traffic is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer network and intelligent operation and maintenance technology, specifically relating to a network traffic prediction method based on periodic decomposition and similarity matching. Background Technology

[0002] With the rapid development of internet technology, numerous network applications have emerged, leading to a continuous increase in network load. Accurate and real-time network traffic prediction can not only help system operators rationally allocate network resources according to actual business needs, but also detect whether the system is under malicious attack by comparing actual traffic with predicted traffic.

[0003] Network traffic forecasting is essentially a time series forecasting problem. Network traffic data exhibits significant time dependence, displaying periodicity, trends, randomness, and seasonality over time. Existing time series forecasting techniques include traditional statistical analysis models, machine learning models, and deep learning models. Traditional statistical analysis models such as AR, MA, and ARIMA have been widely applied in network traffic forecasting. While these algorithms have a solid theoretical foundation and strong interpretability, they are weak in explicitly modeling periodic features and are susceptible to noise. To address this issue, more powerful and flexible machine learning and deep learning algorithms such as RNN, LSTM, and GRU are gradually becoming a new trend in time series forecasting. These algorithms have stronger generalization capabilities and can capture complex patterns in network data, but they require high computational and storage overhead and lack interpretability.

[0004] Chinese patent document CN110740063B, published on October 25, 2019, discloses a method for predicting network traffic characteristic indicators based on signal decomposition and periodicity. It uses the EMD signal decomposition algorithm to perform empirical mode decomposition on the time series of network traffic characteristic indicators, obtaining multiple components and a residual term. The period of each component is calculated using Fast Fourier Transform (FFT). Individual prediction of each component is performed by resampling the periodic points within each component based on its period, forming a new sampled time series. Regression prediction is then conducted on this sampled time series, and finally, the residual term is predicted using a regression method. The predicted outputs of each component and the predicted output of the residual term are summed term by term to obtain the final prediction result. This invention combines signal decomposition technology from the field of digital signal processing with signal period analysis and component regression prediction techniques to achieve prediction of network traffic characteristic indicator time series.

[0005] However, the above technical solutions use the EMD signal decomposition algorithm to perform empirical mode decomposition on the time series of network traffic characteristic indicators. Processing these small signals not only involves a huge amount of computation, but also easily interferes with the accuracy of prediction. Summary of the Invention

[0006] To overcome the defects and shortcomings of the existing technology, this invention provides a network traffic prediction method based on periodic decomposition and similarity matching. By capturing the periodic characteristics of network traffic, the method uses the similarity between different periodic sequences of network data to predict the remaining data of a certain period, thereby improving the interpretability of network traffic prediction and reducing computational overhead.

[0007] This invention is achieved through the following technical solution:

[0008] A network traffic prediction method based on periodic decomposition and similarity matching includes the following steps:

[0009] S1. Collect historical network traffic data to obtain a historical network traffic data sequence;

[0010] S2. Extract time periods from historical network traffic data sequences and divide the historical network traffic data sequences into several periodic subsequences and sequences to be predicted.

[0011] S3. Cluster all periodic subsequences based on dynamic time-stretching distance to obtain median subsequences that can characterize the features of each class;

[0012] S4. Match the median subsequence with the highest similarity to the sequence to be predicted as the benchmark median subsequence;

[0013] S5. Based on the benchmark median subsequence, use the energy amplitude ratio method to predict the remaining sequence value of the sequence to be predicted.

[0014] In step S1, collecting historical network traffic data to obtain a historical network traffic data sequence refers to:

[0015] Set a start timestamp and an end timestamp, and the start timestamp and end timestamp should be spaced out by as many time periods as possible;

[0016] Set a sampling frequency that satisfies the Nyquist sampling theorem.

[0017] Network traffic data is collected using a set start timestamp, end timestamp, and sampling frequency to obtain a historical network traffic data sequence.

[0018] In S2, extracting the time period from the historical network traffic data sequence refers to using the Fast Fourier Transform to extract the time period T of the historical network traffic data sequence.

[0019] In S2, the periodic subsequence refers to a subsequence with a period length that is the same as the extracted time period.

[0020] In S2, the sequence to be predicted refers to an incomplete periodic sequence whose time length is shorter than the extracted time period.

[0021] S3 specifically includes:

[0022] K-means clustering of periodic subsequences is performed based on dynamic time-stretched distance to obtain cluster centroids;

[0023] The cluster centroids are selected as median subsequences to represent the general pattern of each subsequence class.

[0024] S4 specifically includes:

[0025] Extract the sequence of the same length as the sequence to be predicted from each median subsequence;

[0026] Calculate the dynamic time stretch distance between the sequence to be predicted and each truncated median subsequence, and select the median subsequence with the smallest dynamic time stretch distance as the reference median subsequence of the sequence to be predicted.

[0027] S5 specifically includes:

[0028] Calculate the energy amplitude ratio of the sequence value at each same position between the sequence to be predicted and the reference median subsequence, and obtain the average value of the energy amplitude ratio at all positions;

[0029] The remaining portion of the reference median subsequence is extracted and multiplied by the average of the energy amplitude ratios to obtain the sequence value of the remaining portion of the sequence to be predicted.

[0030] The energy amplitude ratio is obtained by the following formula:

[0031]

[0032] In the formula, P i and LC i This represents the i-th sequence value of the corresponding sequence.

[0033] The average value of the energy amplitude ratio at all locations is obtained by the following formula:

[0034]

[0035] Where p is the length of the sequence to be predicted, P i and LC i This represents the i-th sequence value of the corresponding sequence.

[0036] Compared with the prior art, the beneficial technical effects of the present invention are as follows:

[0037] 1. This invention utilizes the periodicity and similarity of network traffic for subsequence matching, which can accurately predict data changes in the short term. The dynamic time stretching distance overcomes the phase shift of time series and improves the accuracy of similarity judgment. The energy amplitude ratio method only requires simple calculations and is more efficient than complex models. It does not require a complicated parameter selection process, thereby improving the prediction efficiency of network traffic. It can not only help system operators to rationally allocate network resources according to actual business needs, but also detect whether the system is under malicious attack by comparing the actual traffic and the predicted traffic.

[0038] 2. In this invention, when collecting historical network traffic data, a time window spanning multiple periods is set to ensure complete coverage of period fluctuations, avoid period extraction errors, and ensure that the sampling frequency satisfies the Nyquist theorem to prevent signal distortion, thus providing high-quality input for subsequent processing.

[0039] 3. This invention, by using Fast Fourier Transform to accurately extract time periods, can efficiently process large-scale data and has strong noise resistance.

[0040] 4. This invention uses K-means clustering of periodic subsequences based on dynamic time-stretching distance to improve the robustness of subsequent matching.

[0041] 5. In this invention, a sequence of the same length as the sequence to be predicted is extracted from each median subsequence; the dynamic time stretch distance between the sequence to be predicted and each extracted median subsequence is calculated; and the median subsequence with the smallest dynamic time stretch distance is selected as the benchmark median subsequence of the sequence to be predicted. This can match the most similar periodic patterns in history and enhance the interpretability of the prediction results.

[0042] 6. This invention uses the energy amplitude ratio method, which avoids the problem of high computational overhead of complex models in the prior art and realizes lightweight real-time prediction. Attached Figure Description

[0043] Figure 1 This is a flowchart of a network traffic prediction method based on periodic decomposition and similarity matching according to an embodiment of the present invention;

[0044] Figure 2 A schematic diagram illustrating the difference between two sequences using Euclidean distance and dynamic time-stretched distance metrics.

[0045] Figure 3 A schematic diagram illustrating the process of segmenting historical network traffic data;

[0046] Figure 4 This is a schematic diagram of the prediction process using the energy amplitude ratio method. Detailed Implementation

[0047] The technical solution of the present invention will be clearly and completely described below with reference to specific embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0048] Example 1

[0049] This embodiment discloses a network traffic prediction method based on periodic decomposition and similarity matching, referring to... Figure 1-4 Network traffic prediction is essentially a time series prediction problem. Specifically, given an input sequence TS = {x1, ..., x...} of length L... i ,…,x L}, x i ∈R L (where R) L (Representing an L-dimensional real space, and similarly for subsequent notations) represents the sequence values ​​observed during a specific process. The goal of time series prediction is to obtain an output sequence S = {y1,…,y...} of length H. i ,…,y H}, y i ∈R H Since network traffic data often exhibits obvious periodicity, and different periodic sequences often show self-similarity, this embodiment collects historical time series data preceding the sequence value to be predicted to predict short-term network data traffic sequence values. Therefore, the method described in this embodiment includes five steps:

[0050] S1. Collect historical network traffic data to obtain a historical network traffic data sequence, specifically:

[0051] Set a start timestamp and an end timestamp, and the start timestamp and end timestamp should be spaced out by as many time periods as possible;

[0052] Set a sampling frequency that satisfies the Nyquist sampling theorem, which states that the sampling frequency of a time series must be greater than twice its highest frequency in order to recover the original signal without distortion.

[0053] Network traffic data is collected using a set start timestamp, end timestamp, and sampling frequency to obtain a historical network traffic data sequence.

[0054] In this embodiment, the preset start and end timestamps can be spaced two months apart, and the network traffic data period may be 24 hours. The two-month duration ensures that the historical network traffic data sequence contains enough periodic patterns for prediction. The sampling frequency can be understood as the duration of network service traffic collection in actual applications. In principle, it can be set to seconds, minutes, hours, and days, etc. However, in order to ensure accurate perception of traffic changes and balance the performance consumption of data processing and analysis, the sampling frequency f in this embodiment is... T This can be done once per minute; resulting in a historical network traffic data sequence TS = {x1,…,x}. i ,…,x L}, x i ∈R L .

[0055] S2. Extract the time period from the historical network traffic data sequence and divide the historical network traffic data sequence into several periodic subsequences and the sequence to be predicted; specifically: use Fast Fourier Transform to extract the time period T of the historical network traffic data sequence, divide the historical network traffic data sequence into subsequences with the same period length as the extracted time period, and the part with the remaining time length after dividing the periodic subsequences is less than the extracted time period as the incomplete periodic sequence to be predicted, i.e. the sequence to be predicted;

[0056] Network traffic exhibits periodicity by displaying similar data patterns at different times, specifically manifesting in the following ways: 1) Daily cycle: Network traffic data typically follows a 24-hour cycle, with usage generally higher during the day and significantly lower at night; 2) Weekly cycle: On weekdays, users frequently use network services, causing traffic to peak during the day, while on weekends, due to daytime travel, traffic peaks at night; 3) Seasonal cycle: Network traffic may fluctuate significantly during holidays, such as when certain gaming services hold events that lead to a marked increase in traffic. Therefore, segmenting historical network traffic sequences into periodic subsequences, each potentially containing different pattern characteristics, can be used to predict sequence values ​​for incomplete periodic sequences.

[0057] In this embodiment, the historical network traffic data sequence TS={x1,…,x i ,…,x L Performing a Fast Fourier Transform to extract the time period can be represented as follows:

[0058] CA = Fast Fourier Transform (TS)

[0059] Wherein, Fast Fourier Transform represents the Fast Fourier Transform function, and CA is the complex array of the output;

[0060] By mapping the indices of the output array CA to their corresponding frequencies, the amplitude and phase of each frequency component can be obtained. The relationship between frequency and index is as follows:

[0061]

[0062] Among them, f k It is the actual frequency of the frequency component, k is the index of the output array CA, and f T Where L is the sampling frequency and L is the sequence length;

[0063] Finding f using amplitude spectrum k The peak value, and the corresponding index k, is the main frequency component k. peak ;

[0064] The period of TS is calculated using the main frequency components. The relationship between the period and the main frequency components is as follows:

[0065]

[0066] Where T is the calculated period of TS, L is the length of the time series, and f T It is the sampling frequency, k peak It is the main frequency component.

[0067] The Fast Fourier Transform (FFT) is a fast algorithm developed based on the Discrete Fourier Transform (DFT) and its odd, even, imaginary, and real characteristics. It can be used to transform time series from the time domain to the frequency domain, thereby extracting their periodic components.

[0068] The historical network traffic data sequence TS is segmented, with the segmentation step size determined by the parameter T. Specifically, for an input sequence TS of length L, it is segmented into several periodic subsequences of length T, with the number of subsequences being... The subsequences with a remaining length of less than T are the sequences to be predicted, P. Figure 3 This illustrates the process of segmenting historical network traffic data. Thus, the original input sequence TS∈R L After the partitioning operation, a matrix form X∈R is generated. q*T (where R) q*T Let X represent a q*T real matrix space (subsequence notation follows similarly) and the sequence P to be predicted, where the i-th row of matrix X represents the subsequence of the i-th period. This process can be represented as:

[0069] [X,P] = Patch(TS,T),

[0070] Here, Patch represents the sequence segmentation operation, TS is the input sequence, and T is the segmentation step size.

[0071] S3. Cluster all periodic subsequences based on dynamic time-stretching distance to obtain the median subsequence that can characterize the features of each class, specifically:

[0072] K-means clustering is performed on the periodic subsequences based on the dynamic time-stretched distance to obtain the cluster centroids; the cluster centroids are selected as the median subsequences to represent the general pattern of each subsequence class.

[0073] Although periodic time subsequences are similar in shape, they are not perfectly aligned in time. Dynamic time stretching distance (Dynamic Time Stretch Distance) is insensitive to temporal shifts and scaling, and is therefore widely used for time series clustering. This invention uses Dynamic Time Stretch Distance to measure the similarity of periodic subsequences, thereby constructing a median subsequence set.

[0074] Dynamic time-stretched distance is an extension of the Euclidean distance (ED). For two subsequences of length n, U = {u1, ..., u...} i ,…,u n}, u i ∈R n and V = {v1,…,v} i ,…,v n}, v i ∈R n The calculation process for ED can be expressed as follows:

[0075]

[0076] The calculation process for the dynamic time stretch distance can be represented by the following recursive process:

[0077]

[0078] That is, the dynamic time-stretched distance between two sequences is the sum of the pointwise minimum Euclidean distances between the two sequences. Figure 2 This illustrates the difference between Euclidean distance and dynamic time-stretched distance metrics for two sequences. Compared to Euclidean distance, for any point 'a' in U, its distance to sequence V is not the distance to the corresponding time point 'b', but rather the distance to the nearest point 'c' in sequence V. Similarly, summing the distance metrics for all pairs of nearest points yields the distance between the two sequences. Therefore, using dynamic time-stretched distance to determine the similarity of sequences can better reflect the similarity of their shapes.

[0079] In this embodiment, K-means clustering is performed on all periodic subsequences in matrix X based on the similarity between subsequences using dynamic time-stretched distance metric, and the centroid is selected as the median subsequence. The detailed steps are as follows:

[0080] Initializing centroids: The matrix X contains q periodic subsequences, and the goal is to classify these q periodic subsequences into K classes. Since network traffic data may exhibit three different patterns on weekdays, weekends, and holidays, K can be set to 3, or adjusted based on the clustering results. To reduce subsequent iteration time and improve clustering accuracy, the initial centroids are selected as follows: a periodic subsequence is randomly chosen as the first centroid C1; the second centroid C2 is the periodic subsequence with the largest dynamic time stretch distance from C1; and the third centroid C3 is the periodic subsequence with the largest sum of dynamic time stretch distances from both C1 and C2.

[0081] Subsequence allocation: Calculate the dynamic time stretch distances of the remaining periodic subsequences to C1, C2, and C3, and assign the periodic subsequences to the nearest cluster centroids. If s is a periodic subsequence in X, its distance from each cluster centroid C is calculated. i The distance can be expressed as:

[0082] D i (s) = Dynamic time stretch distance (s, C) i ), i = 1, 2, 3,

[0083] Among them, D i (s) is the sequence s to the cluster centroid C. i The dynamic time stretch distance.

[0084] Assign s to the smallest D. i (s) represents the class of the centroid corresponding to the matrix X. After performing the assignment operation on all periodic subsequences in matrix X, the periodic subsequences are divided into 3 classes.

[0085] Update centroid: Assume M = {m1, ..., m} i ,…,m n}∈R n This cluster contains all periodic subsequences from the first class, and the centroid of the new cluster is the sum of all subsequences m contained in this class. i The average of the sequence values ​​at the same position (i = 1, 2, ..., n). This process can be represented as:

[0086]

[0087] The remaining clusters are traversed to perform the update centroid operation, resulting in new cluster centroids C1, C2, and C3.

[0088] Iteration: Repeatedly perform the operations of allocating subsequences and updating centroids. For each class, iteration stops when the dynamic time stretch distance between the latest centroid and the previous centroid is less than a predefined value. The final centroids C1, C2, and C3 are used as median subsequences, which can represent the general pattern of the sequences contained in the three classes.

[0089] S4. The median subsequence with the highest similarity to the sequence to be predicted is used as the baseline median subsequence. Specifically:

[0090] Extract the sequence of the same length as the sequence to be predicted from each median subsequence;

[0091] Calculate the dynamic time stretch distance between the sequence to be predicted and each truncated median subsequence, and select the median subsequence with the smallest dynamic time stretch distance as the reference median subsequence of the sequence to be predicted.

[0092] The incomplete periodic sequence to be predicted, segmented from historical data, can be understood as the sequence values ​​already generated within one period. By comparing the sequence values ​​of the overlapping portion of the sequence to be predicted with those of the median subsequence, the median subsequence most similar to the incomplete sequence to be predicted can be found. This median subsequence can then be used as the benchmark for prediction, thereby improving prediction accuracy.

[0093] In this embodiment, the preorder portions of three median subsequences are extracted, with the length of the extracted portion being the same as the length of the sequence P to be predicted. The dynamic time stretching distance between each subsequence and P is calculated to obtain the median subsequence most similar to the incomplete sequence to be predicted.

[0094] Extract the overlapping portion (EC) of the median subsequence and the incomplete sequence to be predicted of the same length. i This process can be represented as:

[0095] EC i =Extract(C i (1, p)), i = 1, 2, 3,

[0096] Where Extract represents the sequence extraction operation, p is the length of the sequence P to be predicted, and (1,p) means that the extracted part is from the first sequence value to the pth sequence value.

[0097] Calculate the dynamic time stretching distance r between the sequence to be predicted and the three truncated sequences respectively. i This process can be represented as: r i =Dynamic time stretch distance (P, EC) i ), i = 1, 2, 3,

[0098] Where P is the incomplete periodic sequence to be predicted, and EC i It corresponds to the part of the neutron subsequence that overlaps with the P position.

[0099] The smallest r iThe corresponding truncated sequence is the truncated sequence that is most similar to the sequence to be predicted. Accordingly, its corresponding median subsequence can be used as the benchmark for predicting incomplete periodic sequences, denoted as SC; at the same time, the overlapping part of the benchmark median subsequence with the position of equal length to P is denoted as LC.

[0100] S5. Based on the baseline median subsequence, the residual sequence value of the sequence to be predicted is predicted using the energy amplitude ratio method, specifically:

[0101] Calculate the energy amplitude ratio of the sequence value at each same position between the sequence to be predicted and the reference median subsequence, and take the average of the energy amplitude ratios at all positions;

[0102] The remaining portion of the reference median subsequence is extracted and multiplied by the average of the energy amplitude ratios to obtain the sequence value of the remaining portion of the sequence to be predicted.

[0103] This step utilizes the pattern similarity between the benchmark median subsequence and the sequence to be predicted for prediction. Although the specific values ​​of the benchmark median subsequence and the sequence to be predicted are different, their fluctuation trends and morphological characteristics are similar. Therefore, the amplitude ratio of the overlapping portion of the two sequences can be calculated. This ratio is then applied to the remaining portion of the benchmark subsequence, scaling the remaining portion of the benchmark median subsequence to predict the sequence value of the remaining portion of the sequence to be predicted within one period.

[0104] In this embodiment, the average value of the energy amplitude ratio at the same position of the truncated portion LC of the target median subsequence P and the reference median subsequence SC is calculated to quantify the amplitude difference between the target sequence P and the reference median subsequence SC at the same position.

[0105] For each corresponding position i in the sequences P and LC to be predicted, calculate the amplitude ratio. The global energy amplitude ratio w is obtained by averaging the amplitude ratios across all locations. This process can be expressed as:

[0106]

[0107] Where p is the length of the incomplete periodic sequence P to be predicted, P i and LC i This represents the i-th sequence value of the corresponding sequence.

[0108] The median subsequence SC is truncated to its remaining portion RC after removing the overlap with the predicted sequence P. This process can be represented as:

[0109] RC = Extract(SC,(p+1,T)),

[0110] Where SC is the baseline median subsequence, p is the length of the incomplete periodic sequence to be predicted, T is the period length, and (p+1,T) indicates that the truncated part is from the (p+1)th sequence value to the Tth sequence value.

[0111] Multiply each sequence value of the remaining RC by the global energy amplitude ratio w to obtain the prediction result F of the remaining sequence value of the incomplete periodic sequence to be predicted. Figure 4 The prediction process using the energy amplitude ratio method is shown. This process can be represented as:

[0112] F = w * RC

[0113] Here, multiplication represents pointwise operation, w is the global energy amplitude ratio, and RC is the non-positional overlap between the reference neutron subsequence SC and the sequence to be predicted P.

[0114] Given the input sequence TS={x1,…,x i ,…,x L}, x i ∈R L The output sequence obtained using the network traffic prediction method described above is F = {y1, ..., y...} i ,…,y T-p ], y i ∈R T-P That is, the length of the predicted sequence is Tp.

Claims

1. A network traffic prediction method based on periodic decomposition and similarity matching, characterized in that, Includes the following steps: S1. Collect historical network traffic data to obtain a historical network traffic data sequence; S2. Extract time periods from historical network traffic data sequences and divide the historical network traffic data sequences into several periodic subsequences and sequences to be predicted. S3. Cluster all periodic subsequences based on dynamic time-stretching distance to obtain median subsequences that can characterize the features of each class; S4. Match the median subsequence with the highest similarity to the sequence to be predicted as the benchmark median subsequence; S5. Based on the benchmark median subsequence, use the energy amplitude ratio method to predict the remaining sequence value of the sequence to be predicted.

2. The network traffic prediction method based on periodic decomposition and similarity matching according to claim 1, characterized in that: In S2, extracting the time period from the historical network traffic data sequence refers to using the Fast Fourier Transform to extract the time period T of the historical network traffic data sequence.

3. The network traffic prediction method based on periodic decomposition and similarity matching according to claim 5, characterized in that: In S2, the periodic subsequence refers to a subsequence with a period length that is the same as the extracted time period.

4. The network traffic prediction method based on periodic decomposition and similarity matching according to claim 3, characterized in that: In S2, the sequence to be predicted refers to an incomplete periodic sequence whose time length is shorter than the extracted time period.

5. A network traffic prediction method based on periodic decomposition and similarity matching according to any one of claims 1-4, characterized in that: S3 specifically includes: K-means clustering of periodic subsequences is performed based on dynamic time-stretched distance to obtain cluster centroids; The cluster centroids are selected as median subsequences to represent the general pattern of each subsequence class.

6. The network traffic prediction method based on periodic decomposition and similarity matching according to claim 5, characterized in that: S4 specifically includes: Extract the sequence of the same length as the sequence to be predicted from each median subsequence; Calculate the dynamic time stretch distance between the sequence to be predicted and each truncated median subsequence, and select the median subsequence with the smallest dynamic time stretch distance as the reference median subsequence of the sequence to be predicted.

7. The network traffic prediction method based on periodic decomposition and similarity matching according to claim 6, characterized in that: S5 specifically includes: Calculate the energy amplitude ratio of the sequence value at each same position between the sequence to be predicted and the reference median subsequence, and take the average of the energy amplitude ratios at all positions; The remaining portion of the reference median subsequence is extracted and multiplied by the average of the energy amplitude ratios to obtain the sequence value of the remaining portion of the sequence to be predicted.

8. The network traffic prediction method based on periodic decomposition and similarity matching according to claim 7, characterized in that: The energy amplitude ratio is obtained by the following formula: In the formula, P i and LC i This represents the i-th sequence value of the corresponding sequence.

9. A network traffic prediction method based on periodic decomposition and similarity matching according to claim 8, characterized in that: The average value of the energy amplitude ratio at all locations is obtained by the following formula: Where p is the length of the sequence to be predicted, P i and LC i This represents the i-th sequence value of the corresponding sequence.

10. A network traffic prediction method based on periodic decomposition and similarity matching according to claim 1, characterized in that: In step S1, collecting historical network traffic data to obtain a historical network traffic data sequence refers to: Set a start timestamp and an end timestamp, with a certain number of time periods between them; Set a sampling frequency that satisfies the Nyquist sampling theorem. Network traffic data is collected using a set start timestamp, end timestamp, and sampling frequency to obtain a historical network traffic data sequence.

Citation Information

Patent Citations

  • A Network Traffic Feature Indicator Prediction Method Based on Signal Decomposition and Periodicity

    CN110740063B

Cited By

  • Main and standby network switching method and system for server and intelligent terminal

    CN121984917A

  • A method, system, and smart terminal for primary / backup network switching of a server.

    CN121984917B