Multi-scale time sequence prediction method and device, terminal and medium
Through the combination of multi-scale decomposition and asymmetric cross-scale attention mechanism, information interaction and feature fusion between coarse-fine grain size are achieved, solving the problem of insufficient timing prediction accuracy in traditional methods and improving the accuracy of prediction results.
Patent Information
- Application Number
- CN202510873242.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-07-29
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional time series models can only be analyzed and predicted on a single particle scale, and it is difficult to effectively capture the long-term dependencies in time series information at different scales, resulting in insufficient accuracy of time series prediction.
Features of different particle sizes are extracted through multi-scale decomposition, combined with asymmetric cross-scale attention mechanisms to achieve bidirectional information interaction between coarse-fine grain sizes, and through asymmetric cross-scale cross-attention mechanisms, the time series feature fusion is achieved to capture short-term fluctuations and long-term trends in the time series.
The model's modeling ability of complex time dependencies is enhanced, the accuracy of prediction results in variable scenarios is improved, and the feature loss problem caused by the traditional method is solved due to the single scale.
Smart Images

Figure CN120387140A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular, to a multi-scale time series prediction method, device, terminal and medium. Background Art
[0002] Time series prediction is a technology for predicting future values based on historical data values of time series, and it plays an important role in aspects such as economic and financial markets, supply chain management, transportation, energy consumption, weather and industrial production. In these fields, time series prediction is usually used as an auxiliary tool for decision-making, which can effectively reduce costs and improve efficiency.
[0003] Traditional time series models often can only analyze and predict at a single granularity scale, without fully exploiting the complementarity of multi-scale time series data, and it is difficult to effectively capture the long-term dependence relationships in time series information at different scales, resulting in the technical problem of insufficient accuracy in existing time series prediction. Summary of the Invention
[0004] The present application provides a multi-scale time series prediction method, device, terminal and medium, which is used to solve the technical problem of insufficient accuracy in existing time series prediction.
[0005] To solve the above technical problem, a first aspect of the present application provides a multi-scale time series prediction method, including:
[0006] Obtain an original time series;
[0007] Perform sequence decomposition on the original time series to obtain multiple sub-time series with different granularity scales;
[0008] Based on the order of the granularity scales of each sub-time series from coarse to fine, take the first two sub-time series as the first sequence and the second sequence in turn, and then calculate the QKV features and the attention weights corresponding to the QKV features based on the first sequence and the second sequence, where the Q feature is calculated based on the first sequence, and the K feature and the V feature are calculated based on the second sequence feature;
[0009] Perform feature fusion processing on the second sequence and the attention weights to obtain a fused time series. If the next sub-time series corresponding to the current second sequence is not empty, update the first sequence based on the fused time series, and update the second sequence based on the next sub-time series, so as to obtain the latest fused time series based on the updated first sequence and second sequence. If the next sub-time series corresponding to the current second sequence is empty, construct a multi-scale time series according to the first sub-time series and each fused time series;
[0010] Based on the multi-scale time series, perform time series prediction to obtain a prediction result.
[0011] Preferably, it further includes:
[0012] Perform non-linear transformation preprocessing on the original time series through a preset multi-layer perceptron.
[0013] Preferably, the multi-layer perceptron includes: a layer normalization module, a fully connected layer, a GELU activation function, a Dropout calculation module, and a residual connection processing module.
[0014] Preferably, the sequence decomposition of the original time series to obtain multiple sub-time series with different granularity scales includes:
[0015] Perform sequence decomposition on the original time series through an average pooling downsampling processing method to obtain multiple sub-time series with different granularity scales.
[0016] Preferably, the feature fusion processing based on the second sequence and the attention weights to obtain a fused time series includes:
[0017] Perform layer normalization on the sum of the second sequence and the attention weights to obtain a preliminary fusion result;
[0018] Perform high-order feature extraction on the preliminary fusion result through a preset feed-forward neural network, and then perform residual connection and layer normalization processing on the extracted high-order features and the preliminary fusion result in sequence to obtain a fused time series.
[0019] Preferably, the time series prediction based on the multi-scale time series to obtain a prediction result includes:
[0020] Input the multi-scale time series into a preset linear neural network to determine the prediction result based on the output result of the linear neural network.
[0021] Preferably, the determination of the prediction result based on the output result of the linear neural network includes:
[0022] According to the multi-scale prediction result sequence output by the linear neural network, combined with a preset learnable weight coefficient sequence, perform weighted summation on each element in the multi-scale prediction result sequence to obtain the final prediction result according to the summation result.
[0023] The second aspect of this application provides a multi-scale time series prediction device, including:
[0024] An original sequence acquisition unit for acquiring an original time series;
[0025] A multi-scale sequence decomposition unit for decomposing the original time series to obtain multiple sub-time series with different granularity scales;
[0026] A cross-scale attention processing unit for, based on the order of the granularity scales of each sub-time series, sequentially taking the first two sub-time series as the first sequence and the second sequence, and then calculating QKV features and the attention weights corresponding to the QKV features based on the first sequence and the second sequence, where the Q feature is calculated based on the first sequence, and the K feature and the V feature are calculated based on the second sequence feature;
[0027] A cross-scale feature fusion unit for performing feature fusion processing based on the second sequence and the attention weights to obtain a fused time series. If the next sub-time series corresponding to the current second sequence is not empty, the first sequence is updated based on the fused time series, and the second sequence is updated based on the next sub-time series, so as to obtain the latest fused time series based on the updated first sequence and second sequence. If the next sub-time series corresponding to the current second sequence is empty, a multi-scale time series is formed according to the first sub-time series and each fused time series;
[0028] A time series prediction unit for performing time series prediction based on the multi-scale time series to obtain a prediction result.
[0029] The third aspect of the present application provides a multi-scale time series prediction terminal, including: a memory and a processor;
[0030] The memory is used to store program code, and the program code is used to implement a multi-scale time series prediction method provided in the first aspect of the present application;
[0031] The processor is used to read and execute the program code.
[0032] The fourth aspect of the present application provides a computer-readable storage medium, characterized in that program code is stored in the computer-readable storage medium, and the program code is used to be read and executed by a processor to implement a multi-scale time series prediction method provided in the first aspect of the present application.
[0033] From the above technical solutions, it can be seen that the present application has the following advantages:
[0034] The solution provided by this application extracts features of different granularities through multi-scale decomposition, realizes two-way information interaction between coarse and fine granularities by combining an asymmetric cross-scale attention mechanism, and simultaneously realizes cross-scale temporal feature fusion through an asymmetric cross-attention mechanism. It can effectively reduce noise and enhance feature complementarity. Through the above technical solutions, this application can simultaneously capture short-term fluctuations and long-term trends in time series, and solve the problem of feature loss caused by a single scale in traditional methods. The cross-scale attention mechanism promotes information fusion between subsequences of different granularities, enhances the model's ability to model complex time-dependent relationships, and thus improves the accuracy of prediction results in variable scenarios, thereby solving the problem that traditional methods are difficult to capture long-term dependence relationships and multi-scale dynamic features. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0036] Figure 1 It is a schematic flowchart of the process of an embodiment of a multi-scale time series prediction method provided by this application.
[0037] Figure 2 It is a schematic flowchart of the processing logic of a multi-layer perceptron in a multi-scale time series prediction method provided by this application.
[0038] Figure 3 It is a flow chart of the fusion logic of multi-scale subsequences in a multi-scale time series prediction method provided by this application.
[0039] Figure 4 It is a schematic structural diagram of an embodiment of a multi-scale time series prediction device provided by this application.
[0040] Figure 5 It is a schematic structural diagram of an embodiment of a multi-scale time series prediction terminal provided by this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] In the prior art, time series prediction technology is widely used in fields such as economic finance, supply chain management, and energy consumption. Traditional methods usually analyze based on a single time scale and are difficult to effectively capture dynamic features at different time granularities. For example, in the scenario of power load prediction, short-term fluctuations and long-term trends often intertwine, and a single-scale model cannot simultaneously take into account hourly changes and quarterly periodic patterns, resulting in prediction results deviating from actual needs.
[0042] To solve the above problems, researchers found that there is complementarity among multi-scale time series data, but existing methods lack an effective cross-scale feature interaction mechanism. Through analysis, it is found that subsequences with different time granularities may contain independent information patterns. For example, high-frequency subsequences reflect detailed fluctuations, and low-frequency subsequences reflect macroscopic trends. How to establish cross-scale correlations has become a technical difficulty, and a computational framework that can adaptively fuse multi-scale information needs to be designed.
[0043] In view of this, embodiments of the present application provide a multi-scale time series prediction method, device, terminal, and medium, which are used to solve the technical problem of insufficient accuracy in existing time series predictions.
[0044] To make the objectives, features, and advantages of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the embodiments described below are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0045] First, a detailed description of an embodiment of the multi-scale time series prediction method provided by the present application is as follows:
[0046] Please refer to Figure 1 , a multi-scale time series prediction method provided by an embodiment of the present application includes:
[0047] Step 101, obtain the original time series;
[0048] Step 102, decompose the original time series to obtain multiple sub-time series with different granularity scales;
[0049] Step 103, based on the order of the granularity scales of each sub-time series, sequentially use the first two sub-time series as the first sequence and the second sequence, and then calculate the QKV features and the attention weights corresponding to the QKV features based on the first sequence and the second sequence;
[0050] Among them, the Q feature is calculated based on the first sequence, and the K feature and the V feature are calculated based on the second sequence feature;
[0051] Step 104: Perform feature fusion processing based on the second sequence and the attention weights to obtain a fused time series. If the next sub-time series corresponding to the current second sequence is non-empty, update the first sequence based on the fused time series and update the second sequence based on the next sub-time series, so as to obtain the latest fused time series based on the updated first sequence and second sequence. If the next sub-time series corresponding to the current second sequence is empty, construct a multi-scale time series according to the first sub-time series and each fused time series.
[0052] Step 105: Perform time series prediction based on the multi-scale time series to obtain a prediction result.
[0053] Among them, the original time series refers to a set of observed data arranged in chronological order, which can be specifically collected through sensors, databases or logging systems and is used to provide basic prediction inputs. Sequence decomposition refers to converting the original sequence into multiple sub-sequences with different time granularities, which can be specifically implemented by means of average pooling downsampling, adaptive multi-scale decomposition, max or min pooling, and multi-resolution convolution. Its role is to separate feature patterns at different time scales. The order of granularity scales refers to arranging the sub-sequences according to the time window size, which can be specifically achieved by comparing the downsampling multiples of the pooling layer and is used to establish a processing order from coarse-grained to fine-grained or vice versa.
[0054] In the QKV feature calculation, the Q feature is generated based on the first sequence, and the K feature and V feature are generated based on the second sequence, which can be specifically implemented through a linear transformation matrix. Its role is to establish a query-key value pair for cross-scale attention interaction. Feature fusion processing refers to weighted superposition of the attention weights and the second sequence, which can be specifically implemented through residual connection and layer normalization and is used to retain the original features and enhance cross-scale information transmission. For example, the multi-scale time series after decomposition is , first, take two time series with different scales , as the inputs of the asymmetric cross-scale cross-attention module. The sequence is linearly projected into a space to generate a query matrix (Query); the sequence is linearly projected into two different spaces to generate a key matrix (Key) and a value matrix (Value). Then, the obtained feature matrix is used to calculate the corresponding attention weights.
[0055] Specifically, after the original time series is downsampled and decomposed into high-frequency, medium-frequency, and low-frequency subsequences, according to the order of the granularity of each sub-time series, the coarsest or finest-grained subsequence is used as the first sequence, and the second coarsest-grained subsequence is used as the second sequence. Through the asymmetric attention mechanism, the first sequence generates a Q vector to query the K vector and V vector of the second sequence, and calculates the cross-scale correlation weight. After this weight is superimposed on the second sequence, if there are still remaining un-fused subsequences, the current fusion result is used as the new first sequence, and the next subsequence is used as the new second sequence for continuous iteration. Finally, the coarsest or finest-grained subsequence is combined with all fusion results to form a multi-scale time series containing multi-level features, which is input into the prediction module to output the result.
[0056] This solution extracts features of different granularities through multi-scale decomposition, combines the asymmetric cross-scale attention mechanism to achieve two-way information interaction between coarse and fine granularities, and at the same time realizes cross-scale temporal feature fusion through the asymmetric cross-attention mechanism, which can effectively reduce noise and enhance feature complementarity. Through the above technical solution, this application can capture both short-term fluctuations and long-term trends in the time series, and solve the problem of feature loss caused by a single scale in traditional methods. The cross-scale attention mechanism promotes information fusion between different granularity subsequences, enhances the model's ability to model complex time-dependent relationships, and thus improves the accuracy of prediction results in variable scenarios.
[0057] Based on the above basic embodiment, the method provided by this application further proposes that before performing step 102, the original time series is preprocessed by a preset multi-layer perceptron for non-linear transformation.
[0058] More specifically, the multi-layer perceptron includes a layer normalization module, a fully connected layer, a GELU activation function, a Dropout calculation module, and a residual connection processing module. The overall formula is as follows:
[0059]
[0060] In the formula, is the output of the multi-layer perceptron, is the input of the multi-layer perceptron (the original time series), is the output after layer normalization, is the weight matrix, is the bias term.
[0061] Among them, the layer normalization module is a unit that normalizes the input time series features. Specifically, it can be implemented by calculating the mean and variance along the feature dimension, which is used to eliminate the dimensional difference of the input features and enhance the stability of the training process. The formula is as follows:
[0062]
[0063] In the formula, is the mean value, is the variance, is usually a very small constant used to prevent the denominator from being zero, is a learnable scaling factor, is the translation parameter.
[0064] A fully connected layer refers to a neural network layer that performs a linear transformation. Specifically, it can be implemented by multiplying a weight matrix with an input vector and adding a bias term. It is used to map the standardized time series features to a new dimensional space to extract potential information. The formula is as follows, where is the output of the fully connected layer.
[0065]
[0066] The GELU activation function refers to an activation function based on the Gaussian error linear unit. Specifically, its non-linear transformation characteristics can be implemented by an approximate calculation method, which is used to enhance the network's fitting ability for complex time series patterns.
[0067] The Dropout calculation module refers to a unit that randomly masks the outputs of some neurons during the training phase. Specifically, it can be implemented by generating a mask matrix using the Bernoulli distribution, which is used to reduce the co-adaptability between neurons and suppress the overfitting phenomenon. Dropout will randomly "put to sleep" some neurons with a specific probability, that is, they do not participate in the current forward propagation and backward propagation processes. And in the test phase, in order to obtain the most accurate prediction results, Dropout will "activate" all neurons again and use the average value of the fully connected layer for comprehensive judgment to ensure that the model can face new and unseen data in the best state and achieve stable and reliable performance output. The residual connection processing module refers to an arithmetic unit that directly adds the module input and output. Specifically, it can be implemented by adding tensors element by element, which is used to alleviate the vanishing gradient problem in the training of deep networks, not only speeding up the training speed of the network but also significantly improving the training accuracy, enabling deep networks to smoothly learn richer and more complex feature representations and providing a solid architecture foundation for solving highly complex tasks.
[0068] Specifically, such as Figure 2As shown in the figure, after the original time series is input into the multi-layer perceptron, it first undergoes layer normalization to eliminate the dimensional differences of features, and then undergoes linear transformation through a fully connected layer to achieve feature interaction. After introducing non-linear calculation through the GELU activation function, Dropout (random inactivation) is used to randomly mask some neurons to enhance the generalization performance. Finally, the processed features are superimposed with the original input through residual connection to form time series features with high-order expression ability. This processing method enables the explicit modeling of the potential correlation patterns between time steps, and at the same time maps the original feature space to a representation space more suitable for multi-scale decomposition.
[0069] By introducing a non-linear transformation module, this solution can effectively extract complex patterns such as periodicity and trend hidden in the time series, realizing the effective mixing and enhancement of time series features in the transformation space, providing more discriminative input features for subsequent multi-scale decomposition. At the same time, by combining layer normalization, Dropout and residual connection, while retaining the feature interaction ability of the fully connected layer, it effectively controls the stability and generalization ability of the training process. This preprocessing mechanism can capture dynamic patterns in the original data that are difficult to extract by linear methods, realizing the robust mapping of time series features in the non-linear space, and providing a feature representation basis with high information density for subsequent cross-scale attention fusion.
[0070] Furthermore, regarding the sequence decomposition process of step 102 of this application, the specific steps are as follows: When performing sequence decomposition on the original time series, an average pooling downsampling processing method is adopted to obtain multiple sub-time series with different granularity scales.
[0071] Among them, the average pooling downsampling processing method means that the original time series is segmented according to a fixed window length, and the average value of the data within each window is taken as a new data point. Specifically, a moving average operation with a window length of 2 can be used to implement it, and different time-resolution sub-sequences are generated by gradually reducing the sampling rate.
[0072] Among them, the granularity scale refers to the degree of time-resolution difference of the sub-time series, which can be specifically controlled by adjusting the step size parameter of the pooling window. For example, an exponential growth mode with a step size of 2 is set, so that the time step of each subsequent sub-sequence is twice that of the previous sequence.
[0073] Specifically, in the implementation process, the original time series is first input into the downsampling processing module, and a multi-level average pooling operation is used to generate a set of sub-sequences with different time resolutions. For example, for an original sequence of length T, the first-level pooling window is set to 2 to obtain a sub-sequence of length T / 2; the second-level pooling window is set to 4 to obtain a sub-sequence of length T / 4. This hierarchical processing enables high-frequency detail information to be retained in the fine-grained sub-sequences, while the coarse-grained sub-sequences capture long-term trend features.
[0074] This solution directly reduces the dimensionality of time-domain data through average pooling downsampling, which not only has higher computational efficiency but also can effectively suppress noise interference through the local average characteristics of the sliding window, making it more conducive to capturing long-term dynamic patterns in time series. Through the above technical solution, the original time series can be decomposed into multi-scale subsequences with complementary information, where the fine-grained subsequences retain short-term fluctuation details and the coarse-grained subsequences represent overall trend changes. The effective separation of these multi-scale features provides a basis for the subsequent cross-scale attention mechanism.
[0075] Furthermore, regarding the feature fusion process based on the second sequence and attention weights proposed in step 105 of this application to obtain a fused time series, the step process can be referred to the following example: perform layer normalization on the sum of the second sequence and attention weights to obtain a preliminary fusion result; perform high-order feature extraction on the preliminary fusion result through a preset feedforward neural network, and then perform residual connection and layer normalization on the extracted high-order features and the preliminary fusion result in sequence to obtain a fused time series. For a time series decomposed into n scales, n - 1 asymmetric cross-scale cross-attention modules need to be constructed successively, and each module is used to process the output result of the previous module and the next time scale. By continuously repeating the above cross-scale cross-attention mechanism and the process of the feedforward neural network, it is ensured that the last scale can be processed. Finally, the multi-scale time series after the asymmetric cross-scale cross-attention module is used for subsequent multi-scale prediction.
[0076] Among them, layer normalization refers to the standardization operation on feature data, which can be specifically implemented by calculating the mean and variance and performing a normalization transformation. Its role is to eliminate the dimensionality difference of features and improve training stability. A feedforward neural network refers to a multi-layer structure composed of a linear layer and an activation function, which can be specifically implemented by combining a fully connected layer and a GELU activation function. Its role is to extract high-order time series features through non-linear transformation. Residual connection refers to the operation of adding the network input and output, which can be specifically implemented by a skip connection structure. Its role is to alleviate the problem of gradient disappearance and accelerate model convergence. An asymmetric cross-scale cross-attention module refers to an attention mechanism that only allows the coarse-grained sequence to generate query vectors while the fine-grained sequence generates key-value vectors, which can be specifically implemented by generating Q, K, and V features for different scale subsequences respectively. Its role is to establish cross-scale temporal dependence relationships.
[0077] As Figure 3 shown, assume that the multi-scale time series after downsampling decomposition is , where represents the sequence of the original scale, and represents the sequence of the coarsest scale. First, the two time series of different scales , As the input of the asymmetric cross-scale cross-attention module, the sequence is linearly projected into a space to generate a query matrix (Query); the sequence is linearly projected into two different spaces to generate a key matrix (Key) and a value matrix (Value). Then, the obtained matrices are used to calculate the attention weights. Next, the attention output is residually connected to the input scale and the fused representation is obtained through layer normalization . Finally, the fused result extracts high-order features through a feed-forward neural network (FFN), and performs residual connection and layer normalization operations to obtain the final output time series representation . The calculation process is as follows:
[0078]
[0079]
[0080]
[0081]
[0082] Among them, is a learnable weight matrix, is a scaling factor.
[0083] For a time series decomposed into n scales, n - 1 asymmetric cross-scale cross-attention modules need to be constructed, and each module is used to process the output result of the previous module and the next time scale . By continuously repeating the above cross-scale cross-attention mechanism and the process of the feed-forward neural network, it is ensured that the last scale can be processed. Finally, these multi-scale time series representations after passing through the asymmetric cross-scale cross-attention module are , which are used for subsequent multi-scale prediction.
[0084] Specifically, in a single feature fusion process, the second sequence is first added to the attention weight and then input into the normalization module to eliminate feature distribution offset. The normalized result is then nonlinearly transformed through a feedforward neural network, for example, using a structure with a GELU activation function inserted between two fully connected layers to extract a time series representation containing high-order dynamic features. The output of the feedforward network is further residually connected with the preliminary fusion result, and the final fused time series is generated through a second layer normalization. For decomposition results containing multiple scales, it is necessary to sequentially construct multiple asymmetric cross-scale cross-attention modules. For example, when the original sequence is decomposed into three scales, two modules need to be constructed to handle the cross-scale interactions between X1 and X2, O1 and X3, and so on. Each module uses the same processing flow, and gradually fuses the fused time series features of all scales through a chain structure.
[0085] This solution introduces layer normalization to ensure feature distribution stability, utilizes feedforward neural networks to enhance feature expression, and combines residual connections to improve gradient propagation efficiency. An asymmetric attention mechanism effectively distinguishes the roles of sequences of different scales in feature extraction. For example, coarse-grained sequences are more suitable for capturing global patterns as query benchmarks, while fine-grained sequences provide local details as key values. This structured feature fusion approach more accurately establishes multi-scale temporal associations than traditional methods. Through the above technical solution, this application solves the information loss problem caused by insufficient cross-scale feature fusion in the prior art and avoids feature distribution conflicts caused by simple addition or concatenation operations. The combined use of layer normalization and residual connections effectively improves model training stability and prevents training oscillations caused by gradient anomalies. The asymmetric attention mechanism strengthens the targeted nature of cross-scale interactions by distinguishing roles, allowing coarse-grained trend features and fine-grained fluctuation features to participate in prediction in a complementary manner. The introduction of a feedforward neural network further exploits nonlinear temporal patterns in the fused features, providing a highly discriminative multi-scale feature representation for subsequent prediction.
[0086] Furthermore, step 106 of the present application proposes time series prediction based on multi-scale time series, and obtaining the prediction result includes: inputting the multi-scale time series into a preset linear neural network to determine the prediction result based on the output result of the linear neural network; according to the multi-scale prediction result sequence output by the linear neural network, combined with a preset learnable weight coefficient sequence, weighted summing up each element in the multi-scale prediction result sequence to obtain the final prediction result according to the summation result.
[0087] Among them, the linear neural network refers to a mapping model composed of a fully connected layer. Specifically, it can be implemented by using a fully connected layer structure whose input dimension matches the feature dimension of the multi-scale time series. Its function is to transform the multi-scale time series into Each time step feature map in is the predicted value The learnable weight coefficient sequence refers to a set of weight parameters that are dynamically adjusted through the back-propagation algorithm. Specifically, it can be implemented in the form of a trainable vector or matrix. Its function is to dynamically allocate the contribution of prediction results at different time scales to the final result.
[0088] Specifically, after the multi-scale time series is input into the linear neural network, the features of each time scale are passed through the fully connected layer to generate the prediction result sequence of the corresponding scale. These prediction result sequences are concatenated to form a multi-dimensional tensor, which can be used to learn the weight coefficient sequence The prediction results of each scale are weighted by matrix multiplication. In the weighting process, the prediction results of each time step are scaled according to the weight coefficient of the scale to which it belongs, and finally summed along the feature dimension to generate a single-dimensional output sequence as the final prediction result. , the formula is as follows:
[0089]
[0090] Among them, the weight of each scale This module can be automatically learned through model training, ensuring that the fusion results strike a balance between global and local trends. This module effectively integrates forecast information across different time scales, ensuring that the final output combines local mutation response capabilities with global trend perception, significantly improving overall forecast accuracy and making it applicable to a variety of complex scenarios.
[0091] This solution achieves dynamic adaptive fusion of multi-scale prediction results through a combination of learnable weight coefficients and linear neural networks, and can automatically adjust the contribution ratio of each time scale to the final result based on the characteristics of the input data. Through the above technical solution, this application solves the problem of rigid fusion of different scale features in multi-scale time series prediction, and effectively improves the model's ability to capture multi-scale correlation features in complex time series patterns. Through the dynamic weight allocation mechanism, the model can autonomously identify key time scales and suppress noise interference, thereby improving the accuracy and robustness of the final prediction results.
[0092] In order to further demonstrate the technical effect of the multi-scale time series prediction method provided by this application, this application provides simulation test results based on multiple time series data. The test results are shown in Table 1:
[0093]
[0094] Eight publicly available datasets were used in the experiment. Among them, the meteorological dataset contains meteorological indicators, the solar energy dataset involves solar power generation data, the power and traffic datasets cover power consumption and traffic flow information respectively, while the ETTh1, ETTh2, ETTm1, and ETTm2 datasets are the operating states of power transformers. The smaller the values of MSE (Mean Squared Error) and MAE (Mean Absolute Error), the better the model performance.
[0095] By comparing the performance of different time series prediction methods on different datasets, it can be seen that the method proposed in this application is superior to other existing methods.
[0096] The above is a detailed description of an embodiment of a multi-scale time series prediction method provided by this application. Next is a detailed description of an embodiment of a multi-scale time series prediction device corresponding to the above method embodiment.
[0097] Please refer to Figure 4 , an embodiment of a multi-scale time series prediction device provided by this application includes:
[0098] An original sequence acquisition unit 201 for acquiring an original time series;
[0099] A multi-scale sequence decomposition unit 202 for decomposing the original time series to obtain sub-time series of multiple different granularity scales;
[0100] A cross-scale attention processing unit 203 for sequentially taking the first two sub-time series as the first sequence and the second sequence based on the order of the granularity scales of each sub-time series, and then calculating the QKV features and the attention weights corresponding to the QKV features based on the first sequence and the second sequence, where the Q feature is calculated based on the first sequence, and the K feature and the V feature are calculated based on the second sequence feature;
[0101] A cross-scale feature fusion unit 204 for performing feature fusion processing based on the second sequence and the attention weights to obtain a fused time series. If the next sub-time series corresponding to the current second sequence is not empty, the first sequence is updated based on the fused time series, and the second sequence is updated based on the next sub-time series to obtain the latest fused time series based on the updated first sequence and the second sequence. If the next sub-time series corresponding to the current second sequence is empty, a multi-scale time series is formed according to the first sub-time series and each fused time series;
[0102] A time series prediction unit 205 for performing time series prediction based on the multi-scale time series to obtain a prediction result.
[0103] AsFigure 5 As shown, a multi-scale time series prediction terminal provided by the present application includes: a memory 33 and a processor 31, wherein the memory 33 and the processor 31 can be connected through a communication bus 34;
[0104] The memory 33 is used to store program codes, and the program codes are used to implement a multi-scale time series prediction method provided in the above-mentioned embodiment;
[0105] The processor 31 is used to read and execute the program codes.
[0106] Among them, the memory 33 refers to a hardware device for storing computer programs and data, and specifically can be implemented by a solid-state drive or a flash memory chip. Its function is to provide persistent storage support for the multi-scale time series prediction method, ensuring that the algorithm codes and data can still be retained after power-off. The processor 31 refers to an arithmetic control unit that executes program instructions, and specifically can be implemented by a multi-core central processing unit or a graphics processing unit. Its function is to accelerate time series decomposition, cross-scale attention weight calculation, and feature fusion processing through parallel computing, improving the prediction efficiency.
[0107] Specifically, during the operation of this terminal, the program codes stored in the memory are loaded and executed by the processor. First, the original time series data is obtained, and then multiple sub-time series with different granularity scales are generated through sequence decomposition. The processor, according to the order of the granularity scales of the sub-sequences, sequentially takes the first two sub-sequences as inputs, calculates query, key, value features, and corresponding attention weights through an asymmetric cross-scale attention mechanism, and completes cross-scale feature fusion. If there are subsequent sub-sequences, the input sequence is iteratively updated and the fusion process is repeated, finally generating a multi-scale time series set. The processor further calls a linear neural network to predict the multi-scale sequence, and combines the learnable weight coefficients to output the final prediction result.
[0108] The fourth aspect of the present application provides a computer-readable storage medium, characterized in that program codes are stored in the computer-readable storage medium and are used to be read and executed by a processor to implement a multi-scale time series prediction method provided in the above-mentioned embodiment.
[0109] Among them, a computer-readable storage medium refers to a physical carrier that can persistently store data information. Specifically, it can be implemented using a solid-state drive, an optical disc, or a magnetic storage device. Its function is to provide non-volatile storage space for program code. Program code refers to a computer program composed of executable instructions. Specifically, it can be written in Python or C++ language. Its function is to convert the logic of the multi-scale time series prediction method into an instruction sequence executable by a processor. A processor refers to a hardware unit that performs arithmetic and logical operations. Specifically, it can be implemented using a central processing unit or a graphics processing unit. Its function is to control the data loading and calculation process in the storage medium by reading and parsing program code.
[0110] Specifically, when the program code is executed by the processor, it first loads the original time series data from the storage medium and decomposes it into subsequences with different time resolutions through average pooling downsampling operations. The decomposed subsequences form a hierarchical structure in the order of coarsest to finest granularity. The processor uses the subsequences of the first two levels as the initial first sequence and second sequence respectively, generates Q, K, and V feature vectors through matrix operations, and calculates the cross-scale attention weights. In the feature fusion stage, the second sequence and the attention weights achieve information interaction through residual connection and layer normalization. If there are finer-grained subsequences, the fusion result is used as the new first sequence to continue the cross-scale attention calculation with the next level. When all levels are processed, the processor concatenates the initial coarse-grained sequence and the fusion results of each layer into a multi-scale time series, and finally generates a prediction result through a linear neural network.
[0111] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described terminal, device, and unit can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0112] In several embodiments provided in this application, it should be understood that the disclosed terminal, device, and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical, or other forms.
[0113] In the description of the present application and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein, for example, can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0114] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or similar expressions refer to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0115] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0116] In addition, in each embodiment of the present invention, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0117] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0118] As described above, the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of the present application.
Claims
1. A multi-scale time series prediction method, characterized in that Including: Obtain the original time series; Perform sequence decomposition on the original time series to obtain multiple sub-time series with different granularity scales; Based on the order of the granularity scales of each sub-time series, take the first two sub-time series as the first sequence and the second sequence in sequence, and then calculate the QKV features and the attention weights corresponding to the QKV features based on the first sequence and the second sequence, where the Q feature is calculated based on the first sequence, and the K feature and the V feature are calculated based on the second sequence feature; Perform feature fusion processing based on the second sequence and the attention weights to obtain a fused time series. If the next sub-time series corresponding to the current second sequence is not empty, update the first sequence based on the fused time series and update the second sequence based on the next sub-time series to obtain the latest fused time series based on the updated first sequence and second sequence. If the next sub-time series corresponding to the current second sequence is empty, construct a multi-scale time series according to the first sub-time series and each fused time series; Perform time series prediction based on the multi-scale time series to obtain a prediction result.
2. The multi-scale time series prediction method according to claim 1, characterized in that Also including: Perform non-linear transformation preprocessing on the original time series through a preset multi-layer perceptron.
3. A multi-scale time series prediction method according to claim 2, wherein The multi-layer perceptron includes: a layer normalization module, a fully connected layer, a GELU activation function, a Dropout calculation module, and a residual connection processing module.
4. A multi-scale time series prediction method according to claim 1, characterized in that The performing sequence decomposition on the original time series to obtain multiple sub-time series with different granularity scales includes: Perform sequence decomposition on the original time series through an average pooling downsampling processing method to obtain multiple sub-time series with different granularity scales.
5. A multi-scale time series prediction method according to claim 1, wherein The performing feature fusion processing based on the second sequence and the attention weights to obtain a fused time series includes: Perform layer normalization processing on the sum of the second sequence and the attention weights to obtain a preliminary fusion result; Perform high-order feature extraction on the preliminary fusion result through a preset feed-forward neural network, and then perform residual connection and layer normalization processing on the extracted high-order feature and the preliminary fusion result in sequence to obtain a fused time series.
6. A multi-scale time series prediction method according to claim 1, characterized in that The performing time series prediction based on the multi-scale time series to obtain a prediction result includes: Input the multi-scale time series into a preset linear neural network to determine the prediction result based on the output result of the linear neural network.
7. A multi-scale time series prediction method according to claim 6, characterized in that The determining the prediction result based on the output result of the linear neural network includes: According to the multi-scale prediction result sequence output by the linear neural network, combined with a preset learnable weight coefficient sequence, perform weighted summation on each element in the multi-scale prediction result sequence to obtain the final prediction result according to the summation result.
8. A multi-scale time series prediction device, characterized in that, Including: An original sequence acquisition unit for obtaining the original time series; A multi-scale sequence decomposition unit for performing sequence decomposition on the original time series to obtain multiple sub-time series with different granularity scales; A cross-scale attention processing unit, which is used to sequentially take the first two sub-time series as the first sequence and the second sequence based on the order of the granularity levels of each sub-time series, and then calculate the QKV features and the attention weights corresponding to the QKV features based on the first sequence and the second sequence, where the Q feature is calculated based on the first sequence, and the K feature and the V feature are calculated based on the second sequence feature; A cross-scale feature fusion unit, which is used to perform feature fusion processing based on the second sequence and the attention weights to obtain a fused time series. If the next sub-time series corresponding to the current second sequence is not empty, the first sequence is updated based on the fused time series, and the second sequence is updated based on the next sub-time series, so as to obtain the latest fused time series based on the updated first sequence and second sequence. If the next sub-time series corresponding to the current second sequence is empty, a multi-scale time series is constructed according to the first sub-time series and each fused time series; A time series prediction unit, which is used to perform time series prediction based on the multi-scale time series to obtain a prediction result.
9. A multi-scale time series prediction terminal, characterized in that, Comprising: A memory and a processor; The memory is used to store program code, and the program code is used to implement a multi-scale time series prediction method according to any one of claims 1 to 7; The processor is used to read and execute the program code.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program code, and the program code is used to be read and executed by the processor to implement a multi-scale time series prediction method according to any one of claims 1 to 7.