Layered wavelet attention-based non-stationary time series prediction method and system

By decomposing and adaptively fusing time series through the hierarchical wavelet attention mechanism, the problem of insufficient prediction accuracy of existing models when dealing with non-stationary time series is solved, and efficient prediction of complex dynamics is achieved.

CN120763597AActive Publication Date: 2025-10-10POWER SUPPLY SERVICE & MANAGEMENT CENT STATE GRID JIANGXI ELECTRIC POWER CO LTD

Patent Information

Application Number
CN202511295110.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2025-10-10
Estimated Expiration
2045-09-11

AI Technical Summary

Technical Problem

Existing time series prediction models have poor prediction performance when dealing with real-world non-stationary time series, especially mutation points and transient oscillations. Different dynamic modes interfere with each other in modeling, resulting in insufficient accuracy.

Method used

A hierarchical wavelet attention mechanism is used to perform wavelet derivative transformation on the time series, decomposing it into detail coefficient sequences of different frequency scales. The model is then modeled through an independent temporal encoder, adaptively fusing features of different frequency scales to generate wavelet domain feature representation, which is finally combined with exogenous variables for prediction.

Benefits of technology

Effectively decoupling and learning the stationary and non-stationary dynamics in time series improves the prediction accuracy of complex dynamics, avoids the mutual interference of information of different natures, and improves the predictive ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120763597A_ABST
    Figure CN120763597A_ABST
Patent Text Reader

Abstract

The invention discloses a non-stationary time sequence prediction method and system based on layered wavelet attention. The invention provides a refined prediction normal form of decomposition-independent modeling-adaptive fusion. The method comprises the following steps: firstly, effectively separating dynamic characteristics of endogenous time sequence data into a plurality of frequency scales through small waveguide number transformation, strengthening non-stationary components such as abrupt change and the like, carrying out layered independent modeling by adopting an independent time sequence encoder for each frequency scale, and then, carrying out self-adaptive scale fusion through a self-adaptive scale fusion mechanism; weighted fusion is dynamically performed on the coded features of all frequency scales at each time point, so that the model can focus on the current most critical dynamic information. According to the method, the problem of complex dynamic modeling in a non-stationary time sequence is fundamentally solved, and the accuracy and robustness of prediction are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of data processing, and particularly relates to a non-stationary time series prediction method and system based on hierarchical wavelet attention. BACKGROUND

[0002] Time series prediction plays a crucial role in many fields such as power load, weather, financial transactions, and traffic flow. Deep learning models based on the Transformer architecture have achieved remarkable success in capturing long-term dependencies of time series. However, existing models still face significant challenges when dealing with non-stationary time series commonly seen in the real world.

[0003] Real-world time series data, such as power consumption data, often contains abrupt points caused by extreme weather, policy changes, or equipment failures, as well as transient oscillations caused by short-term external factors such as holiday effects and market fluctuations. These non-stationary characteristics are manifested as high-frequency transient components in the signal.

[0004] Traditional Fourier transform-based analysis methods can identify the frequency components of the signal, but they lose the time dimension information and cannot locate the time points of the abrupt changes. Existing mainstream time series prediction models usually treat the time series as a whole or segment them into time slices for processing, trying to learn multiple patterns such as trends, periodicities, abrupt changes, and oscillations with a single model structure. This approach can easily lead to interference between different types of signal patterns during learning. For example, the model may smooth out critical abrupt information to adapt to stable periodicity; conversely, it may overfit in stable areas to capture abrupt changes.

[0005] Specifically, when applying deep learning models based on the Transformer architecture to electricity consumption prediction tasks, it is found that the model's prediction performance is poor in the areas of abrupt points and short-term oscillations. The root cause lies in the model's inability to effectively decouple and model the stationary components (such as seasonal trends) and non-stationary components (such as abrupt changes) of the data.

[0006] Therefore, there is an urgent need for a new technical solution that can effectively distinguish and learn stationary and non-stationary dynamics in time series separately, thereby improving the prediction accuracy of real-world time series containing complex dynamics. SUMMARY

[0007] The present application aims to solve the problem of insufficient modeling of non-stationary characteristics (especially abrupt points and transient oscillations) by existing time series prediction models, as well as the mutual interference of different dynamic patterns in modeling. It provides a non-stationary time series prediction method and system based on hierarchical wavelet attention, which can significantly improve the prediction accuracy of non-stationary time series.

[0008] To achieve the above object, the core idea of the present application is to propose a new paradigm of hierarchical independent modeling and adaptive fusion of non-stationary dynamics. Instead of simply parallel processing time domain and frequency domain, this method goes deep into the wavelet domain, and makes fine and decoupled time sequence dependent modeling of different dynamic characteristics represented by different frequency scales. The present application is implemented by the following technical scheme. A non-stationary time series prediction method based on hierarchical wavelet attention, comprising the following steps: a. receiving endogenous time series data; b. extracting features from the endogenous time series data to obtain an endogenous feature representation at each time slice position, which includes: i. wavelet derivative transformation: differentiating the endogenous time series data, and then performing multi-level discrete wavelet transform on the differentiated endogenous time series data to obtain a series of detail coefficient sequences belonging to different frequency scales; ii. hierarchical independent modeling: for each frequency scale, the corresponding detail coefficient sequence is taken as an independent time sequence and input into a time sequence encoder for time sequence dependent modeling to learn the dynamic pattern within the frequency scale, and an encoded feature sequence containing context information at the frequency scale is obtained; iii. adaptive scale fusion: for each time slice position, a scale fusion attention mechanism is used to dynamically weight and fuse the encoded features from all frequency scales at the time slice position to generate a wavelet domain feature representation at the time slice position, which integrates multi-scale non-stationary information, as part of the endogenous feature representation; c. generating a prediction value for the future time series based on the endogenous feature representation.

[0009] Further preferably, in step b, the feature extraction process further includes: processing the endogenous time series data through a time sequence encoder to obtain a time domain feature representation; and the wavelet domain feature representation generated in step iii and the time domain feature representation are fused through a gated fusion mechanism to jointly constitute the endogenous feature representation.

[0010] Further preferably, in step ii, the time sequence encoder used for time sequence dependent modeling of the detail coefficient sequences of different frequency scales is weight-shared.

[0011] Further preferably, the time sequence encoder is a Transformer architecture-based encoder containing a multi-head self-attention mechanism.

[0012] Further preferably, the scale fusion attention mechanism in step iii is implemented by a self-attention network that calculates the correlation between different frequency scale features and outputs the weighted fused features.

[0013] Further preferably, the gating fusion mechanism dynamically calculates the fusion weights according to the input time-domain feature representation and wavelet domain feature representation, and performs weighted summation on the time-domain feature representation and wavelet domain feature representation.

[0014] Further preferably, it further comprises: receiving exogenous variable data related to the endogenous time series data; embedding the exogenous variable data into exogenous feature representation; Before generating the prediction value, the fused endogenous feature representation and the exogenous feature representation are interacted and integrated.

[0015] The application also provides a non-stationary time series prediction system based on hierarchical wavelet attention, comprising: one or more processors; a memory having computer executable instructions stored thereon; Wherein, when the instructions are executed by the one or more processors, the system can execute the above-mentioned non-stationary time series prediction method based on hierarchical wavelet attention.

[0016] The application has the following characteristics: (1) Wavelet derivative transform and multi-scale dynamic separation: First, the received historical endogenous time series is subjected to first-order difference processing. This operation is approximately equivalent to differentiation on discrete signals, and its purpose is to amplify the "change rate" of the signal, thereby numerically strengthening the non-stationary characteristics such as sudden changes and oscillations. Subsequently, the differentiated sequence is subjected to multi-level discrete wavelet transform. This step is defined as "wavelet derivative transform". Through this transformation, the dynamic characteristics of the original sequence are effectively decomposed into multiple different frequency scales, each scale represented by a set of detail coefficient sequences, corresponding to different frequency wave information respectively.

[0017] (2) Hierarchical independent modeling and time series dependent deep learning: This step is the core innovation of the present application, which is called "hierarchical wavelet attention" mechanism. Unlike the prior art, which mixes all frequency information for processing, the present application independently and deeply models the time series for each frequency scale obtained in the previous step. Specifically, for each frequency scale j, we reorganize all the detail coefficients of the time slices at this scale into a time series sequence that is unique to this scale. Then, a time series encoder (such as a Transformer) is used to independently encode this sequence to fully learn and capture the dynamic patterns and time dependencies specific to this frequency scale. Through this "hierarchical independent modeling" approach, high-frequency sudden dynamics and low-frequency periodic dynamics are learned in separate channels, completely avoiding interference between different types of information.

[0018] (3) Adaptive scale fusion and feature integration: After independent modeling of each frequency layer, the information needs to be effectively integrated. For each time slice position i, the present application uses an "adaptive scale fusion" mechanism. This mechanism collects the feature vectors after encoding at all frequency scales for the time slice, and dynamically calculates and assigns weights through a scale fusion attention network. This allows the model to automatically determine which frequency scale information should be emphasized based on the specific circumstances of the current time slice. For example, when the sequence is stable, the model can automatically increase attention to low-frequency trend features; when the sequence fluctuates sharply, it can automatically increase the weight of high-frequency sudden features. Finally, a single, comprehensive wavelet domain feature representation is generated.

[0019] (4) Feature fusion and final prediction: The generated wavelet domain feature representation can be selectively gated and fused with the time domain features (representing stable components) extracted through traditional methods to obtain the final endogenous feature representation. Combined with the exogenous variable information input externally, the prediction head generates predicted values for future time series.

[0020] Through the above method, the present application transforms the prediction problem of non-stationary time series from a rough modeling of mixed patterns to a detailed processing procedure of decomposition-independent modeling-adaptive fusion, thereby fundamentally improving the model's ability to capture complex dynamics and prediction accuracy. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 Flowchart of the present application's non-stationary time series prediction method based on hierarchical wavelet attention. DETAILED DESCRIPTION

[0022] The following is a clear and complete description of the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of the present invention.

[0023] Reference Figure 1 This embodiment provides a non-stationary time series prediction method based on hierarchical wavelet attention, including: Step 101: Data reception. The system receives historical power load data as endogenous time series data. , L is the length of the endogenous time series data, The data domain that represents endogenous time series data can also receive exogenous variable data such as weather and holidays , represents the sample size dimension of exogenous variable data, Indicates the number of types of exogenous variables.

[0024] Step 102: Wavelet derivative transformation. This step aims to separate and enhance the non-stationary dynamics in the endogenous time series data. First, calculate the endogenous time series data The first-order difference sequence of ,in , to amplify the rate of change of the signal, is the data at time t, is the data at time t-1, is the data at time t in the first-order difference sequence. Then, the first-order difference sequence Also divided into Time slices. For each time slice application The discrete wavelet transform (for example, using the Daubechies wavelet basis) is used to obtain the approximate coefficients and Group detail coefficient In this embodiment, in order to force the model to pay attention to non-stationarity, only detail coefficients are used. Represents dynamic information at a specific frequency scale.

[0025] Step 103: Layered independent modeling. This step is the core of the present invention, which is to decouple the dynamics of each frequency scale and perform in-depth time series modeling. , all The corresponding detail coefficient of each time slice Collect them, j is the scale index, i is the time slice position index, and form a time series of coefficient sets .right each element in apply a linear embedding layer to map it to dimensional token, resulting in independent token sequences where the scale is , represent the data field of the token sequence. Using a shared-weighted temporal Transformer encoder (consisting of multi-head self-attention layers and feed-forward network layers), each token sequence is processed independently to calculate its internal temporal dependencies , resulting in an encoded feature sequence where TemporalEncoder denotes the temporal Transformer encoder function.

[0026] Step 104: Adaptive scale fusion. This step aims to intelligently fuse information from different frequency scales. For each time slice position , the corresponding feature vector is extracted from groups of feature sequences, forming a new sequence where denotes the feature sequence of scale at time slice position , and T denotes the transpose. This sequence containing vectors is input into a small self-attention network (scale fusion attention module) to calculate the correlation between vectors and output a weighted average fusion vector . The fusion vectors of all positions are concatenated to obtain the final wavelet domain feature representation where MeanPool is the mean pooling operation, and SelfAttention is the self-attention mechanism.

[0027] Step 105: Optional feature fusion and prediction. To capture the stationary components in the sequence, the endogenous time series data is divided into non-overlapping time slices of length , where is the length of each time slice. For each time slice, a linear projection layer is applied to map it to a dimensional vector, resulting in a time slice token sequence . At the same time, a learnable global token is initialized for subsequent interaction with exogenous variables.

[0028] For each time slice position, the time domain token and the wavelet domain feature are concatenated. The concatenated vector is input into a linear layer and passed through a Sigmoid activation function to calculate a gating score. Then the final fused feature is calculated. Finally, the fused endogenous feature and the exogenous variable information after embedding are sent to an information interaction module, and the prediction value of the future time step is output through a prediction head .

[0029] Step 106: model training. The entire model is trained end-to-end by minimizing the mean square error (MSE) or other loss function between the predicted value and the true value .

[0030] Another embodiment of the present application provides a non-stationary time series prediction system based on hierarchical wavelet attention, comprising: one or more processors; a memory having computer executable instructions stored thereon; wherein the instructions, when executed by the one or more processors, enable the system to perform the above-described non-stationary time series prediction method based on hierarchical wavelet attention. The method of the present application can be implemented on a system composed of one or more computers. The system includes a processor, a memory and an input / output interface.

[0031] It should be noted that the above only lists specific embodiments of the present application, and obviously the present application is not limited to the above embodiments, and there are many similar changes. All modifications directly derived or inferred from the disclosure of the present application by those skilled in the art shall fall within the scope of the present application.

[0032] The above is only a preferred embodiment of the present application, and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A non-stationary time series prediction method based on hierarchical wavelet attention, characterized in that: The following steps are involved: a. Receive endogenous time series data; b. Extracting features from the endogenous time series data to obtain endogenous feature representations at each time slice position. This process includes: i. Wavelet derivative transform: performing differential processing on the endogenous time series data, and then performing a multi-level discrete wavelet transform on the endogenous time series data after differentiation to obtain a series of detail coefficient sequences belonging to multiple different frequency scales; ii. Hierarchical Independent Modeling: For each frequency scale, the corresponding detail coefficient sequence is treated as an independent time series sequence and fed into a temporal encoder for temporal dependency modeling. This method learns the dynamic pattern within the frequency scale and obtains an encoded feature sequence containing contextual information at that frequency scale. iii. Adaptive scale fusion: For each time slice position, a scale fusion attention mechanism is used to dynamically weight the encoded features from all frequency scales at that time slice position. This generates a wavelet domain feature representation for that time slice position that integrates multi-scale non-stationary information, which serves as part of the endogenous feature representation. c. Generate forecasts for future time series based on the endogenous feature representation.

2. The method according to claim 1, characterized in that In step b, the feature extraction process further includes: processing the endogenous time series data through a temporal encoder to obtain a time domain feature representation; and fusing the wavelet domain feature representation generated in step iii with the time domain feature representation through a gated fusion mechanism to jointly constitute the endogenous feature representation.

3. The method according to claim 1 or 2, characterized in that In the step ii, the temporal encoder for temporal dependency modeling of detail coefficient sequences at different frequency scales is weight-shared.

4. The method according to claim 3, characterized in that The temporal encoder is an encoder based on the Transformer architecture, which includes a multi-head self-attention mechanism.

5. The method according to claim 1, wherein The scale fusion attention mechanism in step iii is implemented through a self-attention network, which calculates the correlation between features of different frequency scales and outputs weighted fused features.

6. The method according to claim 2, characterized in that The gated fusion mechanism dynamically calculates fusion weights based on the input time domain feature representation and wavelet domain feature representation, and performs weighted summation on the time domain feature representation and the wavelet domain feature representation.

7. The method according to claim 1, characterized in that Also includes: receiving exogenous variable data related to the endogenous time series data; embedding the exogenous variable data into exogenous feature representation; Before generating the predicted value, the fused endogenous feature representation and the exogenous feature representation are subjected to information interaction and integration.

8. A non-stationary time series prediction system based on hierarchical wavelet attention, characterized in that: include: one or more processors; a memory having computer-executable instructions stored thereon; When the instructions are executed by the one or more processors, the system is enabled to perform the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Lightweight time sequence prediction method based on discrete wavelet transform

    CN114219027A

  • Long-term time sequence prediction method based on time-frequency composite learning

    CN117951641A

  • Multivariable long-time sequence prediction method based on lifting wavelet transform

    CN118312769A

  • Multivariate time series prediction method based on timestamp and multi-scale modeling

    CN120493203A

  • Method and system for using image analysis for time-series forecasting

    US20250209049A1

Cited By

  • Multivariable time sequence prediction method based on iTransform and computer program product

    CN121614767A

  • Feature prediction method and device for non-stationary time series data

    CN122020139A