Time series prediction method based on variable decomposition and convolution attention modeling
By using variable decomposition and convolutional attention structures in the time series prediction model, the trend and seasonal components in the data are extracted and modeled, and the problem of existing models is difficult to distinguish and model these components is solved, achieving higher prediction accuracy and computational efficiency.
Patent Information
- Application Number
- CN202510494484.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-21
AI Technical Summary
Existing time series prediction models are difficult to effectively distinguish and model seasonal and trendy components in data, especially when complex nonlinear patterns exist in the data, which affect the prediction effect.
The method based on variable decomposition and convolutional attention modeling is adopted to perform Gaussian distribution variable decomposition on historical time series data, trend components and seasonal components are extracted, and long-term nonlinear relationships and periodic features are captured respectively through convolutional attention structure and linear mapping processing to achieve effective integration of information.
By accurately decomposing trend and seasonal components, the model can better capture nonlinear relationships and periodic patterns in the time series, improving prediction accuracy and model computational efficiency.
Smart Images

Figure CN120011722A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of time series prediction, and specifically relates to a time series prediction method based on variable decomposition and convolutional attention modeling. Background Art
[0002] Time series forecasting is the process of using historical time series data to predict values at future time points. Time series data is time-dependent data that is arranged in chronological order. The prediction of time series data of electric transformer temperature is an important problem, which predicts future time changes based on past observations of the time series.
[0003] With the development of deep learning technology, the models of time series prediction have gradually shifted from traditional statistical methods to neural network-based methods. Typical deep learning models include recurrent neural networks (RNNs), which can capture short-term and long-term dependencies in time series; long short-term memory networks (LSTMs) and gated recurrent units (GRUs), which are variants of RNNs that solve the gradient vanishing problem of traditional RNNs by introducing gating mechanisms and can effectively model long-term dependencies; convolutional neural networks (CNNs), which are mainly used to extract local features in time series, especially in processing local time dependencies and feature learning; and Transformer models, which capture long-distance dependencies through self-attention mechanisms and support parallel computing, and have been widely used in time series prediction in recent years. However, they usually ignore the seasonal and trend components in the data, or lack sufficient flexibility in modeling these components. In particular, when there are complex nonlinear patterns in the data, traditional neural network models find it difficult to effectively distinguish and model different components, thus affecting the prediction effect. Summary of the invention
[0004] In order to overcome the problems in the prior art, the present invention proposes a time series prediction method based on variable decomposition and convolutional attention modeling.
[0005] The technical solution of the present invention to solve the above technical problems is as follows: The present invention provides a time series prediction method based on variable decomposition and convolutional attention modeling, comprising the following steps: Obtain historical time series data of electric transformer temperature; Constructing an electric transformer temperature time series prediction model, the processing steps in the constructed electric transformer temperature time series prediction model include: preprocessing the historical time series data; performing Gaussian distribution variable decomposition on the preprocessed time series data to obtain trend components and seasonal components; modeling the trend component to capture the long-term nonlinear relationship in the time series; performing linear mapping processing on the seasonal component to retain its periodic characteristics and avoid introducing too many complex relationships; adjusting the weight according to the contribution of the trend component and the seasonal component to achieve effective integration of information; Training the electric transformer temperature time series prediction model to obtain a trained electric transformer temperature time series prediction model; The electric transformer temperature time series data to be predicted is input into the trained electric transformer temperature time series prediction model to obtain a prediction result.
[0006] Furthermore, the preprocessing of the historical time series data includes: normalizing the historical time series data, and performing affine transformation on the normalized historical time series data.
[0007] Furthermore, the pre-processed time series data is subjected to Gaussian distribution variable decomposition to obtain trend components and seasonal components, including: Initialize the convolution kernel weights using Gaussian distribution; The preprocessed historical time series data is padded to keep the length of the time series data before and after convolution consistent, and the initialized convolution kernel is used to smooth the time series data to extract the trend component: ; In the above formula, Indicates trend component; Represents the preprocessed historical time series data; Represents the initialization of convolution kernel weights; Indicates filling; represents convolution; Subtract the trend component from preprocessed time series data Get seasonal ingredients .
[0008] Furthermore, the trend component is modeled to capture the long-term nonlinear relationship in the time series, including: Construct multiple channels, input the trend component into each channel, use convolution kernels of different sizes to perform convolution operations on each channel, and integrate the features between different channels through point-by-point convolution to form a multi-scale fused feature representation; The fused features are processed using a 1×1 convolution, and the result is then multiplied element-wise with the trend component of the input.
[0009] Furthermore, all convolution operations are in the form of depthwise separable convolutions.
[0010] Furthermore, the weights are adjusted according to the contribution of the trend component and the seasonal component to achieve effective integration of information, including: The prediction results of seasonal components and trend components are weighted and fused to obtain the weighted fusion features: ; in, Represents the weighted fusion feature; The result of feature extraction representing the seasonal component; The result of feature extraction representing the trend component; Represents the weighting coefficient, which is used to adjust the contribution of seasonal components and trend components.
[0011] Furthermore, the constructed electric transformer temperature time series prediction model also includes: Through residual connection, the weighted fused features are added to the preprocessed time series data to obtain the residual fusion features: ; in, Represents the weighted fusion feature; is the weight matrix; is the bias vector; It is the intermediate feature after being processed by the fully connected layer and the ReLU activation function; Represents the residual fusion feature.
[0012] Compared with the prior art, the present invention has the following technical effects: The present invention first performs variable decomposition on the historical time series data of the electric transformer temperature, and uses a learnable one-dimensional convolution operation to accurately decompose the trend component and seasonal component of the time series data, thereby overcoming the problem of insufficient flexibility of the traditional decomposition method. In the modeling of the seasonal component, a simple and efficient linear transformation is adopted, and feature extraction is performed by learning the periodic pattern of seasonal changes. In the modeling of the trend component, a convolutional attention structure is adopted to capture the nonlinear relationship in the long-term trend while ensuring the computational efficiency of the model. Finally, the features of the seasonal component and the trend component are fused and passed through the residual learning module to further improve the prediction accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0014] Figure 1 A flowchart of a time series prediction method based on variable decomposition and convolutional attention modeling of the present invention; Figure 2 Fitting diagram of the prediction results on the ETTm transformer temperature dataset; Figure 3 It is a structural schematic diagram of the time series prediction method based on variable decomposition and convolutional attention modeling of the present invention. DETAILED DESCRIPTION
[0015] In order to further explain the technical means and effects taken by the present invention to achieve the predetermined invention purpose, the specific implementation methods, structures, features and effects of the technical solutions proposed by the present invention are described in detail below in conjunction with the accompanying drawings and preferred embodiments. The specific features, structures or characteristics in one or more embodiments may be combined in any suitable form. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by technicians in the technical field of the present invention.
[0016] In one embodiment of the present invention, referring to Figure 1-Figure 3 , provides a time series prediction method based on variable decomposition and convolutional attention modeling, including the following steps: Obtain historical time series data of the electric transformer temperature; construct an electric transformer temperature time series prediction model, wherein the electric transformer temperature time series prediction model includes: preprocessing the historical time series data; performing Gaussian distribution variable decomposition on the preprocessed time series data to obtain trend components and seasonal components; modeling the trend components to capture the long-term nonlinear relationship in the time series; performing linear mapping processing on the seasonal components to retain their periodic characteristics and avoid introducing too many complex relationships; adjusting weights according to the contribution of the trend components and the seasonal components to achieve effective integration of information; Training the electric transformer temperature time series prediction model to obtain a trained electric transformer temperature time series prediction model; The electric transformer temperature time series data to be predicted is input into the trained electric transformer temperature time series prediction model to obtain a prediction result.
[0017] The following is a detailed explanation of each of the above steps: Step 100: Obtain historical time series data of the electric transformer temperature and construct an electric transformer temperature time series prediction model.
[0018] Construct a time series prediction model for electric transformer temperature, refer to Figure 3 The electric transformer temperature time series prediction model includes a reversible normalization module, a variable decomposition module, a linear modeling module, a convolutional attention module, a feature fusion module, a residual learning module and a prediction layer; the input of the reversible normalization module receives the historical time series data of the electric transformer temperature, the output of the reversible normalization module is connected to the input of the variable decomposition module, the output of the variable decomposition module is respectively connected to the input of the linear modeling module and the input of the convolutional attention module, the output of the linear modeling module and the output of the convolutional attention module are respectively connected to the input of the feature fusion module, the output of the feature fusion module is connected to the input of the residual learning module, the input of the residual learning module is connected to the output of the reversible normalization module, the output of the residual learning module is connected to the prediction layer, and the prediction data is output through the prediction layer.
[0019] Among them, the reversible normalization module is used to preprocess the historical time series data of electric transformer temperature; the variable decomposition module is used to perform Gaussian distribution variable decomposition on the preprocessed data to obtain trend components and seasonal components; the convolutional attention module is used to model the trend components and capture the long-term nonlinear relationships in the time series; the linear modeling module is used to perform linear mapping processing on the seasonal components to retain their periodic characteristics and avoid introducing too many complex relationships; the feature fusion module is used to adjust the weights according to the contribution of trend and seasonal components to achieve effective integration of information; the residual learning module is used to alleviate the gradient vanishing problem in deep networks.
[0020] As a specific example, this step 200 may include the following sub-steps: Step 210: pre-process the historical time series data of the electric transformer temperature.
[0021] In this embodiment, a specific implementation of step 210 may be: Step 2101: The historical time series data of the electric transformer temperature is expressed as , and its normalized calculation formula is as follows: ; Where N is the batch size, D is the number of variables, L is the time step, µ is the mean, is the standard deviation; represents the historical time series data of normalized electric transformer temperature; Represents historical time series data of electrical transformer temperature.
[0022] Step 2102: The normalized value is affine transformed using the following formula: ; in, represents the historical time series data of the electric transformer temperature after affine transformation; γ represents the linear transformation factor of the affine transformation; β is the translation factor of the affine transformation.
[0023] Step 220: Perform Gaussian distribution variable decomposition on the preprocessed time series data to obtain trend components and seasonal components.
[0024] Extracting the long-term trend components of time series data is important for capturing the overall change pattern. To achieve this goal, a convolution-based variable decomposition module is used, which can smooth the time series data through convolution operations and extract the trend components. Trend component extraction helps to better understand the global change characteristics of the time series, while separating short-term cyclical fluctuations and extracting trend components.
[0025] This initialization method ensures that the center position of the convolution kernel has the largest weight, while the weight of the position away from the center gradually decreases. This convolution operation will pay more attention to the local central area of the time series, providing a weighted sliding average to extract smooth trend information. The convolution operation is equivalent to a weighted sliding average on the time axis, and the weight distribution of the Gaussian kernel makes the adjacent time steps have a greater impact on the trend decomposition. Compared with the ordinary sliding average, the Gaussian kernel can more smoothly transition the fluctuations in the time series, especially for the noisy series with better robustness.
[0026] Step 2201: Use Gaussian distribution to initialize the convolution kernel weights so that the weight at the center is the largest and the weight at the edge is smaller.
[0027] Specifically define a convolution kernel, each element of the convolution kernel Initialized by the following formula: ; Among them, K is the convolution kernel size, It is the standard deviation of the weight distribution of the control convolution kernel.
[0028] Step 2202: pad the input preprocessed time series data to keep the sequence length consistent before and after convolution, use the initialized convolution kernel to smooth the time series data, and extract the trend component. The formula is as follows: ; In the above formula, Indicates trend component; Represents the preprocessed historical time series data; Represents the initialization of convolution kernel weights; Indicates filling; Represents convolution.
[0029] Step 2203: Subtract the trend component from the input preprocessed time series data Get seasonal ingredients , the formula is as follows: ; In the above formula, Indicates seasonal component; Represents the preprocessed time series data of the input.
[0030] Step 230: Use the convolutional attention mechanism to model the trend component and capture the long-term nonlinear relationship in the time series.
[0031] The main function of the convolutional attention network is to extract features at different levels through multi-scale convolution. It combines convolution kernels of different sizes and captures patterns and features at different scales through multiple convolution operations. Finally, it fuses features of different scales together and performs dot multiplication with the original input to achieve the effect of the attention mechanism.
[0032] In this embodiment, a specific implementation of step 230 may be: Step 2301: In order to effectively capture the different scale features in the time series data, multiple channels are constructed, the trend component is input into each channel, and convolution operations are performed on each channel using convolution kernels of different sizes. Feature integration between different channels is achieved through point-by-point convolution to form a multi-scale fused feature representation.
[0033] The sizes of convolution kernels include 7×7, 11×11, and 21×21. To improve efficiency, refer to Figure 3 , which are actually 1 × 7, 7 × 1, 1 × 11, 11 × 1, 1 × 21, and 21 × 1. The combination of these convolution kernels ensures that both short-term local features and long-term trends are captured simultaneously.
[0034] All convolution operations are performed in the form of depth-wise separable convolutions, which first perform convolutions independently on each channel and then integrate features between different channels through point-by-point convolutions (1×1 convolutions).
[0035] The features extracted by different convolution kernels are fused together to form a multi-scale feature representation. These different convolution kernels extract multi-level features, including local details and global trends, which can more comprehensively describe the trend pattern in the input time series after fusion. The fusion process is implemented by simple element-by-element addition, allowing the model to retain the feature information of each scale.
[0036] Step 2302: Use a 1×1 convolution to process the fused features, and then multiply the result by the input trend component element by element.
[0037] This operation can be regarded as an attention mechanism, which dynamically adjusts the importance of features at each position by weighting the original features, enabling the model to adaptively and selectively amplify or suppress the trend part of the features.
[0038] Step 240: Perform linear mapping on the seasonal component to preserve its periodic characteristics and avoid introducing too many complex relationships.
[0039] The characteristics of seasonal components usually show relatively regular periodic changes, so for the feature extraction of this part, the present invention adopts a linear transformation. This simple and effective method can well preserve the seasonal signal without introducing additional modeling complexity. Since the characteristics of seasonal components are relatively stable and periodic, complex convolution or attention mechanisms may bring the risk of overfitting. Through linear transformation, periodic signals can be directly identified in the input features, and these features can be maintained in the output, ensuring that seasonal changes are fully reflected in the prediction results.
[0040] Step 250: Adopt a feature fusion strategy to adjust the weights according to the contribution of the trend component and the seasonal component to achieve effective integration of information.
[0041] After completing the modeling of seasonal components and trend components, the prediction results of the two are weighted and fused, and finally the weighted fusion features are obtained. The weighted fusion process can be expressed as: ; in, Represents the weighted fusion feature; The result of feature extraction representing the seasonal component; The result of feature extraction representing the trend component; Represents the weighting coefficient, which is used to adjust the contribution of seasonal components and trend components. Preferably, Take 0.7.
[0042] Step 260: Alleviate the vanishing gradient problem in deep networks through the residual learning module.
[0043] Through residual connection, the weighted fused features are added to the preprocessed time series data to obtain the residual fusion features: ; in, Represents the weighted fusion feature; is the weight matrix; is the bias vector; It is the intermediate feature after being processed by the fully connected layer and the ReLU activation function; Represents the residual fusion feature.
[0044] Step 270: Use the residual fusion features to predict the time series.
[0045] Map the residual fusion features to the output space through a fully connected layer The output prediction sequence lengths are 96, 192, 336, and 720. This prediction layer achieves the mapping of high-dimensional features to prediction values through simple and effective linear transformation. This design not only simplifies the computational complexity, but also improves the prediction accuracy and generalization ability of the model.
[0046] Step 300: training the electric transformer temperature time series prediction model to obtain a trained electric transformer temperature time series prediction model.
[0047] Step 400: input the electric transformer temperature time series data to be predicted into the trained electric transformer temperature time series prediction model to obtain a prediction result.
[0048] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A time series prediction method based on variable decomposition and convolutional attention modeling, characterized in that: The following steps are involved: Obtain historical time series data of electric transformer temperature; Constructing an electric transformer temperature time series prediction model, the processing steps in the constructed electric transformer temperature time series prediction model include: preprocessing the historical time series data; performing Gaussian distribution variable decomposition on the preprocessed time series data to obtain trend components and seasonal components; modeling the trend component to capture the long-term nonlinear relationship in the time series; performing linear mapping processing on the seasonal component to retain its periodic characteristics and avoid introducing too many complex relationships; adjusting the weight according to the contribution of the trend component and the seasonal component to achieve effective integration of information; Training the electric transformer temperature time series prediction model to obtain a trained electric transformer temperature time series prediction model; The electric transformer temperature time series data to be predicted is input into the trained electric transformer temperature time series prediction model to obtain a prediction result.
2. According to claim 1, a time series prediction method based on variable decomposition and convolutional attention modeling is characterized in that: The preprocessing of the historical time series data includes: normalizing the historical time series data, and performing affine transformation on the normalized historical time series data.
3. According to claim 2, a time series prediction method based on variable decomposition and convolutional attention modeling is characterized in that: The preprocessed time series data is subjected to Gaussian distribution variable decomposition to obtain trend components and seasonal components, including: Initialize the convolution kernel weights using Gaussian distribution; The preprocessed historical time series data is padded to keep the length of the time series data before and after convolution consistent, and the initialized convolution kernel is used to smooth the time series data to extract the trend component: ; In the above formula, Indicates trend component; Represents the preprocessed historical time series data; Represents the initialization of convolution kernel weights; Indicates filling; represents convolution; Subtract the trend component from preprocessed time series data Get seasonal ingredients .
4. The time series prediction method based on variable decomposition and convolutional attention modeling according to claim 3 is characterized in that: The trend component is modeled to capture the long-term nonlinear relationship in the time series, including: Construct multiple channels, input the trend component into each channel, use convolution kernels of different sizes to perform convolution operations on each channel, and integrate the features between different channels through point-by-point convolution to form a multi-scale fused feature representation; The fused features are processed using a 1×1 convolution, and the result is then multiplied element-wise with the trend component of the input.
5. The time series prediction method based on variable decomposition and convolutional attention modeling according to claim 4 is characterized in that: All convolution operations are in the form of depthwise separable convolutions.
6. The time series prediction method based on variable decomposition and convolutional attention modeling according to claim 4 is characterized in that: The weights are adjusted according to the contribution of trend components and seasonal components to achieve effective integration of information, including: The prediction results of seasonal components and trend components are weighted and fused to obtain the weighted fusion features: ; in, Represents the weighted fusion feature; The result of feature extraction representing the seasonal component; The result of feature extraction representing the trend component; Represents the weighting coefficient, which is used to adjust the contribution of seasonal components and trend components.
7. The time series prediction method based on variable decomposition and convolutional attention modeling according to claim 6 is characterized in that: The constructed electric transformer temperature time series prediction model also includes: Through residual connection, the weighted fused features are added to the preprocessed time series data to obtain the residual fusion features: ; in, Represents the weighted fusion feature; is the weight matrix; is the bias vector; It is the intermediate feature after being processed by the fully connected layer and the ReLU activation function; Represents the residual fusion feature.
Citation Information
Patent Citations
Time sequence prediction method based on decomposition mechanism and attention mechanism
CN116579447A
Sequence decomposition-subspace aggregation parameter prediction method and system
CN119106396A
Knowledge integration decomposition network for long-term time series prediction
CN119226932A
Fault early warning method for current transformer
CN119310516A
LSTM-SVR subway station temperature prediction method based on characteristic of multiple periods
WO2024077969A1
Cited By
Household appliance wire harness aging resistance prediction method
CN120449726A
Seasonal change monitoring method and system based on artificial intelligence
CN121092959A
STL and deep learning-based sequence prediction method
CN121167154A