Time Series Prediction Method and Device Based on Periodic Embedding and Multi-Scale Features

Through a method based on periodic embedding and multi-scale features, the limitations of the prior art in capturing periodic and frequency domain features are solved, and more efficient and accurate time series prediction is achieved.

CN119622319BActive Publication Date: 2025-06-03LUDONG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510156571.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-06-03
Estimated Expiration
2045-02-13

AI Technical Summary

Technical Problem

Existing time series prediction methods have limitations in capturing periodic and frequency domain features, and it is difficult to effectively adapt to the dynamic changes of data, especially in long-term time series prediction.

Method used

The periodic patterns in the subsequence are extracted by overlapping block processing and separated from the non-periodic residual features using a method based on periodic embedding and multi-scale features. Then, discrete wavelet transforms decompose the non-periodic residual into low-frequency and high-frequency components, and dynamically adjust the weights in combination with the wavelet channel attention mechanism. Finally, the implicit spatiotemporal dependencies in the time series are dynamically captured using adaptive adjacency matrix and graph convolutional networks, and the final prediction results are generated through residual connections and periodic corrections.

Benefits of technology

It improves the stability and accuracy of time series prediction, enhances the feature correlation and prediction accuracy of multivariate sequences, and ensures that the prediction results are more in line with the periodic changes of the data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119622319B_ABST
    Figure CN119622319B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of time series prediction, and specifically relates to a time series prediction method and device based on periodic embedding and multi-scale features. Preprocess the historical observation time series of the temperature change of the electrical transformer to obtain subsequences; extract the periodic patterns in the subsequences and calculate the aperiodic residual features of the subsequences; decompose the aperiodic residual features into low-frequency components and high-frequency components; calculate weights according to the low-frequency components and high-frequency components for weighted optimization to form multi-scale features; construct two embedding parameter matrices to generate an adaptive adjacency matrix; input the multi-scale features and the adaptive adjacency matrix into a graph convolutional network, perform residual connection on the features output by the graph convolutional network and the multi-scale features, and calculate the fused features; generate the predicted values of the time series based on the fused features; correct the predicted values of the time series, and output the corrected time series prediction results, thereby improving the stability and accuracy of the prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of time series prediction, and particularly relates to a time series prediction method and device based on periodic embedding and multi-scale features. Background Art

[0002] Time series forecasting (TSF) plays an important role in fields such as weather forecasting, traffic prediction, financial analysis, and energy management. From the perspective of feature decomposition, time series data usually contains trends, seasonality, periodicity, and other potential patterns. Especially in high-frequency fluctuation data, periodic features are crucial for improving the stability and accuracy of forecasting. However, existing methods have limitations in dealing with complex dynamic patterns and are difficult to effectively adapt to the dynamic changes of data. Especially in long time series forecasting, the intertwining of trends, periodicity, and frequency domain features makes modeling extremely challenging.

[0003] Traditional methods such as autoregressive (AR), exponential smoothing, and seasonal decomposition rely on manually set feature rules and are insufficient in capturing long-term dependencies and dealing with complex nonlinear dynamic changes. In recent years, deep learning methods have gradually been introduced into the field of time series prediction. For example, convolutional neural networks (CNNs) capture time series patterns by extracting local features, long short-term memory networks (LSTMs) model long-term dependencies through memory units, and gated recurrent units (GRUs) reduce computational complexity while maintaining the model's capabilities. These methods have made significant progress in the long sequence prediction task of electrical transformer temperature data in industrial substations, but still have deficiencies in modeling complex dynamic periodic features.

[0004] To make up for the above deficiencies, graph neural networks (GNNs) have attracted attention due to their modeling capabilities in spatial and temporal dependencies. Compared with traditional methods, GNNs effectively capture the causal relationships between time series nodes through a directed graph topology and dynamically adjust the adjacency relationships. However, most existing GNN methods focus on capturing static features such as trends and seasonality, and have limited modeling capabilities for periodic and frequency domain features. For example, methods such as Autoformer (Auto-Correlation Mechanism) handle long-term dependencies through trend decomposition, but are insufficient when facing multi-frequency features and complex periodic patterns.

[0005] Periodic feature modeling is another key challenge in time series prediction, reflecting the long-term dependencies of data. Traditional periodic modeling methods such as Fourier transform and moving average decomposition perform well in dealing with simple periodic patterns, but are limited in modeling complex periodic behaviors, especially when facing highly nonlinear and multi-frequency features, lacking flexibility and self-adaptability.

[0006] In summary, as the complexity of time series prediction tasks increases, the limitations of traditional methods in capturing periodic and frequency domain features become more apparent. Summary of the Invention

[0007] To overcome the problems in the prior art, the present invention proposes a time series prediction method and device based on periodic embedding and multi-scale features.

[0008] The technical solution of the present invention to solve the above technical problems is as follows:

[0009] In a first aspect, the present invention provides a time series prediction method based on periodic embedding and multi-scale features, including the following steps:

[0010] S1: Preprocess the historical observation time series of the temperature change of the electrical transformer in the industrial substation. The preprocessing includes overlapping block processing of the historical observation time series to obtain a number of subsequences of a fixed length;

[0011] S2: Extract the periodic patterns in the subsequences and calculate the aperiodic residual features of the subsequences to separate the periodic features and the aperiodic residual features;

[0012] S3: Decompose the aperiodic residual features into low-frequency components and high-frequency components through discrete wavelet transform;

[0013] S4: Calculate weights according to the importance of the low-frequency components and the high-frequency components to generate weighted and optimized frequency domain features, and fuse the weighted and optimized frequency domain features to obtain multi-scale features;

[0014] S5: Construct two embedding parameter matrices to generate an adaptive adjacency matrix;

[0015] S6: Input the multi-scale features and the adaptive adjacency matrix into a graph convolutional network to obtain the output features of the graph convolutional network;

[0016] S7: Perform residual connection on the output features of the graph convolutional network and the multi-scale features to calculate the fused features;

[0017] S8: Generate the prediction result of the time series based on the fused features, that is, the predicted value of the time series;

[0018] S9: Correct the predicted value of the time series;

[0019] S10: Output the time series prediction result after periodic correction.

[0020] Further, the overlapping block processing of the historical observation time series to obtain several subsequences of a fixed length includes: extracting subsequences of length w from the historical observation time series in a sliding window manner, with partial overlap between each window and the previous window, to obtain several subsequences of a fixed length.

[0021] Further, in S2, extracting the periodic patterns in the subsequences and calculating the aperiodic residual features of the subsequences to separate the periodic features and the aperiodic residual features, specifically including:

[0022] S201: Constructing a period embedding matrix P(t) based on the subsequence:

[0023] ;

[0024] where t is the current time step and τ is the set period length;

[0025] S202: Based on the period embedding matrix, calculating the aperiodic residual part to separate the periodic features from the subsequence. The calculation formula for the aperiodic residual part is as follows:

[0026] ;

[0027] where R(t) represents the aperiodic residual and X(t) represents the original value of the input subsequence.

[0028] Further, in S3, decomposing the aperiodic residual features into low-frequency components and high-frequency components through discrete wavelet transform, specifically including:

[0029] Performing multi-scale decomposition on the time series data of the aperiodic residual R(t) using discrete wavelet transform and selecting wavelet basis functions to convert the aperiodic residual R(t) into low-frequency components and high-frequency components. The low-frequency components represent the long-term trend and global change pattern of the time series, and the high-frequency components include capturing short-term fluctuations and detailed features. The specific calculation formula is:

[0030] ;

[0031] where represents the low-frequency wavelet basis function, represents the high-frequency wavelet basis function; represents the wavelet decomposition coefficient corresponding to the low-frequency wavelet basis function, represents the wavelet decomposition coefficient corresponding to the high-frequency wavelet basis function.

[0032] Further, in S4, weights are calculated based on the importance of the low-frequency components and high-frequency components to generate weighted and optimized frequency-domain features, and the weighted and optimized frequency-domain features are fused to obtain multi-scale features, which specifically include:

[0033] S401: Conduct statistical analysis on the channel values of the low-frequency components and high-frequency components to obtain the energy value of each channel;

[0034] ;

[0035] Among them, and respectively represent the salience of the low-frequency channel j and the high-frequency channel k, and N is the length of each channel;

[0036] S402: Based on the energy value of each channel, assign attention weights to each channel and normalize them; the formula for assigning attention weights to the channels is:

[0037] ;

[0038] Among them, and are the normalized weights of the low-frequency and high-frequency channels respectively;

[0039] S403: Apply the normalized attention weights to the low-frequency features and high-frequency features to perform weighted optimization on the results of the discrete wavelet transform. The formula for the weighted optimization is:

[0040] ;

[0041] Among them, represents the weighted and optimized low-frequency feature, represents the weighted and optimized high-frequency feature;

[0042] S404: Fuse the weighted and optimized low-frequency feature and the weighted and optimized high-frequency feature to obtain the multi-scale feature :

[0043] .

[0044] Further, in S5, construct two embedding parameter matrices to generate an adaptive adjacency matrix, which specifically includes:

[0045] Construct two embedding parameter matrices and , represents the input embedding of the time series node, represents the output embedding of the time series node;

[0046] Generate an initial adjacency matrix through the product of the two embedding parameter matrices :

[0047] ;

[0048] Perform normalization using the Sigmoid function to obtain the final adaptive adjacency matrix :

[0049] .

[0050] Furthermore, in S6, input the multi-scale features and the adaptive adjacency matrix into the graph convolutional network to obtain the output features of the graph convolutional network, specifically including:

[0051] S601: Symmetrically normalize the adaptive adjacency matrix to generate a propagation matrix , and its calculation formula is:

[0052] ;

[0053] where D is the degree matrix of the adjacency matrix , which is used to standardize the connection relationship of nodes;

[0054] S602: Use the propagation matrix and the multi-scale features to update the node features through graph convolution operation. The specific formula is:

[0055] ;

[0056] where represents the output features of the graph convolutional network, W is the learnable weight matrix of graph convolution, and ReLU is the activation function.

[0057] Furthermore, in S7, perform a residual connection between the output features of the graph convolutional network and the multi-scale features to calculate the fused features, specifically including:

[0058] Adopt a residual connection mechanism to add the multi-scale features and the output features after passing through the graph convolutional network: , where represents the fused features.

[0059] Furthermore, in S9, correct the predicted value of the time series, specifically including:

[0060] Use the periodic eigenvalue to correct the predicted value of the time series, that is, add the periodic eigenvalue to the predicted value of the time series.

[0061] In a second aspect, the present invention also provides a time series prediction device based on periodic embedding and multi-scale features, the device comprising: a block module, a periodic embedding module, a wavelet decomposition module, a wavelet channel attention module, an adaptive adjacency matrix generation module, a graph convolution module, a residual connection module, a backbone network prediction module, a periodic feature correction module, and an output module;

[0062] The block module is configured to preprocess the historical observation time series of the temperature change of the electric transformer in the industrial substation, and the preprocessing includes performing overlapping block processing on the historical observation time series to obtain a plurality of subsequences of a fixed length;

[0063] The periodic embedding module is configured to extract the periodic patterns in the subsequences and calculate the aperiodic residual features of the subsequences to separate the periodic features and the aperiodic residual features;

[0064] The wavelet decomposition module is configured to perform multi-scale wavelet decomposition on the aperiodic residual features, decomposing them into low-frequency components and high-frequency components;

[0065] The wavelet channel attention module is configured to calculate weights according to the importance of the low-frequency components and the high-frequency components to generate weighted and optimized frequency domain features; and fuse the weighted and optimized frequency domain features to obtain multi-scale features;

[0066] The adaptive adjacency matrix generation module is configured to generate an adaptive adjacency matrix through learning;

[0067] The graph convolution module is configured to input the multi-scale features and the adaptive adjacency matrix into a graph convolution network to obtain the output features of the graph convolution network;

[0068] The residual connection module is configured to perform residual connection on the output features of the graph convolution network and the multi-scale features to calculate the fused features;

[0069] The backbone network prediction module is configured to generate a prediction result of the time series, that is, the predicted value of the time series, by using the fused features;

[0070] The periodic feature correction module is configured to correct the prediction result in combination with the periodic embedding features;

[0071] The output module is configured to output the final prediction result after periodic correction.

[0072] Compared with the prior art, the present invention has the following technical effects:

[0073] (1) The present invention separates the periodic features from the aperiodic features, enabling the model to separately model the two types of features and improving the stability and accuracy of the prediction;

[0074] (2) The present invention decomposes the time series into low-frequency and high-frequency components through wavelet transform, combines the wavelet channel attention mechanism to dynamically adjust the weights, strengthens the expression of key features, and suppresses noise;

[0075] (3) The present invention utilizes the adaptive adjacency matrix and the graph convolutional network to dynamically capture the implicit spatio-temporal dependence relationship in the time series, enhance the feature correlation, and improve the prediction accuracy of the multivariate series;

[0076] (4) The present invention retains the original features through residual connection to prevent information degradation. At the same time, periodic correction makes the prediction results more in line with the periodic changes of the data, further improving the prediction accuracy and robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0078] Figure 1 It is a flowchart of the time series prediction method based on periodic embedding and multi-scale features of the present invention;

[0079] Figure 2 It is a prediction fitting diagram of the temperature change curve of the electrical transformer;

[0080] Figure 3 It is a schematic structural diagram of the time series prediction method based on periodic embedding and multi-scale features of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0081] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following will describe in detail the specific implementation manners, structures, features and effects of the technical solutions proposed according to the present invention in combination with the accompanying drawings and preferred embodiments. The specific features, structures or characteristics in one or more embodiments can be combined in any suitable form. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.

[0082] In one embodiment of the present invention, referring to Figures 1 - 3 , a time series prediction method based on periodic embedding and multi-scale features is provided, and the method includes:

[0083] S1: Preprocess the historical observation time series of the temperature change of the electrical transformer in the industrial substation. The preprocessing includes performing overlapping block processing on the historical observation time series to obtain a number of subsequences of a fixed length;

[0084] S2: Extract the periodic patterns in the subsequences and calculate the aperiodic residual features of the subsequences to separate the periodic features and the aperiodic residual features;

[0085] S3: Decompose the aperiodic residual features into low-frequency components and high-frequency components through discrete wavelet transform;

[0086] S4: Calculate the weights according to the importance of the low-frequency components and the high-frequency components to generate weighted and optimized frequency-domain features; fuse the weighted and optimized frequency-domain features to obtain multi-scale features;

[0087] S5: Construct two embedding parameter matrices to generate an adaptive adjacency matrix;

[0088] S6: Input the multi-scale features and the adaptive adjacency matrix into the graph convolutional network to obtain the output features of the graph convolutional network;

[0089] S7: Perform residual connection on the output features of the graph convolutional network and the multi-scale features to calculate the fused features;

[0090] S8: Use the fused features to generate the prediction result of the time series, that is, the predicted value of the time series;

[0091] S9: Use the periodic feature values to correct the predicted value of the time series;

[0092] S10: Output the prediction result of the time series after periodic correction.

[0093] The following expands each of the above steps in detail:

[0094] S1: Preprocess the historical observation time series X of the temperature change of the electrical transformer in the industrial substation. The preprocessing includes performing overlapping block processing on the historical observation time series to obtain a number of subsequences of a fixed length; the purpose is to retain the continuity of the time series and generate fixed-length subsequences with temporal context information.

[0095] Specifically, the preprocessing includes: extracting subsequences of length w from the historical observation time series in a sliding window manner, with each window partially overlapping with the previous window, for example, the overlapping step size is s, to capture more local patterns and reduce boundary effects. The preprocessing divides the historical observation time series into multiple overlapping subsequences, laying a foundation for subsequent feature extraction and effectively avoiding information loss.

[0096] In this embodiment, a specific implementation of S1 can be as follows:

[0097] S101: The historical observation time series X is segmented by means of a sliding window with a fixed window length w and a step size s to extract subsequences centered on the time step t, and the range of each subsequence is , where , and T is the total length of the time series data.

[0098] By means of the sliding window, multiple adjacent and overlapping subsequences are generated to ensure the continuity of the sequence and the retention of context information, thus laying a foundation for subsequent feature extraction.

[0099] S102: The subsequences are segmented into several small segments of a fixed length.

[0100] Each subsequence is centered on each time step t + i and segmented with a fixed segment length l into several overlapping small segments, and the range of each small segment is , where l represents the length of the small segment, i = −k,..., k, and k is determined by the length after window subdivision. This segmentation operation can capture local feature patterns in the subsequence and at the same time reduce the errors that may be introduced due to changes in the window scale.

[0101] Specifically, the segmentation operation is:

[0102] ;

[0103] In the above formula, represents the subsequence, , represent the small segments. Through this operation, local patterns in the subsequence are captured, and at the same time, the errors caused by scale changes are reduced.

[0104] S2: Extract the periodic patterns in the subsequences and calculate the aperiodic residual features of the subsequences to separate the periodic features and the aperiodic residual features.

[0105] The subsequences are input into the periodic embedding module, and the period length is defined according to the characteristics of different data sets, such as daily period, weekly period, etc., to extract the periodic patterns in the subsequences.

[0106] The periodic embedding module models the implicit periodic changes in the sequence by constructing a periodic embedding matrix. The periodic embedding matrix maps the time points of the input subsequence through learning parameters to capture the change rules within the periodic time span. Next, the aperiodic residual part is calculated, that is, the periodic pattern is separated from the sequence to retain the residual information of the non-periodic part in the data. Through this operation, the periodic features and the aperiodic features are clearly separated, enabling the model to process these two types of features separately.

[0107] In this embodiment, a specific implementation manner of S2 may be as follows:

[0108] S201: Construct a periodic embedding matrix based on the subsequence.

[0109] Input the subsequence into the periodic embedding module, and set the period length according to the characteristics of the dataset τ , and perform a periodic mapping on the time step t of the input subsequence.

[0110] The periodic embedding module constructs the periodic embedding matrix P(t) through the following formula:

[0111] ;

[0112] where t is the current time step and τ is the set period length.

[0113] The periodic embedding matrix captures the implicit periodic change patterns in the sequence, describes the time characteristics under different periods through parametric sine and cosine functions, and lays a foundation for subsequent feature decomposition.

[0114] S202: Based on the periodic embedding matrix, calculate the non-periodic residual part and separate the periodic features from the subsequence.

[0115] The calculation formula of the non-periodic residual is as follows:

[0116] ;

[0117] where X(t) represents the original value of the input subsequence, and P(t) is the periodic eigenvalue generated by the periodic embedding matrix. Through the above calculation, the periodic pattern P(t) is removed from the original sequence, and only the non-periodic residual feature R(t) is retained. This operation ensures that the periodic features and non-periodic features are clearly separated, providing a more independent feature input for subsequent modeling.

[0118] S3: Decompose the non-periodic residual feature into a low-frequency component and a high-frequency component through discrete wavelet transform.

[0119] Wavelet transform is a mathematical tool for multi-scale analysis, which can decompose the sequence signal into different frequency components. By using discrete wavelet transform (DWT), the non-periodic residual is decomposed into two main components: a low-frequency component and a high-frequency component. During the decomposition process, an appropriate wavelet basis, such as Daubechies or Haar, is selected to ensure capturing the important patterns of the data in the frequency domain while reducing the decomposition error. The purpose of this stage is to analyze the multi-scale features of the data at the frequency domain level and provide support for subsequent feature enhancement and analysis.

[0120] In this embodiment, a specific implementation manner of S3 may be:

[0121] Perform multi-scale decomposition on the time series data using the discrete wavelet transform for the non-periodic residual R(t), select an appropriate wavelet basis function, and convert the non-periodic residual R(t) into a low-frequency component and a high-frequency component. The low-frequency component represents the long-term trend and global change pattern of the time series, and the high-frequency component includes capturing short-term fluctuations and detailed features. The specific calculation formula is:

[0122] ;

[0123] Wherein, represents the wavelet basis function of the low frequency, represents the wavelet basis function of the high frequency; represents the wavelet decomposition coefficient corresponding to the wavelet basis function of the low frequency, represents the wavelet decomposition coefficient corresponding to the wavelet basis function of the high frequency.

[0124] S4: Calculate weights according to the importance of the low-frequency component and the high-frequency component to generate weighted and optimized frequency domain features; fuse the frequency domain features obtained by weighted optimization to obtain multi-scale features.

[0125] After completing the wavelet decomposition, perform weighting according to the significance of the low-frequency component and the high-frequency component to strengthen the expression of important frequency domain features. Specifically, first perform statistical analysis on the channel values of the low-frequency component and the high-frequency component, calculate the energy value of each channel to measure its importance, and the energy value is such as the absolute value mean or the sum of squares; based on the energy value, assign attention weights to each channel, and enhance the channel features that contribute more to the target task in the low frequency and the high frequency through weighted operations. The enhanced features not only retain the frequency domain information but also explicitly emphasize the key patterns, enabling the model to focus on the important feature regions.

[0126] In this embodiment, a specific implementation manner of S4 may be:

[0127] S401: Perform statistical analysis on the channel values of the low-frequency component and the high-frequency component to obtain the energy value of each channel.

[0128] The energy calculation methods include the absolute value mean or the sum of squares of the channels, which are respectively expressed as:

[0129] ;

[0130] Wherein, and respectively represent the significance of the low-frequency channel j and the high-frequency channel k, and N is the length of each channel.

[0131] S402: Based on the energy values of each channel, assign attention weights to each channel and normalize them.

[0132] The attention weight assignment adopts a normalization method to ensure that the weight values are in the range of [0, 1], which is convenient for subsequent weighting operations. The attention weight calculation formula is:

[0133] ;

[0134] where and are the normalized weights of the low-frequency and high-frequency channels respectively. These weights reflect the relative importance of each channel in feature expression, and channels with higher significance will be assigned larger weights.

[0135] S403: Apply the normalized attention weights to the low-frequency features and high-frequency features to perform weighted optimization on the results of the discrete wavelet transform.

[0136] The weighted operation formula is:

[0137] ;

[0138] where represents the low-frequency features after weighted optimization, represents the high-frequency features after weighted optimization.

[0139] Through the weighted operation, it is possible to strengthen the channel features that contribute more to the target task while suppressing noise or unimportant features. The low-frequency and high-frequency features after weighted optimization not only retain the information of the original frequency domain features but also explicitly emphasize the key patterns, providing more effective inputs for subsequent model training.

[0140] S404: Fuse the low-frequency features after weighted optimization and the high-frequency features after weighted optimization to obtain multi-scale features :

[0141] .

[0142] S5: Construct two embedding parameter matrices to generate an adaptive adjacency matrix.

[0143] First, construct two parameter matrices through learning, and then generate an initial adjacency matrix through the matrix product of the parameter matrices. To limit the values of the adjacency matrix within a reasonable range, further use the Sigmoid function to normalize the generated matrix to ensure that all values are within the range of [0, 1]. This adaptive adjacency matrix can dynamically capture the topological relationships between time series nodes, especially those non-linear or implicit dependencies. Continuously optimize the initial adjacency matrix through learning, and finally generate an adaptive adjacency matrix that can reflect the characteristic dependency structure of the time series.

[0144] In this embodiment, a specific implementation of S5 can be:

[0145] S501: Construct two embedding parameter matrices and , represents the input embedding of the time series node, represents the output embedding of the time series node; generate an initial adjacency matrix through the product of the two embedding parameter matrices :

[0146] ;

[0147] To ensure that the matrix values are within a reasonable range, use the Sigmoid function for normalization processing to obtain the final adaptive adjacency matrix :

[0148] ;

[0149] The adaptive adjacency matrix can dynamically capture the implicit relationships between time series nodes, and optimize the embedding matrices and through training, and gradually and accurately reflect the dependency structure of the time series, providing input support for subsequent graph convolutional network modeling.

[0150] S6: Input the multi-scale features and the adaptive adjacency matrix into the graph convolutional network to obtain the output features of the graph convolutional network.

[0151] Input the multi-scale features and the adaptive adjacency matrix into the graph convolutional network (GCN) to capture the complex dependencies between time series nodes. First, perform symmetric normalization on the adaptive adjacency matrix to generate a propagation matrix to ensure the stability of the information propagation process; then, calculate the update of the node features through graph convolution operations. The graph convolutional network can propagate information through the adjacency matrix and model the implicit long-term and short-term dependencies in the time series.

[0152] In this embodiment, a specific implementation of S6 can be:

[0153] S601: Symmetrically normalize the adaptive adjacency matrix to generate a propagation matrix , and its calculation formula is:

[0154] ;

[0155] where D is the degree matrix of the adjacency matrix , which is used to standardize the connection relationship of nodes to ensure the numerical stability of the information propagation process;

[0156] S602: Use the propagation matrix and multi-scale features to update the node features through graph convolution operations. The specific formula is:

[0157] ;

[0158] where W is the learnable weight matrix of graph convolution, and ReLU is the activation function. This operation propagates information between nodes through the adjacency relationship, models the long-term and short-term dependencies of the time series, and provides deeper feature representations for subsequent tasks.

[0159] S7: Perform a residual connection between the output features of the graph convolutional network and the multi-scale features to calculate the fused features.

[0160] To ensure the integrity of the model during the feature processing and no loss of information, a residual connection mechanism is adopted to add the multi-scale features and the output features of the graph convolutional network : ′, indicating the fused features obtained. Residual connections are crucial in the deep structure of the model. It can alleviate the degradation problem of information during multi-layer propagation, while retaining the key information of the original features and improving the learning stability of the model.

[0161] S8: Generate the prediction result of the time series based on the fused features, that is, the predicted value of the time series.

[0162] Input the fused feature R''' into the backbone network for final prediction. If a multi-layer perceptron (MLP) is used, further extract the feature patterns through multi-layer non-linear mapping; if a linear layer is used, directly complete the mapping from the features to the predicted values. Finally, output the time series prediction result Y, with the goal of generating accurate predicted values based on the high-quality features extracted previously.

[0163] S9: Use the periodic eigenvalue to correct the predicted value of the time series.

[0164] After the backbone network outputs the prediction result Y, the prediction result is corrected using the periodic feature P, and the calculation formula is:

[0165] Y' = Y + P;

[0166] This correction operation can reintroduce the periodic pattern into the prediction result, ensuring that the model can capture both the periodic change law of the data and retain the key trends and detailed features in the time series.

[0167] Output the final time series prediction result Y' after periodic correction. This result combines multi-scale frequency domain information, periodic features, and the dependencies modeled by the graph convolutional network, and can achieve excellent performance in terms of prediction accuracy and pattern expression ability.

[0168] S10: Output the time series prediction result Y' after periodic correction.

[0169] This result combines multi-scale frequency domain information, the periodic features of the time series, and the dependencies modeled by the graph convolutional network, and can achieve excellent effects in terms of prediction accuracy and data pattern preservation.

[0170] Refer to Figure 2 , when predicting the temperature of an electrical transformer with a long sequence, the present invention can also ensure accuracy and effective fitting effect.

[0171] Based on the same inventive concept, an embodiment of the present invention also provides a time series prediction device based on periodic embedding and multi-scale features for implementing the above-mentioned time series prediction method based on periodic embedding and multi-scale features. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more system embodiments provided below can refer to the limitations on the time series prediction method based on periodic embedding and multi-scale features in the above text, and will not be repeated here.

[0172] In this embodiment, a time series prediction device based on periodic embedding and multi-scale features is provided. The device includes: a block module, a periodic embedding module, a wavelet decomposition module, a wavelet channel attention module, an adaptive adjacency matrix generation module, a graph convolutional module, a residual connection module, a backbone network prediction module, a periodic feature correction module, and an output module.

[0173] Chunking module, used to preprocess the historical observation time series of the temperature change of the electrical transformer in the industrial substation. The preprocessing includes overlapping chunking of the historical observation time series to obtain a number of subsequences of fixed length. Specifically, the historical observation time series X is chunked according to a fixed window length w and a step size s, and the time series is divided into subsequences with overlapping parts, generating multiple segments of overlapping subsequences of fixed length to ensure the continuity of the sequence and provide context information support for subsequent feature extraction.

[0174] Period embedding module, used to extract the periodic patterns in the subsequences and calculate the aperiodic residual features of the subsequences to separate the periodic features and the aperiodic residual features. Specifically, a period length τ is defined according to the characteristics of the time series, and the periodic features of the time series are modeled using the period embedding matrix P(t); the period embedding matrix maps the time step t through parameterized sine and cosine functions, captures the periodic patterns, and generates the aperiodic residual feature R = X - P to separate the periodic features and the aperiodic features.

[0175] Wavelet decomposition module, used to perform multi-scale wavelet decomposition on the aperiodic residual feature R, decomposing it into a low-frequency component and a high-frequency component. Specifically, through discrete wavelet transform (DWT), the aperiodic residual feature R is decomposed into a low-frequency component cA and a high-frequency component cD. The low-frequency component cA reflects the long-term trend and global change pattern of the time series, and the high-frequency component cD represents the local fluctuations and detailed features; the frequency domain features obtained by weighted optimization are fused to obtain multi-scale features.

[0176] Wavelet channel attention module, used to calculate weights according to the importance of the low-frequency component and the high-frequency component to generate weighted optimized frequency domain features. Specifically, through statistical analysis of the low-frequency component cA and the high-frequency component cD, the energy value of each channel (such as the absolute value mean or the sum of squares) is calculated, and the channel attention weights are generated according to the energy value and ; the attention weights are applied to the low-frequency component cA and the high-frequency component cD to generate weighted optimized frequency domain features.

[0177] Adaptive adjacency matrix generation module, used to generate an adaptive adjacency matrix through learning . This module constructs two embedding parameter matrices and , calculates the initial adjacency matrix ; through the Sigmoid function, is normalized to generate an adaptive adjacency matrix with a value range in [0, 1] to dynamically capture the topological relationship between the time series nodes.

[0178] The graph convolution module is used to process multi-scale features and the adaptive adjacency matrix as inputs into the graph convolutional network (GCN) to obtain the output features of the graph convolutional network, capturing the complex dependencies between nodes in the time series. Specifically, first is symmetrically normalized to generate the propagation matrix , and then the feature update is calculated through graph convolution operations:

[0179] .

[0180] The residual connection module is used to perform a residual connection between the output features of the graph convolutional network and the multi-scale features to calculate the fused features . This module can alleviate the problem of information degradation in deep networks while retaining key original features.

[0181] The backbone network prediction module is used to generate the prediction result of the time series, i.e., the predicted value of the time series, using the fused features. Specifically, the fused feature R''' is input into the backbone network for final prediction. The backbone network can be a multi-layer perceptron (MLP) or a linear layer: if it is an MLP, deep features are further extracted through non-linear transformation; if it is a linear layer, the mapping from features to predicted values is directly completed. The prediction result Y of the time series is output.

[0182] The periodic feature correction module is used to correct the prediction result Y by combining the periodic embedding features. By adding the periodic feature P back to the prediction result, the final output value is calculated:

[0183] Y′ = Y + P.

[0184] This module ensures that the model output can reflect the periodic pattern of the data while retaining the key trends in the time series.

[0185] The output module is used to output the finally predicted result Y' after periodic correction. This result synthesizes multi-scale frequency domain information, periodic features, and the dependencies modeled by the graph convolutional network, having significant advantages in prediction accuracy and the retention of data patterns.

[0186] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included within the protection scope of the present application.

Claims

1. A time series prediction method based on period embedding and multi-scale features, characterized in that: The following steps are involved: S1: preprocessing the input historical observation time series of temperature changes of electric transformers in industrial substations, wherein the preprocessing includes overlapping and blocking the historical observation time series to obtain a plurality of subsequences of fixed length; S2: extract the periodic pattern in the subsequence and calculate the non-periodic residual features of the subsequence to separate the periodic features and the non-periodic residual features; S3: decomposing the non-periodic residual feature into a low-frequency component and a high-frequency component by discrete wavelet transform; S4: Calculate the weights according to the importance of low-frequency components and high-frequency components, generate weighted optimized frequency domain features, and fuse the weighted optimized frequency domain features to obtain multi-scale features; S5: construct two embedding parameter matrices and generate an adaptive adjacency matrix; S6: Input the multi-scale features and the adaptive adjacency matrix into the graph convolutional network to obtain the output features of the graph convolutional network; S7: Perform residual connection between the output features of the graph convolutional network and the multi-scale features to calculate the fused features; S8: Generate the prediction result of the time series based on the fused features, that is, the predicted value of the time series; S9: Correct the predicted value of the time series; S10: Output the time series prediction results after periodic correction.

2. According to claim 1, a time series prediction method based on period embedding and multi-scale features is characterized in that: The overlapping block processing of the historical observation time series to obtain a plurality of subsequences of fixed lengths includes: extracting subsequences of length w from the historical observation time series in a sliding window manner, each window partially overlapping with the previous window, and obtaining a plurality of subsequences of fixed lengths.

3. The time series prediction method based on period embedding and multi-scale features according to claim 1 is characterized in that: S2 extracts the periodic pattern in the subsequence and calculates the non-periodic residual features of the subsequence to separate the periodic features and the non-periodic residual features, specifically including: S201: Construct a periodic embedding matrix P(t) based on the subsequence: ; Among them, t is the current time step, and τ is the set cycle length; S202: Based on the periodic embedding matrix, a non-periodic residual part is calculated to separate the periodic feature from the subsequence. The calculation formula of the non-periodic residual part is as follows: ; Among them, R(t) represents the non-periodic residual, and X(t) represents the original value of the input subsequence.

4. The time series prediction method based on period embedding and multi-scale features according to claim 3 is characterized in that: In S3, the non-periodic residual feature is decomposed into a low-frequency component and a high-frequency component by discrete wavelet transform, which specifically includes: The non-periodic residual R(t) is decomposed into multi-scales using discrete wavelet transform on the time series data, and the wavelet basis function is selected to convert the non-periodic residual R(t) into low-frequency components and high-frequency components. The low-frequency components represent the long-term trend and global change pattern of the time series, and the high-frequency components capture short-term fluctuations and detail features. The specific calculation formula is: ; in, represents the low-frequency wavelet basis function, Wavelet basis functions representing high frequencies; Represents the wavelet decomposition coefficients corresponding to the low-frequency wavelet basis function, The wavelet decomposition coefficients corresponding to the wavelet basis functions representing high frequencies.

5. The time series prediction method based on period embedding and multi-scale features according to claim 4 is characterized in that: In S4, weights are calculated according to the importance of low-frequency components and high-frequency components to generate weighted optimized frequency domain features, and the frequency domain features obtained by weighted optimization are fused to obtain multi-scale features, which specifically includes: S401: Performing statistical analysis on the channel values ​​of the low-frequency component and the high-frequency component to obtain the energy value of each channel; ; in, and Respectively represent the significance of low-frequency channel j and high-frequency channel k, N is the length of each channel; S402: Based on the energy value of each channel, an attention weight is assigned to each channel and normalized; the channel-assigned attention weight calculation formula is: ; in, and are the normalized weights of low-frequency and high-frequency channels, respectively; S403: Apply the normalized attention weights to the low-frequency features and high-frequency features, and perform weighted optimization on the discrete wavelet transform results. The weighted optimization formula is: ; in, represents the low-frequency features after weighted optimization, Represents the high-frequency features after weighted optimization; S404: The weighted optimized low-frequency features are combined with the weighted optimized high-frequency features to obtain multi-scale features : 。 6. The time series prediction method based on period embedding and multi-scale features according to claim 5 is characterized in that: In S5, two embedding parameter matrices are constructed to generate an adaptive adjacency matrix, which specifically includes: Construct two embedding parameter matrices and , represents the input embedding of the time series node, represents the output embedding of a time series node; The initial adjacency matrix is ​​generated by the product of the two embedding parameter matrices : ; Use the Sigmoid function to normalize and get the final adaptive adjacency matrix : 。 7. The time series prediction method based on period embedding and multi-scale features according to claim 6 is characterized in that: In S6, the multi-scale features and the adaptive adjacency matrix are input into the graph convolutional network to obtain the output features of the graph convolutional network, which specifically includes: S601: Symmetrically normalize the adaptive adjacency matrix to generate a propagation matrix , and its calculation formula is: ; Where D is the adjacency matrix The degree matrix is ​​used to standardize the connection relationship of nodes; S602: Using the propagation matrix and multi-scale features , update the node features through graph convolution operation, the specific formula is: ; in, Represents the output features of the graph convolutional network, W is the learnable weight matrix of the graph convolution, and ReLU is the activation function.

8. The time series prediction method based on period embedding and multi-scale features according to claim 7 is characterized in that: In S7, the output features of the graph convolutional network are residually connected with the multi-scale features to calculate the fused features, which specifically includes: Using the residual connection mechanism, multi-scale features And the output features of the graph convolutional network Addition: ,in, Represents the fused features.

9. The time series prediction method based on period embedding and multi-scale features according to claim 8 is characterized in that: The predicted value of the time series is corrected in S9, specifically including: The predicted value of the time series is corrected using the periodic characteristic value, that is, the periodic characteristic value is added to the predicted value of the time series.

10. A time series prediction device based on period embedding and multi-scale features, characterized in that: The device comprises: a block module, a period embedding module, a wavelet decomposition module, a wavelet channel attention module, an adaptive adjacency matrix generation module, a graph convolution module, a residual connection module, a backbone network prediction module, a period feature correction module and an output module; The block module is used to preprocess the input historical observation time series of temperature changes of electric transformers in industrial substations, wherein the preprocessing includes overlapping block processing of the historical observation time series to obtain a plurality of subsequences of fixed length; The periodic embedding module is used to extract the periodic pattern in the subsequence and calculate the non-periodic residual features of the subsequence to separate the periodic features and the non-periodic residual features; The wavelet decomposition module is used to perform multi-scale wavelet decomposition on the non-periodic residual features to decompose them into low-frequency components and high-frequency components; The wavelet channel attention module is used to calculate weights according to the importance of low-frequency components and high-frequency components to generate weighted optimized frequency domain features; the frequency domain features obtained by weighted optimization are fused to obtain multi-scale features; The adaptive adjacency matrix generation module is used to generate an adaptive adjacency matrix through learning; The graph convolution module is used to input multi-scale features and adaptive adjacency matrix into the graph convolution network to obtain graph convolution network output features; The residual connection module is used to perform residual connection on the output features of the graph convolutional network and the multi-scale features to calculate the fused features; The backbone network prediction module is used to generate a prediction result of the time series, that is, a predicted value of the time series, by using the fused features; The period feature correction module is used to correct the prediction result in combination with the period embedding feature; The output module is used to output the final prediction result after periodic correction.

Citation Information

Patent Citations

  • Multivariable time series prediction method for multi-scale adaptive graph learning

    CN114169394A

  • Convolutional sparse self-attention-based irrigation area water demand estimation method, equipment and medium

    CN118395108A