CNN-LSTM-based multi-scale feature fusion period method
By extracting multi-scale features and learning long-term dependencies through the CNN-LSTM model, the problem of low cycle recognition accuracy of a single deep learning model in complex time series data is solved, and accurate cycle analysis of the power system is achieved.
Patent Information
- Application Number
- CN202510755357.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-16
AI Technical Summary
The existing single deep learning model is difficult to effectively integrate local and global information of different scales, resulting in low cycle recognition accuracy of complex time series data, which makes it difficult to support the optimized scheduling of power systems and the absorption of new energy.
A multi-scale feature fusion method based on CNN-LSTM is adopted. The local features of multi-scale subsequences are extracted through parallel convolutional networks. The long-term dependencies are learned by LSTM to generate multi-scale comprehensive feature vectors. The fully connected layer and autocorrelation function are used to verify the period recognition results.
It achieves accurate capture of multi-scale periodic patterns of complex time series data, improves the accuracy and reliability of period identification, and provides an accurate basis for optimized scheduling of power systems and new energy consumption.
Smart Images

Figure CN120654186A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of power system and distribution system control, and relates to a periodic modeling and feature learning method for complex time series data. More specifically, it relates to a multi-scale feature fusion periodic method based on CNN-LSTM. Background Art
[0002] Cycle analysis of time series is a core method for revealing inherent patterns in data and is widely used in fields such as power system load forecasting, financial market volatility analysis, and meteorological pattern mining. Traditional cycle analysis methods rely on manually designed features and struggle to capture the complex patterns of nonlinear and non-stationary time series data. Single deep learning models have significant limitations: while CNNs can extract local features, they lack the ability to model long-term dependencies; while LSTMs excel at capturing long-term dependencies, they lack the ability to capture multi-scale local details.
[0003] The existing single deep learning model cannot effectively integrate local and global information at different scales, resulting in low cycle identification accuracy and making it difficult to support key applications such as power system optimization and scheduling and new energy consumption. Summary of the Invention
[0004] Technical problem: The purpose of this invention is to provide a multi-scale feature fusion periodicity method based on CNN-LSTM, which can accurately capture the multi-scale periodic patterns of time series and effectively improve the periodicity recognition accuracy and reliability of complex time series data.
[0005] Technical solution: The present invention is implemented by the following technical means:
[0006] A multi-scale feature fusion cycle method based on CNN-LSTM, the method comprises the following steps:
[0007] 1) For the original time series X, X∈R T×D Perform standardization to generate a standardized time series, and use at least two sliding windows of different window sizes to sample the standardized time series to generate multi-scale subsequences;
[0008] 2) The local features of the multi-scale subsequences are extracted through a parallel convolutional network. The local features of the multi-scale subsequences are compressed by global pooling and then spliced into a multi-scale comprehensive feature vector;
[0009] 3) Arrange the multi-scale integrated feature vectors in the original time order and input them into the LSTM layer to learn the time sequence to capture the long-term dependencies in the time series;
[0010] 4) The global features output by LSTM are input into the fully connected layer, and the period recognition results are generated through the activation function;
[0011] 5) Use the autocorrelation function to verify the correctness of the output period from the time domain or frequency domain.
[0012] Furthermore, the original time series X, X∈R T×D Perform standardization to generate a standardized time series, and use at least two sliding windows of different window sizes to sample the standardized time series to generate multi-scale subsequences, including:
[0013] 11) For the original time series X, X∈R T×D Perform standardization to generate a standardized time series X', including:
[0014]
[0015] Where μ is the mean of the sequence; σ' is the standard deviation of the sequence;
[0016] 12) At least two sliding windows of different window sizes are used to sample the standardized time series, including:
[0017] Design K windows of size W, W={w1,w2,...,w K}, perform overlapping sliding sampling on the standardized time series X' generated in step 11):
[0018]
[0019] Where w1, w2, ..., w K Represent the 1st to kth windows respectively; Under the kth window width, the subsequence (or "window fragment") intercepted from position i, the superscript k represents the index of the window width (different k corresponds to different window widths w k ), the subscript i indicates the starting position of the interception; X' represents the original sequence (or data carrier such as signal, time series, etc.), which is the intercepted "mother sequence"; X'[i:i+w k ] represents mathematical sequence interception, which means starting from the i-th position of X', the interception length is w k Subsequence (including position i, to i+w k -1 position, left closed and right open logic); i is the starting position index of the window interception, the value range is i=1,2,...,Tw k +1, to ensure that the window does not exceed the bounds (T is the total length of the original sequence X'); w kis the width (length) of the kth window, that is, the length of each subsequence intercepted, and different k corresponds to different window widths (such as window length w1 when k = 1, window length w2 when k = 2, etc.); T represents the total length of the original sequence X' (or the total number of samples, total duration, etc., depending on the specific scenario); k represents the index of the window width, and the value k = 1, 2, ..., K, indicating that there are K windows of different widths for intercepting subsequences; K represents the total number of window widths, that is, there are a total of K different lengths w1, w2, ..., w K The window is captured.
[0020] Generate each scale subsequence X (k) , all scale subsequences X (k) Constructing a multi-scale subsequence set
[0021] where N k =Tw k +1.
[0022] Furthermore, the extracting local features of the multi-scale subsequences by using a parallel convolutional network includes:
[0023] 21) According to step 12), for each scale subsequence X (k) Use M convolution kernels Perform convolution:
[0024]
[0025] Where, Represents the final output (result after activation), the subscript i usually represents the "sample / sequence position index" (such as the i-th step of the time series, the i-th sample in the batch); the k in the superscript (k, m) is generally the "category / level index of the input feature / sequence", and m is the "index of the current calculation channel / neuron / output dimension" (to distinguish different transformation branches); ReLU() is the ReLU activation function (Rectified Linear Unit), the formula is ReLU(x) = max(0,x), which is responsible for introducing nonlinearity to allow the model to fit complex relationships; l m -1 summation upper limit, l m Generally means "the input length / dimension of the mth transform channel" (i.e. for a length of l m The input is weighted summed, with indexes from 0 to l m -1); It represents a weighted summation of the input dimension / length (similar to convolution and fully connected linear transformation), and t is the "dimension / position index" of the summation. Indicates input data. The i (sample / sequence position) of the input is represented by the subscript t, which is the index of the input in the dimension / length direction (corresponding to the summed t), and the superscript (k) represents the category / level identifier of the input (to distinguish different input sources). Represents the weight parameter, subscript t and input The dimension t corresponds to the "weight of the t-th dimension input";
[0026] The superscript (m) represents the weight set of the mth channel / neuron / output dimension (different m corresponds to different weights); b (m) Represents the bias parameter, which adds an offset to the linear transformation result to improve the model's expressiveness. The superscript (m) corresponds to the bias of the mth channel / neuron.
[0027] 22) The local features of the multi-scale subsequences are compressed by global pooling and then spliced into a multi-scale comprehensive feature vector, including:
[0028] According to step 21), the convolution output is globally average pooled and compressed into a feature vector F (k) , and splice to get the multi-scale comprehensive feature vector F CNN =[F (1) ,...,F (K) ].
[0029] Furthermore, the multi-scale integrated feature vectors are arranged in the original time order and input into the LSTM layer to learn the time sequence to capture the long-term dependency in the time series, including:
[0030] 31) According to step 22) F CNN Reconstructed into a time-aligned feature sequence, S∈R T′×(K×M) , where T′ is the time step, K is the number of multi-scales, and M is the single-scale feature dimension;
[0031] 32) According to step 31), long-term dependencies are captured by updating the LSTM cell state:
[0032]
[0033] Where S t represents the input at the current moment (the features of the t-th step in the sequence, such as the word vector of the text, the observation value of the time series); h t-1 Represents the hidden state at the previous moment (the "memory summary" output by the previous step of LSTM, which conveys sequence history information); f t Represents the output of the forget gate (controls the "last moment cell state c t-1 How much to forget”, the closer the value is to 1, the more it will retain); i t Indicates the input gate output (controls the "current input How much to keep in the cell state? The closer the value is to 1, the more it will keep.t Indicates the output gate output (controls the "current cell state c t How much to output to the hidden state h t ", the closer the value is to 1, the more output it will produce); Represents the candidate cell state (“temporary memory” of the current input, not gated and filtered); c t Represents the current cell state ("long-term memory" after being updated by the forget gate and input gate, conveying the core information of the sequence); W f 、W i 、W c 、W o Respectively represent the weight matrix (corresponding to the linear transformation parameters of the forget gate, input gate, candidate cell, and output gate, respectively, model training and learning); b f 、b i 、b c 、b o They represent bias terms (adding offsets to linear transformations to improve model expression and model training and learning); ReLU() is the ReLU activation function (commonly used for gating, with non-negative output, controlling information "on / off"); tanh() represents the tanh activation function (commonly used for cell states, with an output range of [-1,1], scaling information amplitude); ⊙ represents element-by-element multiplication (gating and state are element-wise multiplied to achieve "information filtering").
[0034] where f t ,i t , o t is the gate vector, x t Input feature at time t, dimension is D; h t is the hidden state at time t, with a dimension of H, used to represent the temporal characteristics of the current moment; c t is the cell state at time t, with a dimension of C, used to store long-term dependency information; W f ,W i ,W c ,W o are the weight matrices of the forget gate, input gate, cell state update, and output gate, respectively, with dimensions of (H+D)×H; b f ,b i ,b c ,b o The bias term of the corresponding gate; ⊙ is the element-wise multiplication, which is used for weighted screening of the cell state by the gating signal; the dimension H is used to adjust the threshold of the gating signal.
[0035] Furthermore, the global features output by the LSTM are input into the fully connected layer, and the period recognition result is generated through the activation function, including:
[0036] The classification task uses the Softmax function to output the cycle probability:
[0037]
[0038] Where W p Fully connected layer weight matrix; F LSTM is the time series feature vector output by LSTM; b p is the bias term of the fully connected layer;
[0039] The regression task outputs the periodic parameters through the linear layer
[0040]
[0041] Update parameters through Adam optimizer:
[0042]
[0043] Where θ represents the model parameters (such as the weight W and bias b of the neural network, and the optimization goal is to find θ that minimizes the loss); t represents the iteration step (the tth round of parameter update, starting from the initial t = 0); η represents the learning rate (which controls the "step size" of each parameter update, which is manually set); Indicates a minimum value (such as 10 -8 , to prevent the denominator from being 0 and ensure numerical stability); θ t Represents the current model parameters after the t-th iteration; θ t+1 represents the updated model parameters after the t+1th iteration; m t Represents the first-order momentum of the gradient (the exponential moving average of the historical gradient, reflecting the gradient "direction / trend"); v t represents the second-order momentum of the gradient (the exponential moving average of the square of the historical gradient, reflecting the gradient "amplitude / fluctuation"); β1 represents the attenuation coefficient of the first-order momentum (usually 0.9 to control the influence of the historical gradient); m t-1 represents the first-order momentum (historical value, iterative transfer) at step t-1; g t Represents the original gradient calculated in step t (the loss function of θ t The result of the derivative); β2 represents the attenuation coefficient of the second-order momentum (usually 0.999 to control the influence of the square of the historical gradient); v t-1 Represents the second-order momentum at step t-1 (historical value, iterative transfer); represents the square of the original gradient at step t (capturing the change in gradient magnitude).
[0044] Furthermore, the use of the autocorrelation function to verify the correctness of the output period from the time domain or the frequency domain includes:
[0045] The time domain verification of the correctness of the output period includes calculating the lagged correlation through the autocorrelation function:
[0046]
[0047] Where τ is the time delay, ranging from 1≤τ≤T-1, and T is the length of the time series; X t The original time series is the mean;
[0048] The frequency domain verification of the correctness of the output period includes using fast Fourier transform to identify the period corresponding to the peak frequency:
[0049]
[0050] Where τ is the time delay, ranging from 1≤τ≤T-1, T is the length of the time series, and is the frequency index.
[0051] Beneficial effects: Compared with the prior art, the present invention has the following advantages:
[0052] By extracting local features at multiple time scales through CNN and combining it with LSTM to learn long-term dependencies across scales, we can break through the limitations of the traditional method of "single-scale modeling" and realize hierarchical feature mining of "local details-global trends", effectively improving the accuracy and reliability of period recognition of complex time series data and providing accurate period basis for practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 It is a schematic flow chart of the method of the present invention;
[0054] Figure 2 This is a power load diagram of a certain substation in an embodiment of the present invention;
[0055] Figure 3 is a histogram of the power load cycle distribution in an embodiment of the present invention. DETAILED DESCRIPTION
[0056] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is described in detail below with reference to the accompanying drawings and implementation examples. It should be understood that the specific implementation examples described herein are only used to explain the present invention and are not intended to limit the invention.
[0057] In one embodiment of the method of the present invention, the specific steps are as follows:
[0058] 1) For the original time series X∈R T×D Normalize and use sliding windows of different window sizes to generate multi-scale subsequences;
[0059] 2) The local features of each scale subsequence are extracted through a parallel convolutional network, and then spliced into a multi-scale comprehensive feature vector after global pooling compression;
[0060] 3) Arrange the multi-scale integrated feature vectors in the original time order and input them into the LSTM layer to learn the time sequence to capture the long-term dependencies in the time series;
[0061] 4) The global features output by LSTM are input into the fully connected layer, and the period recognition results are generated through the activation function;
[0062] 5) Use the autocorrelation function to verify the correctness of the model output period from the time domain / frequency domain.
[0063] In step 1), the method for generating multi-scale subsequences is as follows:
[0064] 11) The specific calculation method for multi-scale data preprocessing according to step 1) is as follows:
[0065]
[0066] Where μ is the mean of the sequence and σ' is the standard deviation of the sequence.
[0067] 12) Design K window sizes W = {w1, w2, ..., w K}, perform overlapping sliding sampling on the standardized sequence:
[0068]
[0069] Generate a multi-scale subsequence set where N k =Tw k +1.
[0070] In step 2), the CNN local feature extraction step includes:
[0071] 21) According to step 12), for each scale subsequence X (k) Use M convolution kernels Perform convolution:
[0072]
[0073] In the formula, ReLU is the activation function, b (m) For bias.
[0074] 22) Perform global average pooling on the convolution output according to step 21) and compress it into a feature vector F (k) , and splice to get the multi-scale comprehensive feature vector F CNN =[F (1) ,...,F (K) ].
[0075] In step 3), the LSTM long-term dependency modeling step includes:
[0076] 31) According to step 22) F CNN Reconstructed into a time-aligned feature sequence S∈R T′×(K×M) , where T′ is the time step, K is the number of multi-scales, and M is the single-scale feature dimension;
[0077] 32) According to step 31), long-term dependencies are captured by updating the LSTM cell state:
[0078]
[0079] where f t ,i t , o t is the gate vector, x t is the input feature at time t (feature vector extracted by CNN), with dimension D (feature dimension); h t is the hidden state at time t, with a dimension of H (LSTM hidden layer dimension), representing the temporal characteristics of the current moment; c t W is the cell state at time t, with a dimension of C (usually equal to H), storing long-term dependency information. f ,W i ,W c ,W o are the weight matrices of the forget gate, input gate, cell state update, and output gate, respectively, with a dimension of (H+D)×H. f ,b i ,b c ,b o The corresponding gate bias term, ⊙, is an element-wise multiplication, used to weight the gating signal's effect on the cell state. The dimension, H, adjusts the gating signal's threshold. The dimension, H, determines the LSTM's memory capacity. A larger H allows for more complex long-term dependencies to be learned, but the computational complexity increases quadratically. A smaller H may prevent the capture of long-term features.
[0080] In step 4), the cycle identification and model output steps include:
[0081] The classification task uses the Softmax function to output the cycle probability:
[0082]
[0083] Where W p Fully connected layer weight matrix; F LSTM is the time series feature vector output by LSTM; b p is the bias term of the fully connected layer.
[0084] The regression task outputs the periodic parameters through the linear layer:
[0085]
[0086] Update parameters through Adam optimizer:
[0087]
[0088] Where θ t is the model parameter at time t; η is the learning rate (global step coefficient), which globally controls the update amplitude of all parameters; m t ,v t are the first-order momentum and the second-order momentum respectively. is the smoothing term.
[0089] In step 5), the period validity verification step includes:
[0090] Time domain verification: Calculate the lagged correlation through the autocorrelation function:
[0091]
[0092] Where τ is the time delay (lag order), the value range is 1≤τ≤T-1, T is the length of the time series; X t is the original time series, is the mean.
[0093] Frequency domain verification: Use fast Fourier transform to identify the period corresponding to the peak frequency:
[0094]
[0095] Where τ is the time delay (lag order), ranging from 1≤τ≤T-1, T is the length of the time series, and k is the frequency index.
[0096] Based on the actual power supply area of a certain region, the data taken is the actual operating load data of the power grid between November 1, 2021 and October 31, 2022, with a sampling interval of 15 minutes. Figure 2 As shown; through the period probability distribution output by this application, combined with the autocorrelation function verification, the power load period distribution histogram is generated, as shown Figure 3 As shown in the figure, the histogram shows that the method of the present invention can clearly identify the coexistence relationship of multi-scale cycles, intuitively present the multi-scale periodic characteristics of the power load, realize the accurate identification and quantitative analysis of these complex cycles, and provide core technical support for the intelligent operation of the power system.
Claims
1. A multi-scale feature fusion cycle method based on CNN-LSTM, characterized in that: The method comprises the following steps: 1) For the original time series X, X∈R T×D Perform standardization to generate a standardized time series, and use at least two sliding windows of different window sizes to sample the standardized time series to generate multi-scale subsequences; 2) The local features of the multi-scale subsequences are extracted through a parallel convolutional network. The local features of the multi-scale subsequences are compressed by global pooling and then spliced into a multi-scale comprehensive feature vector; 3) Arrange the multi-scale integrated feature vectors in the original time order and input them into the LSTM layer to learn the time sequence to capture the long-term dependencies in the time series; 4) The global features output by LSTM are input into the fully connected layer, and the period recognition results are generated through the activation function; 5) Use the autocorrelation function to verify the accuracy of the output period from the time domain or frequency domain.
2. A multi-scale feature fusion cycle method based on CNN-LSTM according to claim 1, characterized in that, The original time series X, X∈R T×D Perform standardization to generate a standardized time series, and use at least two sliding windows of different window sizes to sample the standardized time series to generate multi-scale subsequences, including: 11) For the original time series X, X∈R T×D Perform standardization to generate a standardized time series X', including: Where, μ is the mean of the sequence; σ' is the standard deviation of the sequence; 12) At least two sliding windows of different window sizes are used to sample the standardized time series, including: Design K windows of size W, W={w1,w2,...,w K }, perform overlapping sliding sampling on the standardized time series X' generated in step 11): Where w1, w2, ..., w K Represent the 1st to kth windows respectively; Under the kth window width, the subsequence intercepted from position i, the superscript k represents the index of the window width, and different k corresponds to different window widths w k , the subscript i indicates the starting position of the interception; X' indicates the original sequence; X'[i:i+w k ] represents mathematical sequence interception, which means starting from the i-th position of X', the interception length is w k The subsequence of, including position i, to i+w k -1 position, left closed and right open logic; i is the starting position index of the window interception, the value range is i=1,2,...,Tw k +1;w k is the width of the kth window, that is, the length of each subsequence intercepted, and different k corresponds to different window widths; T represents the total length of the original sequence X'; k represents the index of the window width, and the value k=1,2,...,K, indicating that there are K windows of different widths for intercepting subsequences; K represents the total number of window widths, that is, there are a total of K different lengths w1,w2,...,w K The window is intercepted; Generate each scale subsequence X (k) , all scale subsequences X (k) Constructing a multi-scale subsequence set where N k =Tw k +1.
3. A multi-scale feature fusion cycle method based on CNN-LSTM according to claim 2, characterized in that, The method of extracting local features of multi-scale subsequences through a parallel convolutional network includes: 21) According to step 12), for each scale subsequence X (k) Use M convolution kernels Perform convolution: Where, Represents the final output. The subscript i usually represents the "sample / sequence position index", the k in the superscript (k,m) is the "category / level index of the input feature / sequence", and m is the "index of the current calculation channel / neuron / output dimension". ReLU() is the ReLU activation function, and the formula is ReLU(x) = max(0,x), which is responsible for introducing nonlinearity to allow the model to fit complex relationships. m -1 summation upper limit, l m Indicates "the input length / dimension of the mth transformation channel"; Indicates weighted summation of the input dimension / length, where t is the "dimension / position index" of the summation; Indicates input data; subscript i is the same as The subscript t is the index of the input in the "dimension / length direction", and the superscript (k) represents the "category / level identifier" of the input; W t (m) Represents the weight parameter, subscript t and input The dimension t corresponds to "the weight of the t-th dimension input"; The superscript (m) represents the weight set of the mth channel / neuron / output dimension; b (m) represents the bias parameter, and the superscript (m) corresponds to the bias of the mth channel / neuron; 22) The local features of the multi-scale subsequences are compressed by global pooling and then spliced into a multi-scale comprehensive feature vector, including: According to step 21), the convolution output is globally average pooled and compressed into a feature vector F (k) , and splice to get the multi-scale comprehensive feature vector F CNN =[F (1) ,...,F (K) ].
4. A multi-scale feature fusion cycle method based on CNN-LSTM according to claim 3, characterized in that: The multi-scale integrated feature vectors are arranged in the original time order and input into the LSTM layer to learn the time sequence to capture the long-term dependencies in the time series, including: 31) According to step 22) F CNN Reconstructed into a time-aligned feature sequence, S∈R T′×(K×M) , where T′ is the time step, K is the number of multi-scales, and M is the single-scale feature dimension; 32) According to step 31), long-term dependencies are captured by updating the LSTM cell state: Where S t Indicates the current input; h t-1 Indicates the hidden state at the previous moment; f t Represents the output of the forget gate; i t Represents the input gate output; o t Represents the output of the output gate; represents the candidate cell state; c t Indicates the current cell state; W f 、W i 、W c 、W o Respectively represent the linear transformation parameters of the weight matrix corresponding to the forget gate, input gate, candidate cell, and output gate; b f 、b i 、b c 、b o Respectively represent bias terms; ReLU() is the ReLU activation function; tanh() represents the tanh activation function; ⊙ represents element-by-element multiplication.
5. A multi-scale feature fusion cycle method based on CNN-LSTM according to claim 4, characterized in that: The global features output by the LSTM are input into the fully connected layer, and the period recognition results are generated through the activation function, including: The classification task uses the Softmax function to output the cycle probability: Where W p Fully connected layer weight matrix; F LSTM is the time series feature vector output by LSTM; b p is the bias term of the fully connected layer; The regression task outputs the periodic parameters through the linear layer Update parameters through Adam optimizer: Where θ represents the model parameters; t represents the iteration step; η represents the learning rate; represents the minimum value; θ t Represents the current model parameters after the t-th iteration; θ t+1 represents the updated model parameters after the t+1th iteration; m t represents the first-order momentum of the gradient; v t represents the second-order momentum of the gradient; β1 represents the attenuation coefficient of the first-order momentum; m t-1 represents the first-order momentum of step t-1; g t represents the original gradient calculated at step t; β2 represents the attenuation coefficient of the second-order momentum; v t-1 represents the second-order momentum at step t-1; Represents the square of the original gradient at step t.
6. The multi-scale feature fusion cycle method based on CNN-LSTM according to claim 1, characterized in that: The method of verifying the accuracy of the output period by using the autocorrelation function from the time domain or the frequency domain includes: The time domain verification of the correctness of the output period includes calculating the lagged correlation through the autocorrelation function: Where τ is the time delay, ranging from 1≤τ≤T-1, and T is the length of the time series; X t is the original time series, is the mean; The frequency domain verifies the correctness of the output period by using fast Fourier transform to identify the period corresponding to the peak frequency: Where τ is the time delay, ranging from 1≤τ≤T-1, T is the length of the time series, and is the frequency index.