Multivariate time series prediction method and system based on trend signal decoupling and prediction optimization

By using technical means such as short-time Fourier transform and Gaussian hybrid model in multivariate long-term time series prediction, the problem of trend signal capture and prediction is solved, and the learning efficiency and prediction accuracy of the model are improved.

CN119939329APending Publication Date: 2025-05-06SHANGHAI JIAOTONG UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411780658.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-05
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture and predict trend signals in multivariate long-term time series prediction, especially in the presence of information sparseness and non-normal distributed variables, resulting in low model learning efficiency and poor generalization ability.

Method used

By separating the trend signal from the original time series using a short-time Fourier transform and converting it into a cumulative probability density sequence, lightweight fully connected network layer prediction is performed in combination with the Gaussian hybrid model and the inverted attention mechanism, and finally recovering the time domain prediction value using the inverse short-time Fourier transform.

Benefits of technology

It reduces the sparseness of time series information, retains the more fine-grained trend characteristics, improves the learning efficiency of the model and adapts to non-normal inputs, and enhances the accuracy of prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939329A_ABST
    Figure CN119939329A_ABST
Patent Text Reader

Abstract

The invention provides a multivariate time sequence prediction method and system based on trend signal decoupling and prediction optimization. The multivariate time sequence prediction method comprises the following steps: step 1, separating a trend numerical sequence from an original time sequence signal by using short-time Fourier transform; 2, converting the trend numerical value sequence into a cumulative probability density sequence; step 3, sharing weights among different frequencies, variables and real parts or imaginary parts by adopting a light-weight complex-level full-connection network layer, and removing high-frequency signal noise by using a low-pass filter; and 4, performing model training and reasoning, combining trend prediction and multi-cycle prediction, and recovering a time domain prediction value by using reverse short-time Fourier transform. According to the method, the trend signal is separated from the original time sequence by using short-time Fourier transform, so that the information sparsity of the original signal is reduced, the trend characteristic with finer granularity is reserved, and the learning efficiency of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multivariate long-term time series prediction, and in particular to a multivariate time series prediction method and system based on trend signal decoupling and prediction optimization. Background Art

[0002] Multivariate long-term time series forecasting tasks are widely used in various fields such as finance, weather, medical care, transportation, and energy. For example, this task can be used to predict changes in demand for energy such as electricity, which helps to balance supply and demand and improve energy efficiency. Time series signals are usually manifested as a mixture of short-term multi-period fluctuations and long-term trends. In practice, trend signals have a more important impact on time series forecasting results because they determine the coarse-grained changes in the overall signal. Therefore, capturing the changing pattern of learning trends is more important and challenging than short-term multi-period fluctuations. However, there are two difficulties in learning trend signals: on the one hand, the information change density in the original sequence data is low, which affects the spatiotemporal efficiency and generalization ability of model learning, resulting in the model being unable to handle longer sequence inputs and dependencies. On the other hand, time series signals of different variables on the cross section often have diverse and non-normal overall distributions in a data set, such as long or short tails, symmetric or skewed, unimodal or bimodal. Since modern deep neural networks can be approximated as the sum of different random variables, they are often best at inputs and outputs with normal distributions. The specificity of the distribution of different cross-sectional variables makes it more difficult for modern deep neural networks to learn cross-sectional change relationships, thus affecting the accuracy of trend signal predictions.

[0003] This work focuses on reducing the sparsity of time series information and the singular distribution of modeling time series variables. The defects of existing technologies are as follows:

[0004] (1) Most existing trend-cycle decomposition technologies start from a global static perspective and perform a one-time decomposition of the input sequence as a whole. The trend obtained for each input sequence is a scalar, which loses the trend signal of finer granularity and increases the difficulty of model learning. The present invention aims to reduce the sparsity of the original sequence information through trend-cycle decomposition while retaining the trend signal of finer granularity.

[0005] (2) Slicing technology can alleviate information sparsity while retaining finer-grained sequence changes. However, slicing technology is mostly performed in the time domain, and each slice retains short-period fluctuations within the window. Since the short-period fluctuation pattern is relatively simple and stable, retaining this information will waste the overall model resources. The present invention aims to extract and process multi-period components separately and in a lightweight and efficient manner to reduce the model burden.

[0006] (3) Existing normalization-related technologies are unable to accurately characterize complex non-normal distributions. On the one hand, simple slice / cluster division and small sample size in a single input window make it difficult to ensure the statistical approximate normality of the local structure. On the other hand, due to the diverse distribution of different variables, treating all variables in the cross section as a whole lacks flexibility. The present invention aims to use an unsupervised algorithm with stronger characterization capabilities to fit the distribution of the input window on all training data, ensuring that the results have better approximate normality in the local structure, while adopting a channel-independent strategy to ensure that the model can flexibly handle the distribution of different variables.

[0007] (4) Although the normalized flow technology can more accurately characterize various complex distributions through mathematical changes from simple normal distributions, the coordinated optimization of the generation module and the discrimination module also makes the model training optimization more complicated, and puts forward higher requirements for solving potential problems such as mode collapse and mode collapse. The present invention aims to simplify the model design and training optimization while retaining the ability to model complex distributions. Summary of the invention

[0008] In view of the defects in the prior art, the object of the present invention is to provide a multivariate time series prediction method and system based on trend signal decoupling and prediction optimization.

[0009] The multivariate time series prediction method based on trend signal decoupling and prediction optimization provided by the present invention includes:

[0010] Step 1: Use short-time Fourier transform to separate the trend value series from the original time series signal;

[0011] Step 2: Convert the trend value sequence into a cumulative probability density sequence. The value of each time step in the cumulative probability density sequence is the cumulative probability density obtained by transforming the radial basis function based on the Gaussian mixture model components.

[0012] Step 3: Use a lightweight fully connected network layer to share weights between different frequencies, variables, and real or imaginary parts, and use a low-pass filter to remove high-frequency signal noise;

[0013] Step 4: Perform model training and inference, merge the trend forecast and multi-period forecast, and use the inverse short-time Fourier transform to restore the time domain forecast value.

[0014] Preferably, the step 1 comprises:

[0015] Given the original time series input For each variable sequence Use short-time Fourier transform to extract the trend, that is, the frequency component with zero phase; the sliding window length is recorded as N, and the step size is recorded as h, d t,krepresents the fast Fourier transform result of the kth frequency in the window [t, t+N-1], d t,0 is the zero phase trend, and the rest d t,k is a multi-periodic signal. Due to the conjugate symmetry of the fast Fourier transform, for k>0, d t,k =d t,N-k+1 ,get multi-periodic components; bidirectional reflection filling is used to mitigate boundary effects in the short-time Fourier transform, and the weight w (n) By default, it is set to 1, since the number of windows in the short-time Fourier transform is Where L p is the input length after padding, and the shape of the trend signal is The shape of the multi-periodic signal is Then we have:

[0016]

[0017] in, It represents the nth timing point in the timing window [t, t+N-1] with a length of N, starting from the subscript t. j represents the imaginary unit. n represents the number of the timing point in the window, and its value range is {0, 1, ..., N-1}.

[0018] Preferably, step 2 comprises:

[0019] First, all input window data are instance normalized and then used to construct the histogram, then:

[0020]

[0021] in, is the trend sequence of variable d input, μ d and are the mean and variance of the trend sequence respectively, and s represents the time point number in the input sequence; Indicates that it belongs to variable d and has a length of L i The sth time series data point in the input sequence; t′ d represents the normalized input sequence of variable d; σ d represents the standard deviation of the input sequence of variable d;

[0022] The histogram of each variable is composed of all Construct and use the Gaussian mixture model to approximate the histogram. The number of Gaussian components K of all variables is the same, and the result is the mean M∈R of the K Gaussian components D×K and variance V∈R D×K; Construct a Gaussian radial basis function according to the mean and variance of each component. Φ is the cumulative density function of the standard normal distribution, which converts the numerical sequence of the trend into a cumulative probability density sequence Use Page approximation Substituting the standard Φ, we have:

[0023] f d,k,: =Φ((t' d -m d,k ) / v d,k ),

[0024] α1=0.070565992, α2=1.5976

[0025] Among them, f d,k,: represents the result of processing the input sequence of variable d using the cumulative density function of the kth Gaussian component in GMM; m d,k represents the mean of the kth Gaussian component in the GMM fitting result of the global distribution of variable d; v d,k represents the standard deviation of the kth Gaussian component in the GMM fitting result of the global distribution of variable d; Φ(s) represents the cumulative density function of the standard Gaussian distribution; The approximate calculation formula for the cumulative density function of the standard Gaussian distribution; α1 and α2 represent Calculate the required constants;

[0026] The embedding vector dimension is recorded as e, the hidden vector dimension is recorded as h, and the inverted attention mechanism is used as the skeleton model. First, the cumulative probability density sequence F is transformed into Embedding coding is performed, and then the learnable time series vector pe∈R is added 1×1×e , after layer normalization Layer norm and random dropout, we get H0∈R D×K×e As the first layer input, the two dimensions D and K of the variable and Gaussian components are regarded as channels, so the two are merged, and we have:

[0027] H0=Dropout(Layernorm(pe+F*E T ))

[0028] For the i+1th layer, given H i Afterwards, the multi-head inverted attention mechanism is combined with the two-layer fully connected network MLP to model the relationship between different channels. With H i Keeping the same dimensions, we have:

[0029]

[0030] Among them, H i is the input of the i-th layer network, Yes H i The intermediate result after being processed by the inverted attention module of the i-th layer network, yes The output after being processed by the MLP module of the i-th layer network;

[0031] A hierarchical structure is also constructed. If there are multiple encoder layers, between each two adjacent layers, the adjacent vectors output by the previous layer are merged by linear projection, and then layer normalization is performed, then:

[0032]

[0033] Among them, W merge,i-1 represents the projection matrix required to merge adjacent vectors in the output vector of the i-1th layer; b merge,i-1 Represents the projection bias required to merge adjacent vectors in the output vector of the i-1th layer;

[0034] Finally, a logistic regression layer is used as a decoder to predict the output cumulative probability density sequence O T , then:

[0035] O T =σ(H l *W d +b d ),

[0036] Where σ represents the sigmoid function, which maps the output value to the interval [0,1], σ(x) = 1 / (1+exp(-x)); H l Represents the final output representation after the lth layer of the model; l represents the number of layers of the model; W d represents the projection matrix of the linear transformation of the decoder, d is the first letter of the decoder word; b d Represents the bias vector of the linear transformation of the decoder decoder, where d is the first letter of the decoder word.

[0037] Preferably, the step 3 comprises:

[0038] Given a multi-periodic signal S, first extract its real and imaginary parts into a matrix In the example, the low-frequency signal with the lowest r ratio is retained; then the matrix is ​​used to Embedding and batch normalization are performed, followed by a two-layer fully connected neural network processing, including batch normalization, plus residual connection and linear projection to obtain the output prediction of the period module The expression is:

[0039]

[0040] O P =(S1+Dropout(MLP(S1)))*W o +b o

[0041] Where S1 represents the multi-periodic input signal after low-pass filtering, embedded coding and batch normalization; W o Represents the projection matrix of the linear transformation of the multi-cycle prediction output layer, o represents output; b o Represents the bias vector of the linear transformation of the multi-period prediction output layer, and o represents output.

[0042] Preferably, step 4 comprises:

[0043] Multi-task learning is adopted, and the sum of binary cross entropy BCE and mean square error MSE is used as the loss function. and are multi-period and trend training labels respectively, then we have:

[0044] L=L BCE (O T ,Y T )+L MSE (O P ,Y P )

[0045] Where L represents the loss function; L BCE Represents the binary cross entropy loss function L BCE (o T ,y T )=y T logo T +(1-y T )log(1-o T ), o T and T are the predicted cdf trend value and the real cdf trend value respectively; L MSE Represents the mean square error loss function o P and P are the predicted period component value and the actual period component value respectively;

[0046] During the inference process, there are K cumulative probability density outputs from different Gaussian components for each variable and time step. The result with the maximum confidence level is selected from these K outputs, and then restored to the numerical sequence and inverse normalized to obtain the output of trend prediction. Then we have:

[0047]

[0048] Among them, k * It means that each variable d has K different Gaussian components corresponding to the cdf prediction sequence, and k* represents the Gaussian component number corresponding to the best sequence selected; It represents the cdf trend prediction output corresponding to the Gaussian component k at the future time t on the variable d; Indicates that the selected cdf prediction at time t in the future on variable d is converted into a numerical prediction and then output after inverse normalization; The inverse function of the approximate function Φ~(s) representing the cumulative density function of the standard Gaussian distribution; s′ represents the calculation The intermediate result of is calculated by input s;

[0049] Next, after combining the trend forecast and the multi-period forecast, the inverse short-time Fourier transform (iSTFT) is used to restore the time domain forecast value. The filtered high-frequency signal is replaced by 0 value in the output part, then:

[0050]

[0051] The multivariate time series prediction system based on trend signal decoupling and prediction optimization provided by the present invention includes:

[0052] Module M1: Use short-time Fourier transform to separate the trend value series from the original time series signal;

[0053] Module M2: Convert the trend value sequence into a cumulative probability density sequence. The value of each time step in the cumulative probability density sequence is the cumulative probability density obtained by transforming the radial basis function based on the Gaussian mixture model components.

[0054] Module M3: uses a lightweight fully connected network layer to share weights between different frequencies, variables, and real or imaginary parts, and uses a low-pass filter to remove high-frequency signal noise;

[0055] Module M4: Perform model training and inference, merge trend forecast and multi-period forecast, and use inverse short-time Fourier transform to restore the time domain forecast value.

[0056] Preferably, the module M1 comprises:

[0057] Given the original time series input For each variable sequence Use short-time Fourier transform to extract the trend, that is, the frequency component with zero phase; the sliding window length is recorded as N, and the step size is recorded as h, dt,k represents the fast Fourier transform result of the kth frequency in the window [t, t+N-1], d t,0 is the zero phase trend, and the rest d t,k is a multi-periodic signal. Due to the conjugate symmetry of the fast Fourier transform, for k>0, d t,k =d t,N-k+1 ,get multi-periodic components; bidirectional reflection filling is used to mitigate boundary effects in the short-time Fourier transform, and the weight w (n) By default, it is set to 1, since the number of windows in the short-time Fourier transform is Where L p is the input length after padding, and the shape of the trend signal is The shape of the multi-periodic signal is Then we have:

[0058]

[0059] in, It represents the nth timing point in the timing window [t, t+N-1] with a length of N, starting from the subscript t. j represents the imaginary unit. n represents the number of the timing point in the window, and its value range is {0, 1, ..., N-1}.

[0060] Preferably, the module M2 comprises:

[0061] First, all input window data are instance normalized and then used to construct the histogram, then:

[0062]

[0063] in, is the trend sequence of variable d input, μ d and are the mean and variance of the trend sequence respectively, and s represents the time point number in the input sequence; Indicates that it belongs to variable d and has a length of L i The sth time series data point in the input sequence; t′ d represents the normalized input sequence of variable d; σ d represents the standard deviation of the input sequence of variable d;

[0064] The histogram of each variable is composed of all Construct and use the Gaussian mixture model to approximate the histogram. The number of Gaussian components K of all variables is the same, and the result is the mean M∈R of the K Gaussian components D×K and variance V∈R D×K; Construct a Gaussian radial basis function according to the mean and variance of each component. Φ is the cumulative density function of the standard normal distribution, which converts the numerical sequence of the trend into a cumulative probability density sequence Replace the standard Φ with Page's approximate Φ~, and we have:

[0065] f d,k,: =Φ((t' d -m d,k ) / v d,k ),

[0066] α1=0.070565992, α2=1.5976

[0067] Among them, f d,k,: represents the result of processing the input sequence of variable d using the cumulative density function of the kth Gaussian component in GMM; m d,k represents the mean of the kth Gaussian component in the GMM fitting result of the global distribution of variable d; v d,k represents the standard deviation of the kth Gaussian component in the GMM fitting result of the global distribution of variable d; Φ(s) represents the cumulative density function of the standard Gaussian distribution; The approximate calculation formula for the cumulative density function of the standard Gaussian distribution; α1 and α2 represent Calculate the required constants;

[0068] The embedding vector dimension is recorded as e, the hidden vector dimension is recorded as h, and the inverted attention mechanism is used as the skeleton model. First, the cumulative probability density sequence F is transformed into Embedding coding is performed, and then the learnable time series vector pe∈R is added 1×1×e , after layer normalization Layer norm and random dropout, we get H0∈R D×K×e As the first layer input, the two dimensions D and K of the variable and Gaussian components are regarded as channels, so the two are merged, and we have:

[0069] H0=Dropout(Layernorm(pe+F*E T ))

[0070] For the i+1th layer, given H i Afterwards, the multi-head inverted attention mechanism is combined with the two-layer fully connected network MLP to model the relationship between different channels. With H i Keeping the same dimensions, we have:

[0071]

[0072] Among them, H i is the input of the i-th layer network, Yes H i The intermediate result after being processed by the inverted attention module of the i-th layer network, yes The output after being processed by the MLP module of the i-th layer network;

[0073] A hierarchical structure is also constructed. If there are multiple encoder layers, between each two adjacent layers, the adjacent vectors output by the previous layer are merged by linear projection, and then layer normalization is performed, then:

[0074]

[0075] Among them, W merge,i-1 represents the projection matrix required to merge adjacent vectors in the output vector of the i-1th layer; b merge,i-1 Represents the projection bias required to merge adjacent vectors in the output vector of the i-1th layer;

[0076] Finally, a logistic regression layer is used as a decoder to predict the output cumulative probability density sequence O T , then:

[0077]

[0078] Where σ represents the sigmoid function, which maps the output value to the interval [0,1], σ(x) = 1 / (1+exp(-x)); H l Represents the final output representation after the lth layer of the model; l represents the number of layers of the model; W d represents the projection matrix of the linear transformation of the decoder, d is the first letter of the decoder word; b d Represents the bias vector of the linear transformation of the decoder decoder, where d is the first letter of the decoder word.

[0079] Preferably, the module M3 comprises:

[0080] Given a multi-periodic signal S, first extract its real and imaginary parts into a matrix In the example, the low-frequency signal with the lowest r ratio is retained; then the matrix is ​​used to Embedding and batch normalization are performed, followed by a two-layer fully connected neural network processing, including batch normalization, plus residual connection and linear projection, to obtain the output prediction of the periodic module The expression is:

[0081]

[0082] OP =(S1+Dropout(MLP(S1)))*W o +b o

[0083] Where S1 represents the multi-periodic input signal after low-pass filtering, embedded coding and batch normalization; W o Represents the projection matrix of the linear transformation of the multi-cycle prediction output layer, o represents output; b o Represents the bias vector of the linear transformation of the multi-period prediction output layer, and o represents output.

[0084] Preferably, the module M4 comprises:

[0085] Multi-task learning is adopted, and the sum of binary cross entropy BCE and mean square error MSE is used as the loss function. and are multi-period and trend training labels respectively, then we have:

[0086] L=L BCE (O T ,Y T )+L MSE (O P ,Y P )

[0087] Where L represents the loss function; L BCE Represents the binary cross entropy loss function L BCE (o T ,y T )=y T logo T +(1-y T )log(1-o T ), o T and T are the predicted cdf trend value and the real cdf trend value respectively; L MSE Represents the mean square error loss function o P and P are the predicted period component value and the actual period component value respectively;

[0088] During the inference process, there are K cumulative probability density outputs from different Gaussian components for each variable and time step. The result with the maximum confidence level is selected from these K outputs, and then restored to the numerical sequence and inverse normalized to obtain the output of trend prediction. Then we have:

[0089]

[0090] Among them, k* It means that each variable d has K different Gaussian components corresponding to the cdf prediction sequence, and k* represents the Gaussian component number corresponding to the best sequence selected; It represents the cdf trend prediction output corresponding to the Gaussian component k at the future time t on the variable d; Indicates that the selected cdf prediction at time t in the future on variable d is converted into a numerical prediction and then output after inverse normalization; Approximate function representing the cumulative density function of the standard Gaussian distribution The inverse function of The intermediate result of is calculated from the input s;

[0091] Next, after combining the trend forecast and the multi-period forecast, the inverse short-time Fourier transform (iSTFT) is used to restore the time domain forecast value. The filtered high-frequency signal is replaced by 0 value in the output part, then:

[0092]

[0093] Compared with the prior art, the present invention has the following beneficial effects:

[0094] (1) The present invention uses short-time Fourier transform to separate trend signals from the original time series, thereby reducing the information sparsity of the original signal, while retaining more fine-grained trend features and improving the learning efficiency of the model;

[0095] (2) Considering that different variables have diverse and strange non-normal distributions, the present invention uses a Gaussian mixture model to fit the non-normal distribution and constructs basis functions according to its different components, and converts the numerical trend sequence values ​​into multiple cumulative probability density sequences. Since different components in the original distribution have different weights, and the generated sequences corresponding to different components here have the same preference, the model's attention to low-weight components is improved, thereby enhancing the ability of the method to adapt to non-normal inputs. BRIEF DESCRIPTION OF THE DRAWINGS

[0096] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:

[0097] Figure 1 To decouple trend signals from multi-period signals using short-time Fourier transform;

[0098] Figure 2 To transform the trend series into cumulative probability density series using Gaussian mixture model;

[0099] Figure 3 is the inverted attention module for trend prediction;

[0100] Figure 4 A lightweight shared parameter layer for multi-period prediction; Figure 5 Flowchart of the multivariate time series forecasting method based on trend signal decoupling and forecast optimization. DETAILED DESCRIPTION

[0101] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several changes and improvements can also be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0102] Example

[0103] The present invention provides a multivariate time series prediction method based on trend signal decoupling and prediction optimization, comprising:

[0104] In order to reduce the information sparsity of time series signals, the present invention uses short-time Fourier transform to separate trends from the original time series. Different from other studies that use global transform, the short-time Fourier transform directly performs fast Fourier transform in each short-time window to output complex frequency.

[0105] For trend sequences, the present invention first normalizes all input windows of each variable in the training set and aggregates them into a histogram as an approximation of the overall distribution of the variable, then uses a Gaussian mixture model for fitting, and uses the mean and variance of each Gaussian component to construct a Gaussian radial basis function, converts the trend numerical sequence into a cumulative probability density sequence, and finally uses the reverse attention module for learning.

[0106] For the remaining multi-period sequences, based on their sparsity and low rank, the present invention adopts a lightweight shared weight fully connected neural network layer and a low-pass filter to efficiently predict the module to reduce the model burden.

[0107] Specifically, it includes:

[0108] In the multivariate long-term time series forecasting task, given the input data Where D is the number of variables, L i is the time step, and the goal is to predict the multivariate sequence Where L o is the predicted length. Figure 1 In order to decouple trend signals from multi-period signals using short-time Fourier transform, Figure 2 In order to transform the trend series into cumulative probability density series using Gaussian mixture model, Figure 3 is the inverted attention module for trend prediction, Figure 4Lightweight shared parameter layer for multi-period prediction.

[0109] like Figure 5 First, the first part introduces how to use short-time Fourier transform to separate trends from raw time series signals; then, the second part introduces how to improve trend prediction based on Gaussian mixture model; the third part explains how to lightweight predict multi-period components; finally, the fourth part introduces how to integrate these modules during training and inference.

[0110] 1. Decompose the original signal to extract the trend

[0111] Given the original time series input For each variable sequence The present invention uses short-time Fourier transform to extract the trend, that is, the frequency component with zero phase. The sliding window length is recorded as N and the step length is recorded as h. t,k Represents the fast Fourier transform result of the kth frequency in the window [t, t+N-1]. t,0 is the zero phase trend, and the rest is a multi-periodic signal. Due to the conjugate symmetry of the fast Fourier transform, for k>0, d t,k =d t,N-k+1 , we can get In practice, the present invention uses bidirectional reflection filling to mitigate the boundary effect in the short-time Fourier transform. (n) The default setting is 1. Since the number of windows in the short-time Fourier transform is Where L p is the input length after padding, and the shape of the trend signal is The shape of the multi-periodic signal is

[0112]

[0113] in, It represents the nth timing point in the timing window [t, t+N-1] with a length of N, starting from the subscript t. j represents the imaginary unit. n represents the number of the timing point in the window, and its value range is {0, 1, ..., N-1}.

[0114] 2. Prediction of trend signals

[0115] In this module, the present invention converts the trend value sequence into a cumulative probability density sequence, and each time step value in the cumulative probability density sequence is a cumulative probability density obtained by transforming the radial basis function constructed based on the Gaussian mixture model components.

[0116] (1) Instance normalization of input windows. Some variables have long-term distribution shifts on different input windows, and their distribution histograms constructed by training sets may be affected. Therefore, before quantitatively analyzing the overall distribution of each variable, the present invention first performs instance normalization on all input window data, and then uses them to construct histograms. is the trend sequence of variable d input, μ d and are the mean and variance of the trend series respectively.

[0117] t' d =(t d -μ d ) / σ d

[0118] Where s represents the time point number in the input sequence; Indicates that it belongs to variable d and has a length of L i The sth time series data point in the input sequence; t′ d represents the normalized input sequence of variable d; σ d Denotes the standard deviation of the input sequence of variable d.

[0119] (2) Apply the Gaussian mixture model. The histogram of each variable is composed of all The present invention uses a Gaussian mixture model to approximate the histogram. The number of Gaussian components K for all variables is the same. Ideally, for each variable, different Gaussian components will overlap each other but cover different sub-ranges in the histogram. The result is the mean M∈R of the K Gaussian components D×K and variance V∈R D×K Next, the present invention constructs a Gaussian radial basis function (Φ is the cumulative density function of the standard normal distribution) according to the mean and variance of each component, and converts the numerical sequence of the trend into a cumulative probability density sequence In practice, in order to speed up the calculation, the present invention uses Page approximation Φ~ to replace the standard Φ, and its absolute approximation error does not exceed 1.4×10 -4 .

[0120] f d,k,: =Φ((t' d -m d,k ) / v d,k ),

[0121] α1=0.070565992, α2=1.5976

[0122] Among them, f d,k,:represents the result of processing the input sequence of variable d using the cumulative density function of the kth Gaussian component in GMM; m d,k represents the mean of the kth Gaussian component in the GMM fitting result of the global distribution of variable d; v d,k represents the standard deviation of the kth Gaussian component in the GMM fitting result of the global distribution of variable d; Φ(s) represents the cumulative density function of the standard Gaussian distribution; The approximate calculation formula for the cumulative density function of the standard Gaussian distribution; α1 and α2 represent Calculates the required constants.

[0123] (3) Inverted Transformer module with hierarchical merging. The embedding vector dimension is denoted as e, and the hidden vector dimension is denoted as h. The present invention adopts the inverted attention mechanism as the skeleton model. First, the present invention converts the cumulative probability density sequence F into a matrix Embedding coding is performed, and then the learnable time series vector pe∈R is added 1×1×e , after layer normalization (Layernorm) and random dropout (Dropout), we get H0∈R D×K×e As the first layer input, both dimensions (D and K) of variables and Gaussian components are considered as channels, so they are merged.

[0124] H0=Dropout(Layernorm(pe+F*E T ))

[0125] For the i+1th layer, given H i Afterwards, the present invention combines the multi-head inverted attention mechanism (Inverted Attention) and the two-layer fully connected network (MLP) to model the relationship between different channels (variables and Gaussian components). With H i Keep the same dimensions.

[0126]

[0127] Among them, H i is the input of the i-th layer network, Yes H i The intermediate result after being processed by the inverted attention module of the i-th layer network, yes The output after processing by the MLP module of the i-th layer network.

[0128] The present invention also constructs a hierarchical structure. If there are multiple encoder layers, between each two adjacent layers, the method merges adjacent vectors output by the previous layer through linear projection, and then performs layer normalization.

[0129]

[0130] Among them, W merge,i-1 represents the projection matrix required to merge adjacent vectors in the output vector of the i-1th layer; b merge,i-1 Represents the projection bias required to merge adjacent vectors in the i-1th layer output vector.

[0131] Finally, the present invention uses a simple logistic regression layer as a decoder to predict the cumulative probability density sequence of the output T .

[0132] O T =σ(H l *W d +b d ),

[0133] Where σ represents the sigmoid function, which maps the output value to the interval [0,1], σ(x) = 1 / (1+exp(-x)); H l Represents the final output representation after the lth layer of the model; l represents the number of layers of the model; W d represents the projection matrix of the linear transformation of the decoder, where d is the first letter of the word decoder; b d Represents the bias vector of the decoder linear transformation, where d is the first letter of the word decoder.

[0134] 3. Prediction of multi-period signals

[0135] Since multi-periodic signals are sparse and low-rank, the present invention uses a lightweight fully connected network layer to share weights between different frequencies, variables, and real or imaginary parts. In addition, the present invention uses a low-pass filter to remove high-frequency signal noise. Given a multi-periodic signal S, first extract its real and imaginary parts into a matrix The low-frequency signal with the lowest r ratio is retained. Then the matrix Embedding and batch normalization are performed, followed by a two-layer fully connected neural network processing (including batch normalization), plus residual connection and linear projection to obtain the output prediction of the period module

[0136]

[0137] O P =(S1+Dropout(MLP(S1)))*W o +b o

[0138] Where S1 represents the multi-periodic input signal after low-pass filtering, embedded coding and batch normalization; W o Represents the projection matrix of the linear transformation of the multi-cycle prediction output layer, o represents output; b o Represents the bias vector of the linear transformation of the multi-period prediction output layer, and o represents output.

[0139] IV. Model training and reasoning

[0140] Trend prediction during training is a probabilistic regression problem, while multi-period real and imaginary part prediction is a numerical regression problem. Therefore, this method adopts multi-task learning and uses the sum of binary cross entropy (BCE) and mean square error (MSE) as the loss function. and They are multi-period and trend training labels respectively.

[0141] L=L BCE (O T ,Y T )+L MSE (O P ,Y P )

[0142] Where L represents the loss function; L BCE Represents the binary cross entropy loss function L BCE (o T ,y T )=y T logo T +(1-y T )log(1-o T ), o T and T are the predicted cdf trend value and the real cdf trend value respectively; L MSE Represents the mean square error loss function o P and P are the predicted periodic component values ​​and the actual periodic component values, respectively.

[0143] During the inference process, there are K cumulative probability density outputs from different Gaussian components for each variable and time step. The present invention selects the result with the maximum confidence level from these K outputs, that is, the cumulative probability density value closer to 0.5 than to 0 or 1, and then restores it to a numerical sequence (using Calculate) and perform inverse normalization to obtain the output of trend prediction in It can be calculated using the Cartan formula (s1 and s2 are intermediate variables).

[0144]

[0145] Among them, k * It means that each variable d has K different Gaussian components corresponding to the cdf prediction sequence, and k* represents the Gaussian component number corresponding to the best sequence selected; It represents the cdf trend prediction output corresponding to the Gaussian component k at the future time t on the variable d; Indicates that the selected cdf prediction at time t in the future on variable d is converted into a numerical prediction and then output after inverse normalization; Approximate function representing the cumulative density function of the standard Gaussian distribution The inverse function of The intermediate result is calculated from the input s.

[0146] Next, after combining the trend prediction and the multi-period prediction, the present invention uses the inverse short-time Fourier transform (iSTFT) to restore the time domain prediction value By default, the high frequency signals that are filtered out are replaced by 0 values ​​in the output.

[0147]

[0148] In addition, the present invention also introduces two improvement measures:

[0149] (1) Trend label smoothing. In order to reduce the boundary effect between the input window and the output window. During training, this method connects the two and performs short-time Fourier transform on the entire sequence, and intercepts all output-related signals from the second half as training labels. The model output length is the same as the training label. During inference, some redundant outputs near the boundary will be discarded. The input trend signal always comes from the input window, avoiding future data leakage.

[0150] (2) Predicting longer future windows. When the prediction length is short, the future information contained in the training labels may be insufficient, which may hinder the model from learning long-term patterns. A common approach is to use longer output windows as training labels and truncate the predictions during inference. To ensure fairness of the results, this method ensures that the test set sample size remains unchanged under different future window lengths.

[0151] Those skilled in the art know that, in addition to implementing the system, device and its various modules provided by the present invention in a purely computer-readable program code, it is entirely possible to implement the same program in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers and embedded microcontrollers by logically programming the method steps. Therefore, the system, device and its various modules provided by the present invention can be considered as a hardware component, and the modules included therein for implementing various programs can also be considered as structures within the hardware component; the modules for implementing various functions can also be considered as both software programs for implementing the method and structures within the hardware component.

[0152] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. In the absence of conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.

Claims

1. A multivariate time series prediction method based on trend signal decoupling and prediction optimization, characterized in that: include: Step 1: Use short-time Fourier transform to separate the trend value series from the original time series signal; Step 2: Convert the trend value sequence into a cumulative probability density sequence. The value of each time step in the cumulative probability density sequence is the cumulative probability density obtained by transforming the radial basis function based on the Gaussian mixture model components. Step 3: Use a lightweight fully connected network layer to share weights between different frequencies, variables, and real or imaginary parts, and use a low-pass filter to remove high-frequency signal noise; Step 4: Perform model training and inference, merge trend forecast and multi-period forecast, and use inverse short-time Fourier transform to restore the time domain forecast value.

2. The multivariate time series prediction method based on trend signal decoupling and prediction optimization according to claim 1 is characterized in that: The step 1 comprises: Given the original time series input For each variable sequence Use short-time Fourier transform to extract the trend, that is, the frequency component with zero phase; the sliding window length is recorded as N, and the step size is recorded as h, d t,k represents the fast Fourier transform result of the kth frequency in the window [t, t+N-1], d t,0 is the zero phase trend, and the rest d t,k is a multi-periodic signal. Due to the conjugate symmetry of the fast Fourier transform, for k>0, d t,k =d t,N-k+1 ,get multi-periodic components; bidirectional reflection filling is used to mitigate boundary effects in the short-time Fourier transform, and the weight w (n) By default, it is set to 1, since the number of windows in the short-time Fourier transform is Where L p is the input length after padding, and the shape of the trend signal is The shape of the multi-periodic signal is Then we have: in, It represents the nth timing point in the timing window [t, t+N-1] with a length of N, starting from the subscript t. j represents the imaginary unit. n represents the number of the timing point in the window, and its value range is {0, 1, ..., N-1}.

3. The multivariate time series prediction method based on trend signal decoupling and prediction optimization according to claim 2 is characterized in that: The step 2 comprises: First, all input window data are instance normalized and then used to construct the histogram, then: in, is the trend sequence of variable d input, μ d and are the mean and variance of the trend sequence respectively, and s represents the time point number in the input sequence; Indicates that it belongs to variable d and has a length of L i The sth time series data point in the input sequence; t′ d represents the normalized input sequence of variable d; σ d represents the standard deviation of the input sequence of variable d; The histogram of each variable is composed of all Construct and use the Gaussian mixture model to approximate the histogram. The number of Gaussian components K of all variables is the same, and the result is the mean M∈R of the K Gaussian components D×K and variance V∈R D ×K ; Construct a Gaussian radial basis function according to the mean and variance of each component. Φ is the cumulative density function of the standard normal distribution, which converts the numerical sequence of the trend into a cumulative probability density sequence Approximation with Page Substituting the standard Φ, we have: Among them, f d,k,: represents the result of processing the input sequence of variable d using the cumulative density function of the kth Gaussian component in GMM; m d,k represents the mean of the kth Gaussian component in the GMM fitting result of the global distribution of variable d; v d,k represents the standard deviation of the kth Gaussian component in the GMM fitting result of the global distribution of variable d; Φ(s) represents the cumulative density function of the standard Gaussian distribution; Φ~(s) represents the approximate calculation formula of the cumulative density function of the standard Gaussian distribution; α1 and α2 represent the constants required for the calculation of Φ~(s); The embedding vector dimension is recorded as e, the hidden vector dimension is recorded as h, and the inverted attention mechanism is used as the skeleton model. First, the cumulative probability density sequence F is transformed into Embedding coding is performed, and then the learnable time series vector pe∈R is added 1 ×1×e , after layer normalization Layer norm and random dropout, we get H0∈R D×K×e As the first layer input, the two dimensions D and K of the variable and Gaussian components are regarded as channels, so the two are merged, and we have: H0=Dropout(Layernorm(for+F*E T )) For the i+1th layer, given H i Afterwards, the multi-head inverted attention mechanism is combined with the two-layer fully connected network MLP to model the relationship between different channels. With H i Keeping the same dimensions, we have: Among them, H i is the input of the i-th layer network, Yes H i The intermediate result after being processed by the inverted attention module of the i-th layer network, yes The output after being processed by the MLP module of the i-th layer network; A hierarchical structure is also constructed. If there are multiple encoder layers, between each two adjacent layers, the adjacent vectors output by the previous layer are merged by linear projection, and then layer normalization is performed, then: Among them, W merge,i-1 represents the projection matrix required to merge adjacent vectors in the output vector of the i-1th layer; b merge,i-1 Represents the projection bias required to merge adjacent vectors in the output vector of the i-1th layer; Finally, a logistic regression layer is used as a decoder to predict the output cumulative probability density sequence O T , then: Where σ represents the sigmoid function, which maps the output value to the interval [0,1], σ(x) = 1 / (1+exp(-x)); H l Represents the final output representation after the lth layer of the model; l represents the number of layers of the model; W d represents the projection matrix of the linear transformation of the decoder, d is the first letter of the decoder word; b d Represents the bias vector of the linear transformation of the decoder decoder, where d is the first letter of the decoder word.

4. The multivariate time series prediction method based on trend signal decoupling and prediction optimization according to claim 3 is characterized in that: The step 3 comprises: Given a multi-periodic signal S, first extract its real and imaginary parts into a matrix In the example, the low-frequency signal with the lowest r ratio is retained; then the matrix is ​​used to Embedding and batch normalization are performed, followed by a two-layer fully connected neural network processing, including batch normalization, plus residual connection and linear projection to obtain the output prediction of the period module The expression is: O P =(S1+Dropout(MLP(S1)))*W o +b o Where S1 represents the multi-periodic input signal after low-pass filtering, embedded coding and batch normalization; W o Represents the projection matrix of the linear transformation of the multi-cycle prediction output layer, o represents output; b o Represents the bias vector of the linear transformation of the multi-period prediction output layer, and o represents output.

5. The multivariate time series prediction method based on trend signal decoupling and prediction optimization according to claim 4 is characterized in that: The step 4 comprises: Multi-task learning is adopted, and the sum of binary cross entropy BCE and mean square error MSE is used as the loss function. and are multi-period and trend training labels respectively, then we have: L=L BCE (O T ,Y T )+L MSE (O P ,Y P ) Where L represents the loss function; L BCE Represents the binary cross entropy loss function L BCE (o T ,y T )=y T logo T +(1-y T )log(1-o T ), o T and T are the predicted cdf trend value and the real cdf trend value respectively; L MSE Represents the mean square error loss function o P and P are the predicted period component value and the actual period component value respectively; During the inference process, there are K cumulative probability density outputs from different Gaussian components for each variable and time step. The result with the maximum confidence level is selected from these K outputs, and then restored to the numerical sequence and inverse normalized to obtain the output of trend prediction. Then we have: Among them, k * It means that each variable d has K different Gaussian components corresponding to the cdf prediction sequence, and k* represents the Gaussian component number corresponding to the best sequence selected; It represents the cdf trend prediction output corresponding to the Gaussian component k at the future time t on the variable d; Indicates that the selected cdf prediction at time t in the future on variable d is converted into a numerical prediction and then output after inverse normalization; Approximate function representing the cumulative density function of the standard Gaussian distribution The inverse function of The intermediate result of is calculated by input s; Next, after combining the trend forecast and the multi-period forecast, the inverse short-time Fourier transform (iSTFT) is used to restore the time domain forecast value. The filtered high-frequency signal is replaced by 0 value in the output part, then:

6. A multivariate time series prediction system based on trend signal decoupling and prediction optimization, characterized in that: include: Module M1: Use short-time Fourier transform to separate the trend value series from the original time series signal; Module M2: Convert the trend value sequence into a cumulative probability density sequence. The value of each time step in the cumulative probability density sequence is the cumulative probability density obtained by transforming the radial basis function based on the Gaussian mixture model components. Module M3: uses a lightweight fully connected network layer to share weights between different frequencies, variables, and real or imaginary parts, and uses a low-pass filter to remove high-frequency signal noise; Module M4: Perform model training and inference, merge trend forecast and multi-period forecast, and use inverse short-time Fourier transform to restore the time domain forecast value.

7. The multivariate time series prediction system based on trend signal decoupling and prediction optimization according to claim 6, characterized in that: The module M1 comprises: Given the original time series input For each variable sequence Use short-time Fourier transform to extract the trend, that is, the frequency component with zero phase; the sliding window length is recorded as N, and the step size is recorded as h, d t,k represents the fast Fourier transform result of the kth frequency in the window [t, t+N-1], d t,0 is the zero phase trend, and the rest d t,k is a multi-periodic signal. Due to the conjugate symmetry of the fast Fourier transform, for k>0, d t,k =d t,N-k+1 ,get multi-periodic components; bidirectional reflection filling is used to mitigate boundary effects in the short-time Fourier transform, and the weight w (n) By default, it is set to 1, since the number of windows in the short-time Fourier transform is Where L p is the input length after padding, and the shape of the trend signal is The shape of the multi-periodic signal is Then we have: in, It represents the nth timing point in the timing window [t, t+N-1] with a length of N, starting from the subscript t. j represents the imaginary unit. n represents the number of the timing point in the window, and its value range is {0, 1, ..., N-1}.

8. The multivariate time series prediction system based on trend signal decoupling and prediction optimization according to claim 7 is characterized in that: The module M2 comprises: First, all input window data are instance normalized and then used to construct the histogram, then: in, is the trend sequence of variable d input, μ d and are the mean and variance of the trend sequence respectively, and s represents the time point number in the input sequence; Indicates that it belongs to variable d and has a length of L i The sth time series data point in the input sequence; t′ d represents the normalized input sequence of variable d; σ d represents the standard deviation of the input sequence of variable d; The histogram of each variable is composed of all Construct and use the Gaussian mixture model to approximate the histogram. The number of Gaussian components K of all variables is the same, and the result is the mean M∈R of the K Gaussian components D×K and variance V∈R D×K ; Construct a Gaussian radial basis function according to the mean and variance of each component. Φ is the cumulative density function of the standard normal distribution, which converts the numerical sequence of the trend into a cumulative probability density sequence Approximation with Page Substituting the standard Φ, we have: Among them, f d,k,: represents the result of processing the input sequence of variable d using the cumulative density function of the kth Gaussian component in GMM; m d,k represents the mean of the kth Gaussian component in the GMM fitting result of the global distribution of variable d; v d,k represents the standard deviation of the kth Gaussian component in the GMM fitting result of the global distribution of variable d; Φ(s) represents the cumulative density function of the standard Gaussian distribution; The approximate calculation formula for the cumulative density function of the standard Gaussian distribution; α1 and α2 represent Calculate the required constants; The embedding vector dimension is recorded as e, the hidden vector dimension is recorded as h, and the inverted attention mechanism is used as the skeleton model. First, the cumulative probability density sequence F is transformed into Embedding coding is performed, and then the learnable time series vector pe∈R is added 1 ×1×e , after layer normalization Layer norm and random dropout, we get H0∈R D×K×e As the first layer input, the two dimensions D and K of the variable and Gaussian components are regarded as channels, so the two are merged, and we have: H0=Dropout(Layernorm(for+F*E T )) For the i+1th layer, given H i Afterwards, the multi-head inverted attention mechanism is combined with the two-layer fully connected network MLP to model the relationship between different channels. With H i Keeping the same dimensions, we have: Among them, H i is the input of the i-th layer network, Yes H i The intermediate result after being processed by the inverted attention module of the i-th layer network, yes The output after being processed by the MLP module of the i-th layer network; A hierarchical structure is also constructed. If there are multiple encoder layers, between each two adjacent layers, the adjacent vectors output by the previous layer are merged by linear projection, and then layer normalization is performed, then: Among them, W merge,i-1 represents the projection matrix required to merge adjacent vectors in the output vector of the i-1th layer; b merge,i-1 Represents the projection bias required to merge adjacent vectors in the output vector of the i-1th layer; Finally, a logistic regression layer is used as a decoder to predict the output cumulative probability density sequence O T , then: Where σ represents the sigmoid function, which maps the output value to the interval [0,1], σ(x) = 1 / (1+exp(-x)); H l Represents the final output representation after the lth layer of the model; l represents the number of layers of the model; W d represents the projection matrix of the linear transformation of the decoder, d is the first letter of the decoder word; b d Represents the bias vector of the linear transformation of the decoder decoder, where d is the first letter of the decoder word.

9. The multivariate time series prediction system based on trend signal decoupling and prediction optimization according to claim 8, characterized in that: The module M3 comprises: Given a multi-periodic signal S, first extract its real and imaginary parts into a matrix In the example, the low-frequency signal with the lowest r ratio is retained; then the matrix is ​​used to Embedding and batch normalization are performed, followed by a two-layer fully connected neural network processing, including batch normalization, plus residual connection and linear projection to obtain the output prediction of the period module The expression is: O P =(S1+Dropout(MLP(S1)))*W o +b o Where S1 represents the multi-periodic input signal after low-pass filtering, embedded coding and batch normalization; W o Represents the projection matrix of the linear transformation of the multi-cycle prediction output layer, o represents output; b o Represents the bias vector of the linear transformation of the multi-period prediction output layer, and o represents output.

10. The multivariate time series prediction system based on trend signal decoupling and prediction optimization according to claim 9, characterized in that: The module M4 comprises: Multi-task learning is adopted, and the sum of binary cross entropy BCE and mean square error MSE is used as the loss function. and are multi-period and trend training labels respectively, then we have: L=L BCE (O T ,Y T )+L MSE (O P ,Y P ) Where L represents the loss function; L BCE Represents the binary cross entropy loss function L BCE (o T ,y T )=y T logo T +(1-y T )log(1-o T ), o T and T are the predicted cdf trend value and the real cdf trend value respectively; L MSE Represents the mean square error loss function o P and P are the predicted period component value and the actual period component value respectively; During the inference process, there are K cumulative probability density outputs from different Gaussian components for each variable and time step. The result with the maximum confidence level is selected from these K outputs, and then restored to the numerical sequence and inverse normalized to obtain the output of trend prediction. Then we have: Among them, k * It means that each variable d has K different Gaussian components corresponding to the cdf prediction sequence, and k* represents the Gaussian component number corresponding to the best sequence selected; It represents the cdf trend prediction output corresponding to the Gaussian component k at the future time t on the variable d; Indicates that the selected cdf prediction at time t in the future on variable d is converted into a numerical prediction and then output after inverse normalization; Approximate function representing the cumulative density function of the standard Gaussian distribution The inverse function of The intermediate result of is calculated by input s; Next, after combining the trend forecast and the multi-period forecast, the inverse short-time Fourier transform (iSTFT) is used to restore the time domain forecast value. The filtered high-frequency signal is replaced by 0 value in the output part, then: