Multivariable time sequence prediction method, device and system and model training method

By performing vector encoding on the high-frequency and low-frequency components of multivariate time series, extracting and adaptively fusing signal vectors, and utilizing the attention mechanism to enhance variable representation, the problem that existing methods have difficulty distinguishing between valid signals and noise under low signal-to-noise ratios is solved, achieving more accurate and robust prediction results.

CN120632637AActive Publication Date: 2025-09-12CHENGDU EVERIMAGING SCI & TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511093894.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-09-12
Estimated Expiration
2045-08-06

AI Technical Summary

Technical Problem

Existing multivariate time series prediction methods have difficulty distinguishing between valid signals and noise under low signal-to-noise ratio conditions, resulting in a decrease in prediction performance. In addition, traditional frequency domain analysis methods ignore the valuable information that may be carried in high-frequency components.

Method used

The high-frequency and low-frequency components of the time series are vector-encoded through an embedding layer based on parameter sharing, and the low-frequency signal vector, high-frequency signal vector and noise signal vector are extracted. The signal representation is improved through adaptive fusion and attention mechanism, and finally synchronous reasoning is performed to predict the variable indicators of future time steps.

Benefits of technology

It effectively separates high-frequency valuable information and noise in the signal, improves the accuracy and robustness of multivariate time series prediction, and performs particularly well under low signal-to-noise ratio conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632637A_ABST
    Figure CN120632637A_ABST
Patent Text Reader

Abstract

The invention discloses a multivariable time sequence prediction method, device and system and a model training method, relates to the technical field of information, and is used for improving the accuracy of multivariable time sequence prediction. According to the method, a time sequence is decomposed on a high-frequency component and a low-frequency component, a low-frequency signal vector, a high-frequency signal vector and a noise signal vector are further coded and decoupled, an effective signal part is subjected to self-adaptive fusion and enhancement, prediction is carried out through an effective signal and a noise signal, an inference result is fused, and a prediction result is obtained. According to the method, effective signals in high-frequency components are reserved, the signals and noise are isolated and utilized, variable dependence can be stably captured, the prediction accuracy of the effective signals is improved, random components are predicted through the noise, the noise is added into the predicted effective signals, and the method is better matched with the real environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular to a multivariate time series prediction method, device, system and model training method. Background Art

[0002] Energy management, traffic control, and IoT system scenarios involve managing multi-dimensional variables and indicators. These indicators are collected synchronously in a time series (i.e., multiple indicators are collected simultaneously at the same time, repeatedly, and arranged in a time series), resulting in a multivariate time series. Multivariate time series forecasting (MTSF) is crucial for decision-making in various fields. Forecasting involves predicting the values ​​of each variable indicator at a specific time step in the future based on historically collected multivariate time series. For example, in traffic control, variables include vehicle flow, pedestrian flow, vehicle type demographics, precipitation, temperature, and wind speed. Time series forecasting involves predicting vehicle flow, pedestrian flow, vehicle type demographics, precipitation, temperature, and wind speed at each time step (e.g., every 5 minutes) over a period of time (e.g., the next hour). This provides data support for traffic management and control decisions for the next hour.

[0003] Currently, most known time series forecasting methods use deep learning architectures such as RNNs (Recurrent Neural Networks), CNNs (Convolutional Neural Networks), and Transformers to train time series forecasting models and predict variable indicators at future time steps. However, these methods are often limited by the inherently low signal-to-noise ratio (SNR) of real-world time series, making it difficult to distinguish valid signals from noise, significantly reducing forecasting performance.

[0004] To address the SNR challenge, traditional approaches use methods such as sliding window filtering or wavelet decomposition in the time domain to filter out noise. However, signal processing in the time domain is complex and labor-intensive. In contrast, transforming the signal to the frequency domain and separating the valid signal from the noise through spectral analysis can isolate the valid signal with less computational effort, thereby improving prediction accuracy.

[0005] However, even though frequency domain analysis can reduce the workload of denoising, current frequency domain analysis methods often directly select the low-frequency components of the signal, crudely filtering out the high-frequency components as noise. This seriously overlooks the valuable information that these high-frequency components may carry. Research and analysis have shown that high-frequency components are not simply unstructured noise; they often contain valuable signals reflecting short-term changes (such as sudden changes in energy demand) and emergencies (such as IoT system equipment failures or traffic network anomalies). If current methods indiscriminately suppress or eliminate high-frequency components, while discarding noise, they also lose critical information crucial for accurate predictions, resulting in poor multivariate time series forecasting accuracy. Summary of the Invention

[0006] The object of the present invention is to provide a multivariate time series prediction method, device, system and model training method to improve the accuracy of multivariate time series prediction in order to address all or part of the above-mentioned problems.

[0007] The technical solution adopted in the present invention is as follows: A multivariate time series prediction method, wherein the time series is obtained by synchronously collecting multiple variable indicators in a time sequence; the method comprises: Get the time series of the current scenario; Decomposing the acquired time series into high-frequency components and low-frequency components based on a preset frequency threshold; Performing vector encoding on the high-frequency component and the low-frequency component respectively using a parameter-sharing embedding layer; Extracting a low-frequency signal vector, a high-frequency signal vector and a noise signal vector from the encoding vectors of the high-frequency component and the low-frequency component respectively; Adaptively fusing the low-frequency signal vector and the high-frequency signal vector to obtain a fused signal vector; Using the low-frequency signal vector to enhance the attention of the fused signal vector on variable dependency, thereby obtaining an enhanced variable representation vector; Synchronous reasoning is performed on the enhanced variable representation vector and the noise signal vector respectively, and the reasoning results are integrated to obtain the prediction result of the future time step.

[0008] Furthermore, extracting a low-frequency signal vector, a high-frequency signal vector, and a noise signal vector from the encoding vectors of the high-frequency component and the low-frequency component, respectively, includes: According to the principle of maximizing the mutual information between the low-frequency signal vector and the high-frequency signal vector and minimizing the mutual information between the low-frequency signal vector and the noise signal vector, the low-frequency signal vector is extracted from the coding vector of the low-frequency component, and the high-frequency signal vector and the noise signal vector are respectively extracted from the coding vector of the high-frequency component.

[0009] Furthermore, the method of maximizing the mutual information between the low-frequency signal vector and the high-frequency signal vector and minimizing the mutual information between the low-frequency signal vector and the noise signal vector includes: Reconstructing the low-frequency signal vector using the high-frequency signal vector and calculating the first reconstruction loss; Reconstructing the low-frequency signal vector using the noise signal vector and calculating the second reconstruction loss; combining the first reconstruction loss and the second reconstruction loss to obtain a total reconstruction loss; Minimize the total reconstruction loss.

[0010] Furthermore, adaptively fusing the low-frequency signal vector and the high-frequency signal vector to obtain a fused signal vector includes: Splicing the low-frequency signal vector and the high-frequency signal vector in a feature dimension to obtain a spliced ​​vector; Processing the concatenated vector using a multi-layer perceptron, adjusting it using a temperature coefficient, activating the contribution of the low-frequency signal vector, and calculating the contribution of the high-frequency signal vector; The low-frequency signal vector and the high-frequency signal vector are weightedly fused using corresponding contribution degrees to obtain the fused signal vector.

[0011] Furthermore, utilizing the low-frequency signal vector to enhance the attention of the fused signal vector on variable dependency includes: Calculating a Gram matrix of the fused signal vector using the low-frequency signal vector and adaptively optimizing the vector using trainable parameters to obtain a similarity matrix; The fusion signal vector is weighted by using the similarity matrix and is residually connected with the fusion signal vector, and is post-processed by using a feedforward neural network after linear transformation.

[0012] Furthermore, the Gram matrix of the fused signal vector is calculated using the low-frequency signal vector, and adaptive optimization is performed using trainable parameters, including: performing a layer normalization operation on the low-frequency signal vector, and then multiplying the low-frequency signal vector by the layer normalized and transposed low-frequency signal vector to obtain a Gram matrix; After activating the Gram matrix, perform element-wise multiplication with the weight matrix and add the bias vector; An activation operation is performed on the calculation result to obtain the similarity matrix.

[0013] On the other hand, the present application also provides a multivariate time series prediction device, including a processor and a storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by the processor, it executes the above-mentioned multivariate time series prediction method.

[0014] As an intermediate product of multivariate time series prediction, the present application also provides a multivariate time series prediction model training method. The multivariate time series prediction model is used to predict a time series of future time steps based on a current input time series. The time series is obtained by synchronously collecting multiple variable indicators in a time series. The training method includes: Obtain a set of time series samples and divide them into training and test sets; Using the training set to train the multivariate time series prediction model with the goal of minimizing training loss, and using the test set to test the trained multivariate time series prediction model; The multivariate time series forecasting model is configured to: Get the input time series; Decomposing the acquired time series into high-frequency components and low-frequency components based on a preset frequency threshold; Performing vector encoding on the high-frequency component and the low-frequency component respectively using a parameter-sharing embedding layer; Extracting a low-frequency signal vector, a high-frequency signal vector and a noise signal vector from the encoding vectors of the high-frequency component and the low-frequency component respectively; Adaptively fusing the low-frequency signal vector and the high-frequency signal vector to obtain a fused signal vector; Using the low-frequency signal vector to enhance the attention of the fused signal vector on variable dependency, thereby obtaining an enhanced variable representation vector; Performing synchronous reasoning on the enhanced variable representation vector and the noise signal vector respectively, and fusing the reasoning results to obtain a prediction result for the future time step; The training loss is composed of the mean square error (MSE) between the predicted results and the true results of all variables in the time series samples of the training set at each time step, and the weighted total reconstruction loss in the process of extracting the low-frequency signal vector, the high-frequency signal vector, and the noise signal vector.

[0015] On the other hand, the present application also provides another multivariate time series prediction device, which is configured with a multivariate time series prediction model trained using the above-mentioned multivariate time series prediction model training method.

[0016] On the other hand, the present application also provides a multivariate time series prediction system, which includes an input device, an output device and a multivariate time series prediction device; the input device is connected to the input end of the multivariate time series prediction device, for receiving the time series to be predicted and inputting it into the multivariate time series prediction model; the output device is connected to the output end of the multivariate time series prediction device, for outputting the time series predicted by the multivariate time series prediction model.

[0017] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are: This application addresses the need for multivariate time series prediction and provides a multivariate time series prediction method that separates valuable signals and noise from the high-frequency components of the signal, and adaptively fuses and enhances the valuable signals with the low-frequency signals, so that it can not only effectively extract valuable information, but also robustly capture the spatial and temporal dependencies between variables, and more accurately extract signal features. In addition, the noise signal is used to predict future random components and is added to the prediction results of the effective signal, so that the final prediction results are more consistent with the actual environment. This application not only extracts effective signals (especially effective signals in high-frequency components) more effectively, but also utilizes noise signals, thereby improving the prediction accuracy of multivariate time series in real-world application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The present invention will now be described by way of example with reference to the accompanying drawings, in which: Figure 1 FIG. 1 is a flowchart of a multivariate time series forecasting method in one embodiment.

[0019] Figure 2 3 is a construction diagram of a multivariate time series prediction model in one embodiment, wherein sub-figure (a) shows a decomposition method of high-frequency components and low-frequency components, sub-figure (b) shows a decoupling method of low-frequency signal vectors, high-frequency signal vectors, and noise signal vectors, sub-figure (c) shows a construction process of a fusion signal vector, and sub-figure (d) shows a construction process of an enhanced variable representation vector.

[0020] Figure 3 This is a performance comparison chart of this application and the baseline model prediction. Figure 3 Sub-figure (a), sub-figure (b) and sub-figure (c) are test figures showing the mean square error of the proposed method, the iTransformer model and the FilterNet model on the PEMS04 dataset, the ETTm1 dataset and the Traffic dataset as the signal-to-noise ratio changes.

[0021] Figure 4 This is a performance test chart of this application at different prediction lengths. Figure 4 Sub-figures (a) and (b) are test figures showing how the mean square error of the proposed method, the FilterNet model, the iTransformer model, the PatchTST model, and the Crossformer model on the Electricity dataset and the Traffic dataset changes with the length of the input time series.

[0022] Figure 5is a test graph of the regularization coefficient configuration in different embodiments, Figure 5 Sub-graph (a), sub-graph (b) and sub-graph (c) are the mean square error of the proposed method on the ECL dataset, Traffic dataset and Solar-Energy dataset respectively as the regularization coefficient Changing test pattern. DETAILED DESCRIPTION

[0023] All features disclosed in this specification, or all steps in the disclosed methods or processes, except mutually exclusive features and / or steps, can be combined in any manner.

[0024] Any feature disclosed in this specification (including any appended claims and abstract), unless otherwise stated, may be replaced by other equivalent or similar features. In other words, unless otherwise stated, each feature is only an example of a series of equivalent or similar features.

[0025] In response to the problem of poor accuracy in multivariate time series prediction caused by the direct filtering of high-frequency components for denoising in existing frequency domain analysis methods, the embodiments of the present application propose a multivariate time series prediction method, device, system and model training method, which retains the effective signal of the high-frequency components while also utilizing the noise signal, thereby improving the accuracy of the multivariate time series in display application scenarios.

[0026] like Figure 1 As shown, in one embodiment, the multivariate time series prediction method proposed in this application includes the following process: S1. Get the time series of the current scenario.

[0027] The method of this application can be applied in scenarios such as traffic control, medical testing, and weather forecasting. In traffic control scenarios, the variable indicators included in the time series can be traffic flow, pedestrian flow, vehicle type (such as vans, cars, trucks) statistics, rainfall / snowfall, temperature, etc. In medical testing scenarios, the variable indicators included in the time series can be physiological indicators (such as blood pressure, heart rate, etc.), test indicators (platelet count, blood oxygen saturation, white blood cell count, etc.), and detection indicators (ECG trends, nodule size, etc.). In weather forecasting scenarios, the variable indicators included in the time series can be temperature, rainfall / snowfall, wind speed, light intensity, humidity, etc.

[0028] The value of each variable indicator is recorded at the corresponding time step according to the predetermined frequency / step size. The multivariate time series is obtained by following the chronological order of the records. Each variable can be obtained as a one-dimensional vector, and multiple variables can form a two-dimensional vector matrix.

[0029] According to the predicted input window size, the recorded multivariate time series is slidingly sampled to obtain a time series of predetermined length.

[0030] S2. Decomposing the acquired time series into high-frequency components and low-frequency components based on a preset frequency threshold.

[0031] As an optional implementation, see the attached Figure 2 In the sub-figure (a), the decomposition methods of the high-frequency components and low-frequency components of the time series include: S21. Transform the time series from the time domain to the frequency domain.

[0032] The acquired time series is transformed using Fast Fourier Transform (FFT) Transform from time domain to frequency domain and decompose it into frequency points and their corresponding amplitudes, where R represents the real number domain, N represents the number of variable indicators, and L represents the length of the variable time series. Express Round up.

[0033] The above transformation process is expressed as: Formula (1): ; Among them, FFT stands for Fast Fourier Transform, Represents the time series after transformation into the frequency domain.

[0034] S22, passing the set frequency threshold Distinguish between high-frequency and low-frequency components.

[0035] The frequency threshold In some embodiments, it is a ratio threshold. The process of dividing high-frequency components and low-frequency components is: The low-frequency component corresponds to the front frequency points, and the high-frequency components correspond to the remaining frequency points. In addition, in order to ensure that the high-frequency components and low-frequency components are included frequency points to facilitate signal recovery in the time domain, and to recover the high-frequency and low-frequency components respectively. The length portion is filled with zero value.

[0036] definition , represents the number of frequency points (length) of the low-frequency component, then the low-frequency component (frequency domain) It is expressed by the following formula (2): Formula (2): .

[0037] High-frequency components (frequency domain) It is expressed by the following formula (3): Formula (3): .

[0038] S23. Convert the high-frequency component and the low-frequency component from the frequency domain to the time domain respectively.

[0039] The inverse fast Fourier transform (IFFT) is used to restore the high-frequency and low-frequency components in the frequency domain to the time domain, thereby generating a low-frequency component that represents the main signal and a high-frequency component that captures short-term fluctuations and noise.

[0040] Low-frequency component after transformation to time domain and high frequency components Expressed as: Formula (4): ; Where IFFT stands for inverse fast Fourier transform.

[0041] S3. Use the parameter-sharing embedding layer to vector encode the high-frequency component and the low-frequency component respectively.

[0042] Vector coding is to project the time domain signal into the vector space. The low-frequency component and the high-frequency component are linearly transformed to obtain the low-frequency component coding vector in the form of vector representation. and high frequency component encoding vector : Formula (5): ; Where, Indicates a linear transformation. D Represents the feature dimension.

[0043] S4. Extract the low-frequency signal vector, the high-frequency signal vector and the noise signal vector from the encoding vectors of the high-frequency component and the low-frequency component respectively.

[0044] As an optional implementation, the method of extracting the low-frequency signal vector, the high-frequency signal vector, and the noise signal vector includes: According to the principle of maximizing the mutual information between the low-frequency signal vector and the high-frequency signal vector and minimizing the mutual information between the low-frequency signal vector and the noise signal vector, the low-frequency signal vector is extracted from the coding vector of the low-frequency component, and the high-frequency signal vector and the noise signal vector are respectively extracted from the coding vector of the high-frequency component.

[0045] In some feasible embodiments of the present application, three parallel feature extractors (feature encoders) are used to separate three independent potential representations: the low-frequency signal vector , high-frequency signal vector and noise signal vector . Assume that the three feature extractors are , then the three potential representations are expressed as: Formula (6): ; Formula (7):

[0046] Formula (8): .

[0047] in, .

[0048] After analysis, there are also effective signals in the high-frequency components, so it can be assumed that and There is a potential time dependency between In terms of statistical laws Therefore, in this application, according to the maximization of the low-frequency signal vector With high frequency signal vector Mutual information, minimize the low-frequency signal vector With the noise signal vector Based on the mutual information principle, the following optimization objective function is designed to minimize: Formula (9): ; Where, J represents the objective function, Indicates mutual information, for example Express request and The mutual information of Express request and By minimizing J , which can maximize the low-frequency signal vector With high frequency signal vector Mutual information , minimize the low-frequency signal vector With the noise signal vector Mutual information .

[0049] In addition, as a possible optimization method, see Figure 2 In subgraph (b) of the embodiment of the present application, in some optional implementations, the information entropy calculation method is used to optimize the above objective function. The above method of maximizing the mutual information between the low-frequency signal vector and the high-frequency signal vector and minimizing the mutual information between the low-frequency signal vector and the noise signal vector includes: S41. Reconstruct the low-frequency signal vector using the high-frequency signal vector and calculate a first reconstruction loss; reconstruct the low-frequency signal vector using the noise signal vector and calculate a second reconstruction loss.

[0050] This application uses conditional entropy to decompose mutual information. The principle is: Formula (10): ; Where, represents the mutual information between parameter A and parameter B, represents the information entropy of parameter A, It represents the conditional entropy of parameter A under the condition of parameter B.

[0051] Based on the above principles, the objective function can be obtained J The conversion form: Formula (11): .

[0052] In formula (11), and are all conditional entropies, where: Formula (12): ; Where, g For the reconstruction function, E Expressing hope, Express request It can be seen that the method of reducing the reconstruction error will directly reduce the conditional entropy. Therefore, in this application, the mean square error is used as the first reconstruction loss , which is designed to: Formula (13): ; Where, Indicates that the parameter group Parameterized MLP (Multilayer Perceptron) model, Indicates that the MLP model is used by Reconstruct the corresponding low-frequency signal vector, Figure 2 The subgraph (b) of . and They are and The i This application minimizes the conditional entropy when optimizing the first reconstruction loss. This can in turn maximize mutual information This mechanism ensures that the high-frequency signal vector can retain frequency information that is highly consistent with the low-frequency signal vector, while effectively filtering out irrelevant noise.

[0053] The other part after the objective function is converted is the conditional entropy , the optimization goal is to maximize ,maximize This in turn minimizes Since the low frequency signal vector With high-frequency noise vector There is statistical independence between them, so it is necessary to estimate the statistical independence of the two separately.

[0054] Due to minimization This is equivalent to forcing the orthogonality of the two in the vector space. Therefore, in some optional embodiments of the present application, a set constraint based on cosine similarity is designed to construct the second reconstruction loss, thus avoiding Towards The problem of degenerate solution of reconstruction. Specifically, a connection loss is designed on the negative cosine similarity to construct the second reconstruction loss : Formula (14): ; In the formula, Cosine represents the cosine similarity function, (like ) represents the separation boundary established in the angle domain, express The i-th variable in express and The vector dot product of .

[0055] S42. Combine the first reconstruction loss and the second reconstruction loss to obtain a total reconstruction loss.

[0056] Total reconstruction loss Expressed as: Formula (15): .

[0057] S43, minimizing the total reconstruction loss. Minimizing the total reconstruction loss maximizes the mutual information between the low-frequency signal vector and the high-frequency signal vector, and minimizes the mutual information between the low-frequency signal vector and the noise signal vector.

[0058] S5, adaptive fusion of low-frequency signal vectors and high-frequency signal vector , get the fusion signal vector .

[0059] As an optional implementation, see the attached Figure 2 Sub-image (c) in the figure shows adaptive fusion of low-frequency signal vectors. and high-frequency signal vector The methods include: S51, the low frequency signal vector and high-frequency signal vector Splicing is performed on the feature dimension to obtain a splicing vector. The splicing vector is expressed as .

[0060] S52, use the multi-layer perceptron MLP to process the splicing vector, and use the temperature coefficient to adjust it, and activate it to obtain the low-frequency signal vector The contribution of the high-frequency signal vector is calculated contribution.

[0061] by Indicates the temperature coefficient, Represents the low-frequency signal vector The contribution of the high-frequency signal vector The contribution of .

[0062] Formula (16): ; Where MLP represents multi-layer perceptron, Represents the sigmoid activation function.

[0063] S53, using the corresponding contribution to perform weighted fusion on the low-frequency signal vector and the high-frequency signal vector, to obtain a fused signal vector .

[0064] Formula (17): ; Where, Represents element-by-element multiplication. Temperature coefficient Assists in regulating gating behavior during training.

[0065] Through the above-mentioned gating mechanism, the present application realizes the adaptive integration of low-frequency signal vectors and high-frequency signal vector , thereby dynamically learning a more comprehensive time series representation.

[0066] After completing the low-frequency signal vector and high-frequency signal vector After adaptive fusion, the present application also models the dependencies between variables through a learnable attention mechanism. Specifically, the method also includes: S6. Using low-frequency signal vectors Improved fusion signal vector Attention on variable dependencies, to obtain enhanced variable representation vectors .

[0067] As an optional implementation, see Figure 2 In sub-graph (d), in step S6, the low-frequency signal vector Improved fusion signal vector Methods for attention on variable dependencies include: S61, using low-frequency signal vector Calculate the fused signal vector Gram matrix , and utilizes trainable parameters ( ) to perform adaptive optimization and obtain the similarity matrix .

[0068] In some specific embodiments, step S61 obtains the similarity matrix by the following method: : The low-frequency signal vector Perform layer normalization operation and then compare it with the low-frequency signal vector that has been normalized and transposed Multiply them together to get the Gram matrix . Expressed as: Formula (18): ; Where, Representation layer normalization operation.

[0069] Gram matrix After activation, with the weight matrix Perform element-wise multiplication and add a bias vector .

[0070] Perform activation operation on the calculation results to obtain the similarity matrix .

[0071] The above two operations are expressed as: Formula (19): ; Where, Represents the activation function.

[0072] S62. Using Similarity Matrix The fusion signal vector weighted and combined with the fusion signal vector Perform residual connection and use at least one level of feedforward neural network after linear transformation (Feedforward NeuralNetwork) post-processing (multi-level feedforward neural network represented as FFNs) to obtain the enhanced variable representation vector . Expressed as: Formula (20): ; Where, represents a feedforward neural network, represents linearization processing, Represents the residual connection operation.

[0073] S7, respectively enhance the variable representation vector and noise signal vector Perform synchronous reasoning and fuse the reasoning results to obtain prediction results for future time steps.

[0074] like Figure 2 As shown, in the prediction stage, the enhanced variable representation vector and noise signal vector The reasoning mapping is performed through independent prediction layers (Project) respectively, so that the hidden dimensions of the two prediction layers are aligned with the prediction time domain (future time steps) T. By the low frequency signal vector and high-frequency signal vector The fusion and enhancement result contains the main signal features; It captures additional high-frequency components, such as potential anomalies or irregular events, and characterizes the impact of the real world on the effective signal. The final prediction result is expressed as : Formula (21): ; Where, Represent the enhanced variable representation vector and noise signal vector The prediction layer.

[0075] When the network parameters of each layer are determined, the above method can be used to predict a (multivariate) time series of predetermined length T from the acquired time series.

[0076] In this regard, an embodiment of the present application also provides a multivariate time series prediction model training method for training to obtain network parameters of relevant network layers. The multivariate time series prediction model is used to predict the time series of future time steps based on the input current time series.

[0077] Training methods include: Step 1: Get a set of time series samples and divide them into training set and test set.

[0078] Time series samples can be obtained by collecting historical time series. The ratio of the training set to the test set is, for example, 7:3 or 8:2.

[0079] Step 2: Use the training set to train the multivariate time series prediction model with the goal of minimizing the training loss, and use the test set to test the trained multivariate time series prediction model.

[0080] The multivariate time series forecasting model is configured as follows: S1. Get the input time series.

[0081] S2. Decomposing the acquired time series into high-frequency components and low-frequency components based on a preset frequency threshold.

[0082] S3. Use the parameter-sharing embedding layer to vector encode the high-frequency component and the low-frequency component respectively.

[0083] S4. Extract the low-frequency signal vector, the high-frequency signal vector and the noise signal vector from the encoding vectors of the high-frequency component and the low-frequency component respectively.

[0084] S5. Adaptively fuse the low-frequency signal vector and the high-frequency signal vector to obtain a fused signal vector.

[0085] S6. Use the low-frequency signal vector to enhance the attention of the fusion signal vector on the variable dependency, and obtain an enhanced variable representation vector.

[0086] S7. Perform synchronous reasoning on the enhanced variable representation vector and the noise signal vector respectively, and fuse the reasoning results to obtain the prediction results for the future time steps.

[0087] The above training loss is composed of the mean square error (MSE) between the predicted results and the true results of all variables in the time series samples of the training set at each time step, and the weighted total reconstruction loss in the process of extracting low-frequency signal vectors, high-frequency signal vectors, and noise signal vectors.

[0088] The features that can be further designed in each step of step S1 to step S7 in the above-mentioned model training method embodiment can refer to the design of the prediction method in each embodiment above, and will not be described one by one here.

[0089] As for the training loss involved in the training, assuming that the principle of maximizing the mutual information between the low-frequency signal vector and the high-frequency signal vector and minimizing the mutual information between the low-frequency signal vector and the noise signal vector is still followed, the low-frequency signal vector is extracted from the coding vector of the low-frequency component, and the high-frequency signal vector and the noise signal vector are extracted from the coding vector of the high-frequency component respectively. Then, the total reconstruction loss in the process of extracting the low-frequency signal vector, the high-frequency signal vector and the noise signal vector can be used as the total reconstruction loss of the previous formula (15). .

[0090] The mean square error (MSE) between the predicted results and the actual results of all variables in the time series samples of the training set at each time step is calculated by the following method: Formula (22): ; Where, This is the required mean square error, Respectively represent i The variable in the future t The predicted and true values ​​for each time step.

[0091] Therefore, the training error Expressed as: Formula (23): ; Where, is the hyperparameter of the model, which represents the The weighted regularization coefficient is used to adjust the influence of the regularization term to ensure a balanced training objective. It can be determined by hyperparameter optimization algorithm or set based on experience (such as =1, 2), but it is certain that for different application scenarios, it is often necessary to adjust the optimal hyperparameters based on the historical data of the application field. This can be done by setting different coefficient values ​​on the historical data for comparison, and then selecting the one with the best performance. .

[0092] Through the above process, a multivariate time series prediction model can be obtained by training historical time series. This module not only considers the effective components of the high-frequency components of the signal, but also the actual impact of the noise signal in the high-frequency components, making the prediction results more accurate and more consistent with real-world scenarios.

[0093] According to the concept of the present application, an embodiment of the present application also provides a multivariate time series prediction device, including a processor and a storage medium, the storage medium storing a computer program, and when the computer program is run by the processor, executing the multivariate time series prediction method of the above embodiment.

[0094] In addition, an embodiment of the present application also provides another multivariate time series prediction device, which is configured with a multivariate time series prediction model trained using the above-mentioned multivariate time series prediction model training method.

[0095] Based on the above-mentioned prediction device, an embodiment of the present application further provides a multivariate time series prediction system, which includes an input device, an output device, and a multivariate time series prediction device. The input device is connected to the input end of the multivariate time series prediction device and is used to receive the time series to be predicted and input it into the multivariate time series prediction model; the output device is connected to the output end of the multivariate time series prediction device and is used to output the time series predicted by the multivariate time series prediction model.

[0096] Furthermore, the present embodiments also conducted systematic experiments on 12 real-world multivariate time series datasets. These datasets include ECL (electricity load), Traffic, Weather, Solar-Energy, four PEMS (traffic flow) datasets (PEMS03, PEMS04, PEMS07, and PEMS08), and four ETT (electric transformer temperature) datasets (ETTh1, ETTh2, ETTm1, and ETTm2). These datasets are widely used benchmarks in the field of multivariate time series forecasting. The key features of these datasets are shown in Table 1.

[0097] Table 1. Characteristics of multivariate time series dataset

[0098] To fully evaluate the effectiveness of this application's solution, three baseline models were selected for performance comparison in the examples of this application. The baseline models include: (i) Transformer-based architectures, including the iTransformer model, the PatchTST model, and the FEDformer model; (ii) MLP-based architectures, including the FilterNet model, the SOFTS model, the TimeMixer model, the TSMixer model, the FreTS model, the TiDE model, and the DLinear model; and (iii) the CNN-based model, the TimesNet model.

[0099] Experimental Conditions: To ensure consistency across all benchmarks, a consistent history window length of 96 time steps was used across all benchmarks (i.e., the input historical time series length was 96). Forecast durations were configured as follows: The PEMS dataset was designed with four different prediction lengths: {12, 24, 48, and 96}. The remaining benchmark datasets were tested using four different prediction lengths: {96, 192, 336, and 720}.

[0100] In order to evaluate the prediction performance of each prediction method and facilitate quantitative comparison, the embodiment of this application uses the mean square error (MSE) and the mean absolute error (MAE) as performance evaluation indicators. The smaller the value of both indicators, the higher the prediction accuracy. Definition represents the predicted value of time step t, and the true value of the corresponding time step is expressed as , T represents the total number of prediction steps. The calculation method of the above two indicators is as follows: Formula (24): .

[0101] The experimental results are shown in Table 2. Due to the large amount of data, Table 2 only shows the test indicators under the longest prediction length (i.e., 96 or 720).

[0102] Table 2 Performance index test table (excerpt)

[0103] Table 2 Performance test table (continued)

[0104] Table 2 Performance test table (continued)

[0105] Experimental results show that the proposed method demonstrates significant advantages over baseline models across various prediction time domains and datasets, achieving significant improvements in both MSE and MAE (reduced metrics, improved performance). Experimental results across 12 datasets demonstrate that the proposed method achieves the lowest MSE in 37 test cases and the lowest MAE in 44 test cases. While currently recognized state-of-the-art time-domain prediction models (such as SOFTS and iTransformer) also focus on modeling relationships between variables, their noise sensitivity distorts these relationships, leading to incorrect learned dependencies. This proposed method effectively mitigates this problem by isolating noise signals from valid signals, capturing more robust variable dependencies from the valid signal alone, demonstrating superior performance. Meanwhile, the mainstream frequency-domain model, FilterNet, selectively filters time series components. While this approach improves model robustness to a certain extent, it also leads to information loss, particularly in datasets with numerous variables, such as traffic data and PEMS. This proposed method innovatively utilizes a decoupling mechanism to separate signal from noise and independently characterize each component. This design enables this application to retain key information while effectively controlling noise, thereby demonstrating stronger prediction performance in all benchmarks.

[0106] Furthermore, to evaluate the robustness of this application under diverse noise conditions, comparative experiments were conducted on the PEMS04, ETTm1, and Traffic datasets. Two state-of-the-art baseline models, the FilterNet model and the iTransformer model, were selected for benchmarking. To simulate varying noise levels, Gaussian white noise was injected into the input data during training, with target signal-to-noise ratios (SNRs) of {−10, −5, 0, 5, 10, 20} dB. For each SNR value, the signal power was calculated, the corresponding linear noise power was determined, and Gaussian noise was superimposed according to implementation details. This method effectively evaluates the robustness of the model by controlling the gradual degradation of input quality.

[0107] like Figure 3As shown, the present application can maintain a low mean square error in all datasets and signal-to-noise ratio levels. It is worth noting that FilterNet's performance on the ETTm1 dataset drops significantly in a low signal-to-noise ratio environment, which may be related to its dependence on key frequency components, that is, these components are easily masked by noise. Similarly, iTransformer's performance will also decline at low signal-to-noise ratios because excessive noise will interfere with the attention mechanism's ability to capture effective correlations between variables. In contrast, the present application shows more stable performance: compared with FilterNet and iTransformer, its MSE curve remains flat at different signal-to-noise ratio levels, fully demonstrating its excellent anti-interference ability to noise perturbations.

[0108] In addition, in order to verify the scientificity and necessity of the design of each link of this application (such as frequency decoupling), an ablation experiment was also conducted in the embodiment of this application.

[0109] In the embodiment of the present application, the ECL, Traffic and Solar-Energy datasets are removed. Figure 2 In the experiment, the input sequence length was fixed at 96, and the prediction duration was set to 96, 192, 336, and 720 time steps, respectively. As shown in Table 3, the complete model of the present application consistently outperforms the ablation model with the frequency decoupling module removed under all benchmark tests and prediction durations. It is worth noting that the performance gap gradually widens with the extension of the prediction duration. This phenomenon was attributed to the fact that the absence of the frequency decoupling module weakened the model's ability to distinguish between signals and noise. High-frequency noise interferes with low-frequency signals, obscuring potential long-term trends, making it difficult to capture the true correlation between real variables, and ultimately affecting the overall performance.

[0110] Table 3 Ablation experiment test table

[0111] In addition, the embodiments of the present application also target the phenomenon that the prediction accuracy of the currently widely used models based on the Transformer architecture (such as the Informer model, the Autoformer model, and the FEDformer model) decreases when the length of the input time series is extended. The performance of the present application scheme in this regard is experimented to verify the sensitivity of the present application scheme to parameters.

[0112] The embodiment of this application compares the solution of this application with the FilterNet model, iTransformer model, PatchTST model and Crossformer model through experiments. The experiment uses the ECL dataset and Traffic dataset, with input sequence lengths of {48, 96, 192, 336, 512} respectively, and a fixed prediction time of 96 time steps. The experimental results are shown in Figure 2. Figure 4 As shown, the results show that our application consistently outperforms all baseline models across all input sequence lengths in both datasets. Notably, while increasing input length typically increases the signal-to-noise ratio, our application's performance decline is more gradual than that of other models. This smoother performance curve demonstrates that even with shorter input sequences (which inherently have lower noise levels), our application can still effectively distinguish between signal and noise, significantly improving overall prediction robustness.

[0113] As mentioned above, hyperparameters Used to adjust the regularization term in the training loss (i.e. ) to ensure the balance of training objectives. The embodiment of the present application also tests the impact of this hyperparameter on the training effect.

[0114] To evaluate the regularization coefficient in the training loss In order to investigate the impact of the regularization coefficient, this paper conducted a series of experiments using the ECL dataset, Traffic dataset, and Solar-Energy dataset. For each dataset, the input sequence length was set to 96 and the prediction duration was set to 720. The model was adjusted in the range of {0.01, 0.1, 1.0, 2.0, 5.0, 10.0} to evaluate its impact on model performance. The experimental results are shown in the figure. Figure 5 shown.

[0115] Depend on Figure 5 It can be seen that in the ECL data set, the mean square error increases with This indicates that stronger regularization can effectively improve the effect of the decomposition process. In contrast, the mean square error of the Traffic dataset and the Solar-Energy dataset shows a non-monotonic trend. The performance reaches its peak when the value is in the middle range (about 1.0 or 2.0), and then shows a downward trend as the value increases. This shows that excessive regularization may suppress key time features, resulting in information loss and reduced prediction accuracy. Experimental results show that the regularization coefficient needs to be adjusted according to the characteristics of each dataset. Fine-tuning is required to achieve the best performance. For example, for multivariate time series prediction tasks in the transportation field, during the training phase, Designed as 1 or 2, for multivariate time series prediction tasks in the power sector, during the training phase, Designed to be a larger value (such as 5 to 10).

[0116] In summary, this application decomposes high-frequency signals and noise signals separately, fully exploits the characteristics of effective signals, and retains the influence of signals and noise on prediction results, thereby improving prediction accuracy. The decoupling method based on mutual information can not only enhance the anti-noise ability, but also capture the characteristics of short-term fluctuations and emergencies, further improving the matching degree between prediction results and real-world scenarios. Compared with the current method of directly discarding high-frequency components, this application significantly improves the prediction performance of multivariate time series by retaining high-frequency components and decomposing high-frequency signals and noise for prediction.

[0117] It should be noted that the multivariate time series prediction method or multivariate time series prediction model training method of the present application can be applied to a variety of multivariate time series prediction fields, such as the transportation field, weather field, energy field, power field, medical field, etc. introduced in the previous embodiment. In specific applications, the multivariate time series obtained is composed of various variables in the corresponding scenario. For example, when applied in the weather field, it is a weather time series prediction method. The multiple variable indicators include temperature, humidity, rainfall / precipitation, wind force, light intensity, etc. By collecting the historical time series of these variables to construct a time series sample, and inputting it into the aforementioned multivariate time series prediction model for training, a weather time series prediction model can be obtained. The weather time series prediction model can be used to predict the future time steps (the duration is set in advance) of the input weather time series. Another example of this application in the transportation sector is a traffic time series prediction method. Multiple variable indicators, such as traffic flow data, speed data, and density data, are collected from historical time series to construct time series samples. These samples are then fed into the aforementioned multivariate time series prediction model for training. This yields a traffic time series prediction model, which can then be used to predict future time steps of the input traffic time series. The same principle applies to other fields.

[0118] The present invention is not limited to the aforementioned specific embodiments, but extends to any new features or any new combination disclosed in this specification, as well as any new method or process steps or any new combination disclosed.

Claims

1. A multivariate time series prediction method, wherein the time series is obtained by synchronously collecting multiple variable indicators in a time sequence; characterized in that: Methods include: Get the time series of the current scenario; Decomposing the acquired time series into high-frequency components and low-frequency components based on a preset frequency threshold; Performing vector encoding on the high-frequency component and the low-frequency component respectively using a parameter-sharing embedding layer; Extracting a low-frequency signal vector, a high-frequency signal vector and a noise signal vector from the encoding vectors of the high-frequency component and the low-frequency component respectively; Adaptively fusing the low-frequency signal vector and the high-frequency signal vector to obtain a fused signal vector; Using the low-frequency signal vector to enhance the attention of the fused signal vector on variable dependency, thereby obtaining an enhanced variable representation vector; Synchronous reasoning is performed on the enhanced variable representation vector and the noise signal vector respectively, and the reasoning results are integrated to obtain the prediction result of the future time step.

2. The multivariate time series prediction method according to claim 1, wherein: Extracting a low-frequency signal vector, a high-frequency signal vector, and a noise signal vector from the encoding vectors of the high-frequency component and the low-frequency component, respectively, comprising: According to the principle of maximizing the mutual information between the low-frequency signal vector and the high-frequency signal vector and minimizing the mutual information between the low-frequency signal vector and the noise signal vector, the low-frequency signal vector is extracted from the coding vector of the low-frequency component, and the high-frequency signal vector and the noise signal vector are respectively extracted from the coding vector of the high-frequency component.

3. The multivariate time series prediction method according to claim 2, wherein: Methods for maximizing the mutual information between the low-frequency signal vector and the high-frequency signal vector and minimizing the mutual information between the low-frequency signal vector and the noise signal vector include: Reconstructing the low-frequency signal vector using the high-frequency signal vector and calculating the first reconstruction loss; Reconstructing the low-frequency signal vector using the noise signal vector and calculating the second reconstruction loss; combining the first reconstruction loss and the second reconstruction loss to obtain a total reconstruction loss; Minimize the total reconstruction loss.

4. The multivariate time series prediction method according to claim 1, wherein: Adaptively fusing the low-frequency signal vector and the high-frequency signal vector to obtain a fused signal vector, including: Splicing the low-frequency signal vector and the high-frequency signal vector in a feature dimension to obtain a spliced ​​vector; Processing the concatenated vector using a multi-layer perceptron, adjusting it using a temperature coefficient, activating the contribution of the low-frequency signal vector, and calculating the contribution of the high-frequency signal vector; The low-frequency signal vector and the high-frequency signal vector are weightedly fused using corresponding contribution degrees to obtain the fused signal vector.

5. The multivariate time series prediction method according to claim 1, wherein: Utilizing the low-frequency signal vector to enhance the attention of the fused signal vector on variable dependency includes: Calculating a Gram matrix of the fused signal vector using the low-frequency signal vector and adaptively optimizing the vector using trainable parameters to obtain a similarity matrix; The fusion signal vector is weighted by using the similarity matrix and is residually connected with the fusion signal vector, and is post-processed by using a feedforward neural network after linear transformation.

6. The multivariate time series prediction method according to claim 5, wherein: Calculating the Gram matrix of the fused signal vector using the low-frequency signal vector and adaptively optimizing using trainable parameters, including: performing a layer normalization operation on the low-frequency signal vector, and then multiplying the low-frequency signal vector by the layer normalized and transposed low-frequency signal vector to obtain a Gram matrix; After activating the Gram matrix, perform element-wise multiplication with the weight matrix and add the bias vector; An activation operation is performed on the calculation result to obtain the similarity matrix.

7. A multivariate time series prediction device, comprising a processor and a storage medium, wherein the storage medium stores a computer program, characterized in that: When the computer program is executed by a processor, the computer program executes the multivariate time series prediction method according to any one of claims 1 to 6.

8. A method for training a multivariate time series prediction model, wherein the multivariate time series prediction model is used to predict a time series of future time steps based on a current input time series, wherein the time series is obtained by synchronously collecting multiple variable indicators in a time series; characterized in that: Training methods include: Obtain a set of time series samples and divide them into training and test sets; Using the training set to train the multivariate time series prediction model with the goal of minimizing training loss, and using the test set to test the trained multivariate time series prediction model; The multivariate time series forecasting model is configured to: Get the input time series; Decomposing the acquired time series into high-frequency components and low-frequency components based on a preset frequency threshold; Performing vector encoding on the high-frequency component and the low-frequency component respectively using a parameter-sharing embedding layer; Extracting a low-frequency signal vector, a high-frequency signal vector and a noise signal vector from the encoding vectors of the high-frequency component and the low-frequency component respectively; Adaptively fusing the low-frequency signal vector and the high-frequency signal vector to obtain a fused signal vector; Using the low-frequency signal vector to enhance the attention of the fused signal vector on variable dependency, thereby obtaining an enhanced variable representation vector; Performing synchronous reasoning on the enhanced variable representation vector and the noise signal vector respectively, and fusing the reasoning results to obtain a prediction result for the future time step; The training loss is composed of the mean square error (MSE) between the predicted results and the true results of all variables in the time series samples of the training set at each time step, and the weighted total reconstruction loss in the process of extracting the low-frequency signal vector, the high-frequency signal vector and the noise signal vector.

9. A multivariate time series prediction device, characterized in that: The device is configured with a multivariate time series prediction model trained using the multivariate time series prediction model training method as claimed in claim 8.

10. A multivariate time series prediction system, characterized in that: It includes an input device, an output device and the multivariate time series prediction device as described in claim 9; the input device is connected to the input end of the multivariate time series prediction device, for receiving the time series to be predicted and inputting it into the multivariate time series prediction model; the output device is connected to the output end of the multivariate time series prediction device, for outputting the time series predicted by the multivariate time series prediction model.

Citation Information

Patent Citations

  • Daily peak load prediction method, computer equipment and readable storage medium

    CN115169232A

  • Construction method of complex multivariable system network prediction model based on Informer architecture

    CN118690170A

  • Time sequence prediction method based on improved Autoformer model

    CN120197747A

  • Multivariable time series prediction method based on Patching and multi-scale feature extraction

    CN120316427A

  • Methods and systems for sampling and storing machine signals for analytics and maintenance using the industrial internet of things

    US20200133256A1