Time sequence prediction method based on patch-level differential prediction and correction

By introducing patch-level differential prediction and correction methods in time series prediction, the full attention mechanism and the tanh loss-weighted differential loss function are used to solve the problem of fluctuation prediction distortion in long-term prediction, and achieve higher prediction accuracy and generalization performance.

CN120067576APending Publication Date: 2025-05-30TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510125722.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing time series prediction methods have problems with fluctuation prediction distortion and lack of effective utilization of predicted values ​​in long-term predictions, especially when there are distribution drifts and fluctuations in the sequence.

Method used

A time series prediction method based on patch-level differential prediction and correction is proposed. By patching and characterizing the first-order difference sequence of the baseline model prediction sequence and the backtracking sequence, the timing dependency is extracted using the full attention mechanism, and patch alignment and accumulation and operation are performed through the differential prediction sequence and the difference sequence of the baseline model prediction sequence to obtain the final prediction result. At the same time, a tanh loss-weighted differential loss function is introduced to enhance the robustness of the model.

Benefits of technology

This method not only improves the baseline model's prediction ability for fluctuations, but also improves the overall prediction accuracy and shows good generalization performance on multiple baseline models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067576A_ABST
    Figure CN120067576A_ABST
Patent Text Reader

Abstract

The invention discloses a time sequence prediction method based on patch-level difference prediction and correction. The time sequence prediction method mainly comprises the following steps: dividing a difference sequence of a baseline model prediction sequence and a backtracking sequence into non-overlapped patches and representing the patches; modeling a time sequence dependency relationship between patches by using a full attention mechanism; a representation corresponding to the prediction interval is selected and mapped into a patch window prediction value through a linear layer; patch alignment and difference operation are carried out on the predicted difference sequence and the difference sequence of the baseline model prediction sequence, the cumulative sum of difference values in a patch window is calculated, and the cumulative sum and the baseline model prediction sequence are added point by point to obtain a final prediction result; in order to improve the robustness to larger fluctuation, a difference loss function of tanh weighting is additionally introduced; the method serves as an auxiliary plug-in of the baseline model, the differential information in the prediction sequence is introduced, the problem that an existing model is insufficient in sequence fluctuation prediction is solved, and meanwhile the overall prediction precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of time series prediction, and particularly relates to a method based on patch-level differential prediction and correction. Background Art

[0002] Time series prediction refers to using the observed values of a historical time window to predict the unknown values of a future time window. According to the different lengths of the prediction window, it can be further divided into short-term prediction and long-term prediction. Traditional time series prediction methods usually focus on short-term prediction. However, in many practical applications, long-term prediction has a more profound guiding significance for decision-making and has a wider application in many fields.

[0003] Time series data usually has characteristics such as long-term dependence, seasonality, trend, and sudden events. As the time span increases, the uncertainty of the data also increases, and the accuracy of long-term prediction gradually decreases. Therefore, constructing an accurate long-term time series prediction model is very challenging. In recent years, the development of deep learning has promoted the progress of long-term time series prediction research. Although the accuracy of long-term time series prediction has been greatly improved at the present stage, it is found from the prediction curve that there are problems of prediction distortion in the fluctuation of most segments. Due to the potential distribution drift phenomenon in the sequence, two sequence segments with similar fluctuation conditions may fall into different value range spaces. Without explicit prior guidance, it is difficult for the model to discover the correlation between the two segments. The first-order difference can well describe the fluctuation relationship in the sequence by calculating the difference between adjacent moments. For two sequence segments with similar fluctuation changes, their first-order difference curves are also highly similar, and the difference sequence has good stationarity. However, the process of restoring the difference sequence to the original sequence requires cumulative sum operations, and directly applying it to long-term prediction tasks will cause large cumulative errors. Therefore, how to effectively utilize the difference information in the prediction sequence and improve the model's prediction ability for sequence fluctuations is a meaningful but challenging research issue, and there is no relevant research on using the difference sequence for auxiliary prediction in existing time series prediction methods. Summary of the Invention

[0004] Aiming at the problems of distorted fluctuation prediction and lack of display and utilization of predicted values in existing time series prediction methods, the present invention proposes a time series prediction method based on patch-level differential prediction and correction. This method constructs a new end-to-end differential enhancement model on the basis of a baseline model. Specifically, this method introduces the first-order difference sequences of the baseline model prediction sequence and the backtracking sequence, and through the patch division representation method, uses the full attention mechanism to extract the temporal dependence relationships between historical patches, between historical patches and predicted patches, and between predicted patches, effectively enriching the feature representation of the predicted patches. By aligning the patch window differential prediction sequence obtained by linear layer mapping with the differential sequence of the baseline model prediction sequence and taking the point-by-point difference, and adding the cumulative sum of the patch window to the baseline model prediction sequence point by point to obtain the final prediction result. At the same time, a differential loss function weighted by tanh loss is introduced to enhance the robustness of the model to large fluctuations. This method not only improves the prediction ability of the baseline model for fluctuations, but also improves the overall prediction accuracy, and has good generalization performance on multiple baseline models.

[0005] To solve the above technical problems, a time series prediction method based on patch-level differential prediction and correction proposed by the present invention includes the following steps:

[0006] Step 1) Patch division and representation of difference sequences: Input the backtracking sequence into the baseline model to obtain the baseline model prediction sequence, perform first-order differences on the backtracking sequence and the baseline model prediction sequence respectively to obtain two difference sequences, perform non-overlapping patch divisions on the two difference sequences respectively to obtain two difference patch sequences, splice the two difference patch sequences in the time dimension, and obtain the patch representation through linear mapping and embedding position encoding;

[0007] Step 2) Differential prediction based on the full attention mechanism, including: 2-1) Input the patch representation obtained in Step 1) into the Transformer Encoder, and use the multi-head self-attention mechanism to extract the long-term and short-term dependence features between historical patches, between historical patches and predicted patches, and between predicted patches. The output of the Transformer Encoder is the patch representation with the same shape as the input; 2-2) Intercept the patch representation corresponding to the prediction interval in the output of Step 2-1), and input it into the shared patch-level linear layer to map to obtain the differential prediction sequence with the length of the patch window;

[0008] Step 3) Patch-level differential correction: Align the differential prediction sequence obtained in Step 2) and the differential sequence of the baseline model prediction sequence in Step 1) by patch to calculate the difference, then calculate the cumulative sum of the differences within the patch and add it point by point to the baseline model prediction sequence obtained in Step 1) to obtain the differentially corrected prediction sequence;

[0009] Step 4) Calculation of differential loss with tanh weight introduction, including: 4-1) Input the true differential sequence of the prediction window into the tanh function with absolute value to calculate the point-level differential loss weight, calculate the mean squared error loss point by point according to the differential prediction sequence output in Step 2) and the true differential sequence of the prediction window, multiply the differential loss weight and the above mean squared error loss point by point and take the mean to obtain the differential prediction loss; 4-2) Calculate the mean squared error loss according to the differentially corrected prediction sequence obtained in Step 3) and the true sequence of the prediction window to obtain the final prediction loss; 4-3) Sum the differential prediction loss obtained in Step 4-1) and the final prediction loss calculated in Step 4-2) to obtain the overall optimization loss term.

[0010] Further, in the time series prediction method of the present invention, where:

[0011] The specific content of the differential sequence patch division and characterization in Step 1) is as follows:

[0012] For the input retrospective sequence X ∈ R C×L , the baseline model first outputs the baseline model prediction sequence where C represents the number of variables. In the present invention, C is 1, that is, unit variable prediction is performed. L represents the length of the retrospective window, and H represents the length of the prediction window. The retrospective sequence and the baseline model prediction sequence are subjected to first-order difference processing to obtain two difference sequences and The two difference sequences are respectively divided into two difference patch sequences with a length of S in a non-overlapping manner and where: represent the number of patches in the retrospective window and the number of patches in the prediction window respectively. Then, the two difference patch sequences are concatenated in the time series dimension to obtain the difference patch sequence D p :

[0013]

[0014] In Equation (1), N = N x + N y ;

[0015] Subsequently, the difference patch sequence D pInput a patch - granularity linear layer to perform a characterization mapping on each patch and encode its embedding position to obtain a patch representation:

[0016]

[0017] In Equation (2), d is the hidden - layer dimension.

[0018] Step 2) The differential prediction based on the full - attention mechanism is as follows:

[0019] In Step 2 - 1), in each of the multiple layers of the Transformer Encoder, each layer includes a multi - head self - attention layer, a normalization layer, and a feed - forward network layer; in the multiple layers, the input of the first layer is the patch representation obtained from Equation (2), and the calculation process of the patch representation of the l - th layer is as follows:

[0020] Step 2 - 1 - 1) l = 2;

[0021] Step 2 - 1 - 2) For the input tensor of the l - th layer The multi - head self - attention layer first maps Z l to tensors Q l , K l , V l :

[0022]

[0023] In Equation (3), F q F k F v are all fully - connected layers in the self - attention layer;

[0024] The tensors Q l , K l , V l perform multi - head attention calculation on the tensor matrices, and use to represent the tensor slices corresponding to each attention head in the tensor matrices of Q l , K l , V l respectively, where H and d h represent the number of attention heads and the dimension of each attention head respectively; the attention score is calculated for each variable channel in a variable - channel - independent manner:

[0025]

[0026] In Equation (4), is the attention score matrix, and the weighted attention output is calculated as:

[0027]

[0028] The final output of the multi-head self-attention layer of this layer is: Subsequently, the input and output of the multi-head self-attention layer are subjected to residual connection and layer normalization to obtain:

[0029]

[0030] Subsequently is input into the feed-forward network layer, and the output of the feed-forward network layer will be subjected to residual connection and layer normalization with the input to obtain the output of the l-th layer:

[0031]

[0032] The output of the l-th layer is used as the input of the (l + 1)-th layer; reassign l + 1 to l;

[0033] Step 2-1-3) Return to the above Step 2-1-2) until the patch representation calculation of all layers in the Transformer Encoder is completed, and the final output of the Transformer Encoder is Thus, the patch representation Z of the input of the first layer passes through the Transformer Encoder to obtain a patch representation Z with the same shape as the patch representation Z enc .

[0034] Step 2-2) First, slice the patch representation Z output in Step 2-1 enc in the time series dimension into and Subsequently is input into another patch-granularity linear layer for prediction:

[0035]

[0036] In Equation (8), is the differential prediction patch sequence of the patch window length, and this differential prediction patch sequence is the differential patch sequence of the prediction sequence of the baseline model in Step 1 's calibration prediction result.

[0037] Step 3) The specific content of patch-level differential correction is as follows:

[0038] First, for the differential prediction sequence of the patch window length obtained in Step 2 Differential patch sequence with the baseline model prediction sequence Perform patch alignment and take the difference:

[0039]

[0040] In Equation (9), δ p is the differential prediction patch sequence and the differential patch sequence of the baseline model prediction sequence of the difference patch sequence. Perform a cumulative sum operation within each patch of this difference patch sequence, and then flatten the difference patch sequence δ p into a long sequence:

[0041]

[0042] Finally, add the long sequence δ point - by - point to the baseline model prediction sequence to obtain the differentially corrected prediction sequence:

[0043]

[0044] The specific content of step 4) for calculating the differential loss with tanh weights is as follows:

[0045] In step 4 - 1), perform a first - order difference operation on the true sequence of the prediction window to obtain the true differential sequence of the prediction window Perform a calculation on the true differential sequence D of the prediction window y through the tanh function with absolute value to obtain the point - level differential loss weight W:

[0046]

[0047] Subsequently, according to the differential prediction sequence output in step 2) and the true differential sequence D of the prediction window y calculate the squared error loss point - by - point, multiply the differential loss weight W and the above - mentioned squared error loss point - by - point and take the average to obtain the differential prediction loss

[0048]

[0049] In Equation (13),

[0050] In step 4 - 2), for the differentially corrected prediction sequence in step 3) use the mean squared error loss to measure the difference between the prediction sequence and the true sequence:

[0051]

[0052] In formula (14), Take as the final prediction loss;

[0053] 4-3) Take the differential prediction loss obtained in step 4-1) and the final prediction loss calculated in step 4-2) Sum them to obtain the overall optimization loss term:

[0054]

[0055] In formula (15), ρ is a constant used to balance the difference in the order of magnitude of the two loss terms, and ρ = 10.

[0056] Compared with the prior art, the beneficial effects of the present invention are:

[0057] The present invention proposes a time series prediction method based on patch-level differential prediction and correction. By performing patch partitioning and characterization on the first-order difference sequences of the baseline model prediction sequence and the backtracking sequence, and using the full attention mechanism to extract the temporal dependence relationship between patches, the feature representation of the prediction patches is effectively enriched. By performing patch alignment and subtraction on the differential prediction sequence obtained by linear layer mapping and the difference sequence of the baseline model prediction sequence, and adding the cumulative sum of the patch window to the baseline model prediction sequence point by point to obtain the final prediction result. At the same time, the model introduces a differential loss function weighted by the tanh loss to enhance the robustness of the model to large fluctuations. This method not only improves the prediction ability of the baseline model for fluctuations, but also improves the overall prediction accuracy, and has good generalization performance on multiple baseline models. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 is a flowchart of the time series prediction method described in the present invention;

[0059] Figure 2 is a schematic structural diagram of the model in the embodiment of the present invention; DETAILED DESCRIPTION OF THE INVENTION

[0060] The design idea of the time series prediction method based on patch-level differential prediction and correction proposed by the present invention is as follows: By performing patch partitioning and characterization on the first-order difference sequences of the baseline model prediction sequence and the retrospective sequence, a full attention mechanism is used to extract the temporal dependencies between historical patches, between historical patches and prediction patches, and between prediction patches, effectively enriching the feature representation of the prediction patches. By performing patch alignment and subtraction on the differential prediction patch sequence obtained by linear layer mapping and the difference sequence of the baseline model prediction sequence, and adding the cumulative sum of the patch window back to the baseline model prediction sequence, the final prediction result is obtained. At the same time, the model introduces a differential loss function weighted by tanh loss to enhance the robustness of the model to large fluctuations. This method not only improves the prediction ability of the baseline model for fluctuations, but also improves the overall prediction accuracy, and has good generalization performance on multiple baseline models.

[0061] The following further describes the present invention with reference to the accompanying drawings and specific embodiments, but the following embodiments are by no means any limitation to the present invention.

[0062] Definition of the problem in the present invention: Input a retrospective sequence X = [x t-L+1 , x t-L+2 , …, x t , its first-order difference sequence is defined as D x = [d t-L+1 , d t-L+2 , …, d t , and the differential prediction sequence of the differential enhancement model is The final prediction sequence of the model is The overall optimization goal is to reduce the difference between the prediction sequence and and the true sequence Y and D y .

[0063] The time series prediction method of the present invention is implemented through a differential enhancement model. Figure 2As shown in the figure, the differential enhancement model includes a differential sequence patch division and characterization module, a differential prediction module based on the full attention mechanism, a patch-level differential correction module, and a differential loss module that are connected in sequence. The differential loss calculation with tanh weights is introduced to optimize the overall model. Among them, the differential sequence patch division and characterization module includes a baseline model and a patch-granularity linear layer. The baseline model includes, but is not limited to, (PatchTST, TimeSNet, SegRNN, Dlinear, etc.); the differential prediction module based on the full attention mechanism includes a Transformer Encoder (encoder) and a patch-granularity linear layer. In each layer of the multiple layers of the Transformer Encoder, each layer includes a multi-head self-attention layer, a normalization layer, and a feed-forward network layer.

[0064] As Figure 1 and Figure 2 shown, a time series prediction method based on patch-level differential prediction and correction proposed by the present invention is as follows:

[0065] Step 1: Differential sequence patch division and characterization:

[0066] For the input backtracking sequence X ∈ R C×L , the baseline model will first output the baseline model prediction sequence where C represents the number of variables. In the present invention, C is 1, that is, unit variable prediction is performed. L represents the backtracking window length, and H represents the prediction window length. The backtracking sequence and the baseline model prediction sequence are subjected to first-order difference processing to obtain two difference sequences and The two difference sequences are respectively divided into two difference patch sequences with a length of S in a non-overlapping manner and where: respectively represent the number of backtracking window patches and the number of prediction window patches. Then, the two difference patch sequences are concatenated in the time series dimension to obtain the difference patch sequence D p :

[0067]

[0068] In formula (1), N = N x + N y ; Subsequently, the difference patch sequence D p is input into a patch-granularity linear layer to perform characterization mapping on each patch and embed position encoding for it to obtain the patch characterization:

[0069]

[0070] In formula (2), d is the dimension of the hidden layer.

[0071] Step 2, differential prediction based on the full attention mechanism, including:

[0072] Step 2-1) Input the patch representation obtained in Step 1 into the Transformer Encoder, and use the multi-head self-attention mechanism to extract the long-term and short-term dependence features among historical patches, between historical patches and the predicted patch, and among predicted patches. Since this method extracts three kinds of temporal dependence relationships simultaneously through the self-attention mechanism, it is defined as the full attention mechanism. Specifically, in the present invention, the differential enhancement model will adopt stacked Encoder Layers in the temporal dimension to learn the feature dependence relationship among patches. In each layer of the multi-layer Transformer Encoder, each layer includes a multi-head self-attention layer, a normalization layer, and a feed-forward network layer; the Transformer Encoder will output a patch representation with the same shape as the input. The specific content is as follows:

[0073] In the multi-layer, the input of the first layer is the patch representation obtained from formula (2), and the calculation process of the patch representation of the l-th layer is as follows:

[0074] 2-1-1) l = 2;

[0075] 2-1-2) For the input tensor of the l-th layer The multi-head self-attention layer first maps Z l to the tensors Q l , K l , V l :

[0076]

[0077] In formula (3), F q F k F v are all fully connected layers in the self-attention layer;

[0078] Subsequently, the Q l , K l , V l tensor matrices will perform multi-head attention calculations, and use to represent Q l , K l , V lThe tensor slices corresponding to each attention head in the tensor matrix, where H and d h represent the number of attention heads and the dimension of each attention head respectively; in the present invention, an attention calculation is performed in a variable-channel independent manner. For each variable channel, the model will calculate the attention score:

[0079]

[0080] In formula (4), is the attention score matrix, and the weighted attention output is calculated as:

[0081]

[0082] The final output of the multi-head self-attention layer of this layer layer is: Subsequently, the input and output of the multi-head self-attention layer are subjected to residual connection and layer normalization to obtain:

[0083]

[0084] Subsequently, is input into the feed-forward network layer, and the output of the feed-forward network layer will be subjected to residual connection and layer normalization with the input to obtain the output of the l-th layer layer:

[0085]

[0086] The output of the l-th layer layer is used as the input of the (l + 1)-th layer layer; reassign l + 1 to l;

[0087] 2-1-3) Return to the above step 2-1-2) until the patch representation calculation of all layers layer in the Transformer Encoder is completed. The final output of the Transformer Encoder is Thus, the patch representation Z of the input of the first layer layer is passed through the Transformer Encoder to obtain a patch representation Z with the same shape as the patch representation Z enc .

[0088] Step 2-2) First, slice the patch representation Z output in step 2-1 enc in the time series dimension into and Subsequently, is input into another patch-granularity linear layer for prediction:

[0089]

[0090] In formula (8), is the differential prediction patch sequence of the patch window length, and this differential prediction patch sequence is the differential patch sequence of the prediction sequence of the baseline model in step 1) Calibration prediction result

[0091] Step 3, Patch-level differential correction:

[0092] Since the process of restoring the differential sequence to the original sequence requires cumulative sum operations, a large cumulative error will be caused when the prediction window is long. To reduce the influence of the cumulative error, this method performs cumulative sum restoration operations within each patch

[0093] First, for the differential prediction sequence of the patch window length obtained in step 2 and the differential patch sequence of the prediction sequence of the baseline model in step 1 Perform patch alignment and subtraction:

[0094]

[0095] In formula (9), δ p is the difference patch sequence between the differential prediction patch sequence and the differential patch sequence of the prediction sequence of the baseline model Perform cumulative sum operations on this difference patch sequence within each patch, and then flatten this difference patch sequence δ p into a long sequence:

[0096]

[0097] Then, add the long sequence δ to the prediction sequence of the baseline model point by point to obtain the differentially corrected prediction sequence:

[0098]

[0099] Step 4, Optimize the model after the above patch-level differential correction by introducing a differential loss module with tanh weights

[0100] Considering that larger fluctuations in the sequence have a greater impact on the sequence trend, based on the numerical characteristics of the first-order difference sequence, the greater the absolute value of the first-order difference value at a certain moment, the greater the fluctuation at that moment. The tanh function has an axisymmetric characteristic, and its value range is between (-1, 1). When the absolute value of the difference value is larger, the absolute value of the tanh function value of this difference value is also larger. In this study, the tanh function with absolute value is used to calculate the weight of the difference loss, enhancing the robustness of the model to larger fluctuations. Specifically, the calculation of the difference loss introducing the tanh weight includes:

[0101] First, perform a first-order difference operation on the true sequence of the prediction window to obtain the true difference sequence of the prediction window Perform the following operation on the true difference sequence D of the prediction window y to calculate the point-level difference loss weight W through the tanh function with absolute value:

[0102]

[0103] Subsequently, according to the difference prediction sequence output in step 2) and the true difference sequence D of the prediction window y calculate the squared error loss point by point, multiply the difference loss weight W and the above squared error loss point by point, and take the average to obtain the difference prediction loss

[0104]

[0105] In formula (13),

[0106] For the prediction sequence after the difference correction in step 3 use the mean squared error loss to measure the difference between the prediction sequence and the true sequence:

[0107]

[0108] In formula (14), Take as the final prediction loss;

[0109] Finally, sum the difference prediction loss obtained from the above formula (13) and the final prediction loss calculated by formula (14)

[0110]

[0111] In formula (15), ρ is a constant used to balance the difference in the order of magnitude of the two loss terms, and ρ = 10.

[0112] Research materials:

[0113] To verify the effectiveness of the time series prediction method based on patch - level differential prediction and correction proposed by the present invention, unit variable prediction experiments are conducted on four publicly available transformer datasets, namely ETTh1, ETTh2, ETTm1, and ETTm2, for the prediction window length set (96, 192, 336, 720). The present invention constructs a new end - to - end trainable differential enhancement model based on the baseline model. As can be seen from the above specific description, steps 1 - 3 are, in sequence, the differential sequence patch division and characterization module, the differential prediction module based on the full attention mechanism, and the patch - level differential correction module. In step 4, a weighted differential prediction loss is calculated for the differential prediction sequence obtained by the differential prediction module by introducing the differential loss calculation with tanh weights, and at the same time, the final prediction loss is calculated for the final prediction sequence obtained by the differential correction module. The sum of the two losses is the overall optimization loss term of the model, realizing the optimization of the model.

[0114] As Figure 2 shown, simply input the baseline model prediction sequence and the input sequence into the overall model according to the process described in step 1, and a new end - to - end trainable model can be combined. Steps 1 - 4 show the specific training process of the overall module.

[0115] The present invention can be adapted to multiple baseline models for time series prediction, including but not limited to (PatchTST, TimeSNet, SegRNN, Dlinear, etc.).

[0116] In the present research invention, PatchTST is used as the baseline model to verify the effectiveness of the differential enhancement model. Table 1 (including the continued Table 1) shows the comparison of the prediction mean squared error (MSE) and mean absolute error (MAE) results between the method proposed by the present invention and six baseline model methods. The data in bold in the table represent the optimal prediction effects.

[0117] Table 1

[0118]

[0119] Continued Table 1

[0120]

[0121] From the experimental results in Table 1 (including the continued Table 1), it can be seen that the present invention has achieved better prediction effects compared to PatchTST on the four datasets. Compared with other baseline methods, the present invention shows the best prediction effects on three of the datasets, namely ETTh1, ETTh2, and ETTm1, proving the effectiveness of the time series prediction method based on patch - level differential prediction and correction of the present invention.

[0122] Although the present invention has been described above in conjunction with the accompanying drawings, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many improvements and changes without departing from the spirit of the present invention, and all of these fall within the scope of protection of the present invention.

Claims

1. A time series prediction method based on patch-level differential prediction and correction, characterized in that: The method comprises the following steps: Step 1) Differential sequence patch division and characterization: Input the backtracking sequence into the baseline model to obtain the baseline model prediction sequence, perform first-order difference on the backtracking sequence and the baseline model prediction sequence to obtain two difference sequences, perform non-overlapping patch division on the two difference sequences to obtain two difference patch sequences, concatenate the two difference patch sequences in the time series dimension, perform linear mapping and embed position coding to obtain patch representation; Step 2) Differential prediction based on the full attention mechanism, including: 2-1) Input the patch representation obtained in step 1) into the Transformer Encoder, and use the multi-head self-attention mechanism to extract the long-term and short-term dependency features between historical patches, between historical patches and predicted patches, and between predicted patches. The Transformer Encoder outputs a patch representation with the same shape as the input; 2-2) intercepting the patch representation corresponding to the prediction interval in the output of step 2-1), inputting the shared patch granularity linear layer mapping to obtain the differential prediction sequence of the patch window length; Step 3) Patch-level differential correction: Align the difference prediction sequence obtained in step 2) and the difference sequence of the baseline model prediction sequence in step 1) by patch and calculate the difference, then calculate the cumulative sum of the differences within the patch and add them point by point to the baseline model prediction sequence obtained in step 1) to obtain the prediction sequence after difference correction; Step 4) Introduce the differential loss calculation of tanh weights, including: 4-1) Input the real difference sequence of the prediction window into the tanh function with an absolute value to calculate the point-level difference loss weight, calculate the square error loss point by point based on the difference prediction sequence output from step 2) and the real difference sequence of the prediction window, multiply the difference loss weight and the above square error loss point by point and calculate the average to obtain the difference prediction loss; 4-2) Calculate the mean square error loss based on the difference-corrected prediction sequence obtained in step 3) and the actual sequence of the prediction window to obtain the final prediction loss; 4-3) The differential prediction loss obtained in step 4-1) and the final prediction loss calculated in step 4-2) are summed to obtain the overall optimization loss term.

2. The time series prediction method according to claim 1, characterized in that: The specific contents of the differential sequence patch division and characterization described in step 1) are as follows: For the input backtracking sequence X∈R C×L , the baseline model first outputs the baseline model prediction sequence Where C represents the number of variables, C is 1, L represents the length of the lookback window, and H represents the length of the prediction window. Perform first-order difference processing to obtain two difference sequences and Divide the two differential sequences into two differential patch sequences of length S in a non-overlapping manner and in: They represent the number of patches in the lookback window and the number of patches in the prediction window respectively. Then, the two differential patch sequences are concatenated in the time dimension to obtain the differential patch sequence D p : In formula (1), N = N x +N y ; Then the differential patch sequence D p Input a patch-sized linear layer to represent and map each patch and encode its embedding position to obtain the patch representation: In formula (2), d is the hidden layer dimension.

3. The time series prediction method according to claim 2, characterized in that: In step 2-1), in the multi-layer layer of the Transformer Encoder, each layer includes a multi-head self-attention layer, a normalization layer and a feed-forward network layer; in the multi-layer layer, the input of the first layer is the patch representation obtained by formula (2), and the patch representation calculation process of the lth layer is as follows: Step 2-1-1) l = 2; Step 2-1-2) For the input tensor of layer l The multi-head self-attention layer first transforms Z l Mapped to tensor Q l ,K l ,V l : In formula (3), F q ,F k ,F v They are all fully connected layers in the self-attention layer; The Q l ,K l ,V l Tensor matrix for multi-head attention calculation, use Respectively represent Q l ,K l ,V l The tensor slice corresponding to each attention head in the tensor matrix, where H and d h Represent the number of attention heads and the dimension of each attention head respectively; each variable channel is independent of the variable channel. Calculate the attention score: In formula (4), is the attention score matrix, and the weighted attention output is calculated as: The final output of the multi-head self-attention layer of this layer is: Then the input and output of the multi-head self-attention layer are residually connected and layer normalized to obtain: Then will Input the feedforward network layer, the output of the feedforward network layer will be residually connected with the input and layer normalized to obtain the output of the lth layer: The output of the lth layer is used as the input of the l+1th layer; l+1 is reassigned to l; Step 2-1-3) Return to step 2-1-2) above until the patch representation calculation of all layers in the Transformer Encoder is completed. The final output of the Transformer Encoder is Thus, the input patch representation Z of the first layer is passed through the Transformer Encoder to obtain a patch representation Z with the same shape as the patch representation Z. enc .

4. The time series prediction method according to claim 3, characterized in that: The specific contents of step 2-2) are as follows: First, the patch representation Z output from step 2-1) enc Slicing on the time series dimension is and Then will Enter another patch-sized linear layer for prediction: In formula (8), The differential prediction patch sequence is the difference patch sequence of the patch window length, which is the difference patch sequence of the baseline model prediction sequence in step 1). The calibration prediction results.

5. The time series prediction method according to claim 1, characterized in that: The specific contents of step 3) are as follows: First, the differential prediction sequence of the patch window length obtained in step 2) The difference patch sequence between the baseline model prediction sequence Perform patch alignment and subtraction: In formula (9), δ p is the differential prediction patch sequence The difference patch sequence between the baseline model prediction sequence and The difference patch sequence is accumulated and operated inside each patch, and then the difference patch sequence δ p Flatten to a long sequence: Finally, the long sequence δ is compared with the baseline model prediction sequence Add point by point to get the prediction sequence after difference correction:

6. The time series prediction method according to claim 1, characterized in that: The specific contents of step 4) are as follows: In step 4-1), the real sequence of the prediction window Perform first-order difference processing to obtain the true difference sequence of the prediction window The true difference sequence D of the prediction window y The point-level differential loss weight W is calculated by the tanh function with an absolute value: Then, according to the differential prediction sequence output in step 2) and the forecast window true difference sequence D y Calculate the square error loss point by point, multiply the differential loss weight W and the above square error loss point by point and calculate the average to get the differential prediction loss In formula (13), In step 4-2), for the prediction sequence after differential correction in step 3) The mean square error loss is used to measure the difference between the predicted sequence and the true sequence: In formula (14), Will As the final predicted loss; 4-3) The differential prediction loss obtained in step 4-1) The final prediction loss calculated in step 4-2) The sum is the overall optimization loss term: In formula (15), ρ is a constant used to balance the difference in the magnitude of the two loss terms, ρ = 10.