Time series prediction method based on extreme loss and hybrid expert structure

By introducing extreme value loss function and mixed expert structure in time series prediction, the problems of insufficient extreme value prediction and limited model expression ability in the prior art are solved, and higher prediction accuracy and flexibility are achieved.

CN120067968AInactive Publication Date: 2025-05-30TIANJIN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510021340.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing time series prediction methods have problems such as insufficient extreme value prediction and limited model expression capabilities when dealing with complex time patterns.

Method used

Using a method combining extreme loss and mixed expert structure, the model predicts the extreme loss function to improve the model's prediction accuracy for extreme values, and dynamically process complex and diverse time series data through the mixed expert structure.

Benefits of technology

It significantly improves the flexibility and expression ability of the model, improves the prediction accuracy of extreme cases and the generalization ability of the model, and is suitable for application scenarios that require high-precision prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067968A_ABST
    Figure CN120067968A_ABST
Patent Text Reader

Abstract

The invention discloses a time series prediction method based on extreme value loss and a mixed expert structure, and the method mainly comprises the steps: carrying out the preprocessing of time series data, and carrying out the calculation through a moving average value, and obtaining trend term data and seasonal term data; predicting the trend term data through a normalization layer, a multi-layer perceptron and a linear projection layer by using an RLinear model to obtain a trend term prediction result; selecting experts by using a gating mechanism, and predicting seasonal items based on a mixed expert structure; and combining prediction results of the trend item and the season item to generate a final time sequence prediction result. In order to improve the accuracy of extreme value prediction, an extreme value loss function is introduced, and the extreme value loss is calculated by setting an upper and lower threshold division sequence. Finally, the loss function combines a time sequence prediction error and an extreme value error, and the model dynamically adjusts the attention to the extreme value and overall prediction. The method effectively improves the trend and extreme value prediction precision of the time sequence, and is suitable for a prediction task of complex time sequence data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of time series prediction, and particularly relates to a method combining extreme value loss and a mixture of experts structure. Background Art

[0002] Time series prediction aims to predict the future data change trend by analyzing the internal laws of historical data sequences. These data usually contain seasonal, long-term trend and non-stationary characteristics. Time series prediction technology has been widely applied in fields such as economy, finance, meteorology, etc. Accurate prediction is crucial for decision support systems. For example, in power supply management, accurate load prediction can help optimize resource allocation and energy procurement; in stock market analysis, accurate prediction of stock price trends can help investors make more informed investment decisions.

[0003] Traditional time series prediction methods, such as ARIMA models, exponential smoothing methods, and machine learning methods, although achieving certain effects in practice, show obvious limitations when dealing with complex time patterns. In recent years, with the rapid development of deep learning technology, more and more artificial intelligence algorithms have been applied to the field of time series prediction, including advanced models such as long short-term memory networks (LSTM), Transformer, etc.

[0004] However, due to the structural limitations of a single model, its model capacity is limited. When facing data with high diversity or strong complexity, a single model is difficult to flexibly capture all features. The mixture of experts structure (MoE) can dynamically select the most suitable expert by introducing multiple experts, each of which focuses on processing different parts of the data, thus enhancing the model's expressive ability and flexibility. At the same time, most current models use mean squared error (MSE) as the loss function. MSE is highly sensitive to outliers with large residuals, resulting in the model being more inclined to optimize most of the intermediate values and ignoring extreme values, thus having problems of insufficient prediction in these extreme cases. Summary of the Invention

[0005] Aiming at the problems of insufficient extreme value prediction and limited model expressive ability in existing time series prediction methods, the present invention proposes a time series prediction method based on extreme value loss and a mixture of experts structure. This method effectively alleviates the sensitivity problem of mean squared error (MSE) when dealing with extreme values by introducing an extreme value loss function, and improves the prediction accuracy of the model for extreme values. At the same time, by adopting a mixture of experts structure, multiple experts are used to dynamically process complex and diverse time series data, significantly enhancing the flexibility and expressive ability of the model. This method can not only capture complex time series patterns, but also improve the generalization ability and prediction accuracy of the model, and is particularly suitable for application scenarios that require high-precision prediction, such as financial market trend prediction and energy load management and other fields.

[0006] To solve the above technical problems, a time series prediction method based on extreme value loss and a mixture of experts structure proposed by the present invention includes the following steps:

[0007] Step 1) Data preparation: Preprocess the obtained time series data and divide it into a training set, a validation set, and a test set according to a predetermined ratio; and process missing values and outliers in the obtained historical data;

[0008] Step 2) Sequence decomposition: Decompose the time series data preprocessed in Step 1), and calculate the trend term data using a moving average; subtract the trend term data from the preprocessed time series data to obtain the seasonal term data;

[0009] Step 3) Predict the trend term: Use the RLinear model to predict the trend term data through a normalization layer, a multi-layer perceptron, and a linear projection layer to obtain the trend term prediction result;

[0010] Step 4) Predict the seasonal term: Include:

[0011] 4-1) Decompose the seasonal term data into different time patterns using a multi-scale decomposition method, mix the different time patterns to obtain enhanced seasonal term data;

[0012] 4-2) Input the enhanced seasonal term data into the mixture of experts model;

[0013] 4-3) Each expert in the mixture of experts model will output its own prediction result; at the same time, input the enhanced seasonal term data into the Top-K gating mechanism to obtain the weight assigned to each expert, multiply the result output by each expert by the weight to obtain the prediction result of the mixture of experts model;

[0014] 4-4) Calculate the extreme value loss based on the prediction result of the mixture of experts model obtained in Step 4-3) and the enhanced seasonal term data obtained in Step 4-1), and automatically update the network parameters inside the mixture of experts model through the extreme value loss; after multiple updates, until the number of iterations meets the set requirements, obtain the seasonal term prediction result;

[0015] Step 5) Add the trend term prediction result obtained in Step 3) and the seasonal term prediction result obtained in Step 4) to obtain the final time series prediction result.

[0016] Further, in the time series prediction method of the present invention, where:

[0017] The specific content of Step 2) is as follows:

[0018] Step 2-1) Decompose the time series data preprocessed in Step 1), and calculate the trend term data using the moving average:

[0019]

[0020] In formula (1), x trend (t) represents the t-th trend term, K is the window size of the moving average, and x(t-i) represents the (t-i)-th time series data;

[0021] Step 2-2) Subtract the trend term data from the preprocessed time series data to obtain the seasonal term data, which represents the periodic fluctuations, and its formula is:

[0022] x seasonal (t) = x(t) - x trend (t) (2)

[0023] In formula (2), x seasonal (t) represents the t-th seasonal term, and x(t) represents the t-th time series data.

[0024] In the said Step 3), the RLinear model includes a normalization layer, a multi-layer perceptron, and a linear projection layer. The specific content of Step 3) is as follows:

[0025] Step 3-1) Perform normalization processing on the trend term data through the normalization layer;

[0026] Step 3-2) Input the normalized trend term data into the multi-layer perceptron to extract the features of the time series;

[0027] Step 3-3) Use the linear projection layer to project the features extracted by the multi-layer perceptron;

[0028] Step 3-4) Denormalize the projection result of Step 3-3) to obtain the trend term prediction result:

[0029]

[0030] In formula (3), X trend represents the trend term, represents the trend term prediction result, RevIN norm represents the normalization process, RevIN denorm represents the denormalization process, and W represents the weight matrix of the linear projection layer.

[0031] The specific content of Step 4-1) is as follows:

[0032] Step 4-1-1) Downsample the seasonal term data to extract the first-layer time pattern:

[0033] τ 1 = AvgPooling(x seasonal )(4)

[0034] In Equation (4), τ 1 represents the first-layer time pattern, and AvgPooling represents the average pooling operation;

[0035] Step 4-1-2) For each subsequent layer of time pattern, it is extracted by further performing an average pooling operation on the previous layer of time pattern:

[0036] τ i = AvgPooling(τ i-1 ), i = 2, 3, …, h (5)

[0037] In Equation (5), h represents the total number of layers of the downsampling operation;

[0038] Step 4-1-3) The time patterns are fused layer by layer into u i through a multi-layer perceptron:

[0039] u i = τ i + MLP(τ i+1 ), i = 1, 2, …, h - 1 (6)

[0040] In Equation (6), u i represents the i-th time series information after mixing;

[0041] Finally, the enhanced seasonal term data u 1 is obtained, which is briefly denoted as u.

[0042] The specific content of Step 4-3) is as follows:

[0043] 4-3-1) Input the enhanced seasonal term data obtained in Step 4-1) into the Top-K gating mechanism to obtain the scores of each expert. After normalizing the scores through the Softmax function, select the largest k scores and perform normalization again:

[0044] S = Softmax(TopK(Softmax(u), k)) (7)

[0045] In Equation (7), k is the number of selected experts, and S is the assigned weight of each expert;

[0046] 4-3-2) Each expert in the mixture of experts model receives the enhanced seasonal term data obtained in Step 4-1) as input and inputs it into a deep neural network with an activation layer and an output layer. The output result obtained by each expert is:

[0047] Ei (u) = Linear outp ut(GELU(Linear hidden (u))) (8)

[0048] In Equation (8), and r is the dimension of the low-rank matrix;

[0049] The output results of the k selected expert models in 4-3-3) are multiplied by weights to obtain the prediction result of the mixture-of-experts model:

[0050]

[0051] The specific content of step 4-4) is as follows:

[0052] Step 4-4-1) Calculate the predicted value of the mixture-of-experts model and the true value Y t = {y t+1 , y t+2 , …, y t+H} to obtain the prediction loss

[0053]

[0054] In Equation (10), H represents the predicted time step, is the mean squared error loss;

[0055] 4-4-2) Segment according to the upper and lower quartiles of the data. The data points greater than the upper quartile, less than the upper quartile and greater than the lower quartile, and less than the lower quartile are called the left extreme value, the middle value, and the right extreme value respectively; the left extreme value is assigned 1, the middle value is assigned 0, and the right extreme value is assigned -1; the results after assignment of the true value and the predicted value are denoted as e t and That is:

[0056]

[0057] In Equations (11) and (12), θ high and θ low are the upper quartile and the lower quartile respectively;

[0058] 4-4-3) Calculate the extreme value loss The extreme value loss is the loss between the predicted extreme value sequence and the true extreme value sequence:

[0059]

[0060] 4-4-4) Calculate the difference between the predicted extreme value sequence and the true extreme value sequence, and record the proportion of non-zero terms in this difference as p, that is:

[0061]

[0062] In formula (14), 1 is an indicator function used to calculate the number of inconsistent terms between the predicted extreme value sequence and the true extreme value sequence;

[0063] 4-4-5) Combine the prediction loss and the extreme value loss to calculate the final loss function according to the proportion of non-zero terms p

[0064]

[0065] 4-4-6) According to the result of the extreme value loss, the mixture of experts model automatically updates the internal network parameters. When the number of updates meets the set requirements, the result output by the mixture of experts model is the seasonal term prediction result.

[0066] Compared with the prior art, the beneficial effects of the present invention are:

[0067] Traditional time series prediction methods often use the mean squared error (MSE) as the loss function. This method is often not sensitive enough when dealing with extreme values in the data, resulting in insufficient accuracy in predicting maximum or minimum values. The present invention effectively alleviates this problem by introducing an extreme value loss function, enabling the model to be more sensitive to react and predict extreme values. This improvement not only improves the prediction accuracy for extreme situations but also enhances the adaptability of the model in the face of emergencies or abnormal fluctuations.

[0068] Existing prediction models often rely on a single prediction framework, which is often difficult to capture the diversity and dynamic changes of data when dealing with time series data with high nonlinearity and complex seasonality. The present invention adopts a mixture of experts structure, integrating multiple expert models, and each expert is responsible for capturing different patterns and features in the data. This structure not only significantly improves the flexibility and expressive ability of the model but also can more accurately adapt to and predict the diverse characteristics of time series data by dynamically adjusting the weights of each expert. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 is a flowchart of the time series prediction method described in the present invention;

[0070] Figure 2 is a schematic structural diagram of the model in the embodiment of the present invention;

[0071] Figure 3It is a schematic diagram of the calculation process of extreme value loss provided in the embodiments of the present invention. Detailed implementation manners

[0072] The design idea of the time series prediction method based on extreme value loss and mixture of experts structure proposed by the present invention is as follows: the trend term is predicted by a linear model, and the seasonal term is predicted by a mixture of experts structure. Seasonal data is decomposed into time patterns through downsampling and mixed to enhance the data. A gating mechanism is used to select experts and prediction is performed based on the mixture of experts structure. Combining the prediction results of the trend term and the seasonal term, the final time series prediction value is generated. To improve the accuracy of extreme value prediction, an extreme value loss function is introduced. By setting upper and lower thresholds to divide the predicted value and the true value, the extreme value loss is calculated. The final loss function combines the time series prediction error and the extreme value error to dynamically adjust the model's attention to extreme values and overall prediction. This method effectively improves the trend and extreme value prediction accuracy of time series and is applicable to the prediction task of complex time series data.

[0073] The following further describes the present invention with reference to the accompanying drawings and specific embodiments, but the following embodiments are by no means limiting to the present invention.

[0074] Definition of the problem in the present invention: Let the form of the time series be X = {x 1 , x 2 , …, x T}, assuming that given the historical observation data X = {x 1 , x 2 , …, x L} of L time steps, the prediction task objective is to predict the data Y L+1:L+H = {x L+1 , x L+2 , …, x L+H} of the next H time steps.

[0075] Refer to Figure 1 , Figure 2 and Figure 3 , taking the weather data of a certain area as an example to illustrate the time series prediction method proposed by the present invention. The method includes the following steps:

[0076] Step 1) Data preparation: First, obtain historical data and use it for preprocessing operations before model training. This preprocessing process includes constructing a training set, a validation set, and a test set to ensure the reasonable distribution of data. In addition, it is also necessary to process missing values and outliers in the data to ensure the integrity and accuracy of the input data. Missing values are filled by interpolation method, and outliers are removed to improve the robustness of the model.

[0077] Step 2) Sequence decomposition: Decompose the time series data preprocessed in Step 1) into a trend term and a seasonal term. Calculate the trend term data using a moving average; subtract the trend term data from the preprocessed time series data to obtain the seasonal term data; the trend component x trend (t) is calculated by taking the moving average of the sequence x(t) and represents the long-term change trend of the sequence; the seasonal component x seasonal (t) is obtained through residual calculation and expresses the periodic fluctuations in the sequence.

[0078] Specifically, decompose the time series data preprocessed in Step 1) and calculate the trend term data using a moving average:

[0079]

[0080] In Equation (1), x trend (t) represents the t-th trend term, K is the window size of the moving average, and x(t - i) represents the (t - i)-th time series data;

[0081] Subtract the trend term data from the preprocessed time series data in Step 1) to obtain the seasonal term data to represent periodic fluctuations, and its formula is:

[0082] x seasonal (t) = x(t) - x trend (t) (2)

[0083] In Equation (2), x seasonal (t) represents the t-th seasonal term, and x(t) represents the t-th time series data.

[0084] Through this decomposition method, the model can better capture the long-term trend and periodic pattern in the time series, thereby improving the prediction accuracy.

[0085] Step 3) Predict the trend term: Use the RLinear model to predict the trend term data through a normalization layer, a multi-layer perceptron, and a linear projection layer to obtain the trend term prediction result; the RLinear model includes a normalization layer, a multi-layer perceptron, and a linear projection layer. Specifically, predicting the trend term is as follows:

[0086] Perform normalization processing on the trend term. The normalization processing uses a reversible normalization layer, whose function is to remove the distribution differences between different time periods of the data and ensure the robustness of the model when dealing with long-term trends.

[0087] Input the normalized trend term into the multi-layer perceptron to extract the features in the time series. The multi-layer perceptron can extract important time features from the input data for subsequent trend term prediction.

[0088] The extracted features are projected using a linear projection layer to generate the final trend term prediction result. The linear projection layer maps the high-dimensional features to the target dimension in order to generate a suitable prediction output.

[0089] The above projection prediction result is denormalized through an invertible normalization layer to restore the original distribution of the data, obtaining the trend term prediction result:

[0090]

[0091] In Equation (3), X trend represents the trend term, represents the trend term prediction result, RevIN norm represents the normalization process, and RevIN denorm represents the denormalization process. W represents the weight matrix of the linear projection layer.

[0092] Step 4) Predict the seasonal term: including:

[0093] 4-1) Perform an initial downsampling on the seasonal data, and extract the first-layer time pattern τ 1 in the seasonal data through an average pooling operation. The seasonal term data is decomposed into different time patterns using a multi-scale decomposition method, and different time patterns are mixed to obtain enhanced seasonal term data, specifically as follows:

[0094] Downsample the seasonal term data to extract the first-layer time pattern:

[0095] τ 1 = AvgPooling(x seasonal ) (4)

[0096] In Equation (4), τ 1 represents the first-layer time pattern, and AvgPooling represents the average pooling operation;

[0097] Further process the first-layer time pattern, and use the average pooling operation to downsample each layer of time pattern to extract multi-layer time pattern information. That is, for each subsequent layer of time pattern, it is extracted by further performing an average pooling operation on the previous layer of time pattern:

[0098] τ i = AvgPooling(τ i-1 ), i = 2, 3, …, h (5)

[0099] In Equation (5), h represents the total number of layers of the downsampling operation;

[0100] Mix the time patterns of different granularities. By mixing the time pattern τ of the previous layeri+1 Input the multi-layer perceptron and combine it with the time pattern τ of the current layer i to generate a hybrid time series with rich temporal information: fuse the time pattern layer by layer through the multi-layer perceptron into u i :

[0101] u i = τ i + MLP(τ i+1 ), i = 1, 2, …, h - 1 (6)

[0102] In Equation (6), u i represents the i-th time series information after mixing;

[0103] Finally fuse the time patterns of all layers to finally obtain the enhanced seasonal term data u 1 , briefly denoted as u. Through multi-level decomposition and fusion, this information contains a complete representation of multi-scale time patterns and can effectively capture the time series characteristics at different scales.

[0104] 4 - 2) Input the enhanced seasonal term data into the mixture of experts model;

[0105] 4 - 3) Each expert in this mixture of experts model will output their respective prediction results; at the same time, input the enhanced seasonal term data into the Top-K gating mechanism to obtain the weights assigned to each expert, multiply the results output by each expert by the weights to obtain the prediction result of the mixture of experts model; the specific content is as follows:

[0106] Input the enhanced seasonal term data u obtained in step 4 - 1) into the Top-K gating mechanism to calculate the scores of each expert. First, perform preliminary processing on the input data through a linear transformation to generate the scores of each expert. Then, use the Softmax function to normalize the scores to ensure that the sum of the scores is 1, representing the weights of the experts. Next, perform normalization again by selecting the k experts with the highest scores. The scores S of the finally selected experts are calculated by the following formula:

[0107] S = Softmax(TopK(Softmax(u), k)) (7)

[0108] In Equation (7), k is the number of experts selected, and S is the weight assigned to each expert.

[0109] Perform predictions using a mixture of experts structure. Each expert in the mixture of experts model receives the enhanced seasonal term data obtained in step 4-1) as input and processes it through a deep neural network with an activation layer and an output layer. Specifically, the input first extracts features through a linear layer, then undergoes a non-linear transformation through the GELU activation function, and finally generates the final prediction result through the output layer. The output result obtained by each expert is:

[0110] E i (u) = Linear output (GELU(Linear hidden (u))) (8)

[0111] In equation (8), and r is the dimension of the low-rank matrix;

[0112] Using the outputs of the selected k experts, generate the final seasonal time series prediction result using their weighted sum:

[0113]

[0114] 4-4) Calculate the extreme value loss based on the prediction result of the mixture of experts model obtained in step 4-3) and the enhanced seasonal term data obtained in step 4-1), and automatically update the network parameters inside the mixture of experts model through the extreme value loss; After multiple updates, until the number of iterations meets the set requirements, obtain the seasonal term prediction result. The specific content is as follows:

[0115] Calculate the predicted value of the mixture of experts model and the true value Y t = {y t+1 , y t+2 , …, y t+H}, and obtain the prediction loss

[0116]

[0117] In equation (10), H represents the predicted time step, is the mean squared error loss;

[0118] According to the upper and lower thresholds θ high and θ lowSegment the time series data, and divide the true values and predicted values into left extremes (greater than the upper threshold), middle values (between the upper and lower thresholds), and right extremes (less than the lower threshold) respectively. Specifically: Segment according to the upper and lower quartiles of the data, and call the data points greater than the upper quartile, less than the upper quartile and greater than the lower quartile, and less than the lower quartile the left extreme, middle value, and right extreme respectively; assign 1 to the left extreme, 0 to the middle value, and -1 to the right extreme; the results after assignment of the true value and the predicted value are denoted as e t and That is:

[0119]

[0120] In equations (11) and (12), θ high and θ low are the upper quartile and the lower quartile respectively.

[0121] Calculate the extreme value loss By measuring the difference between the predicted extreme value sequence and the true extreme value sequence E t calculate the loss between the two:

[0122]

[0123] Calculate the difference between the predicted extreme value sequence and the true extreme value sequence, and denote the proportion of non-zero terms in this difference, that is, the inconsistency proportion, as p. This p represents the proportion of the terms where the predicted value and the true value are not equal in the total number of terms.

[0124]

[0125] In equation (14), 1 is an indicator function used to calculate the number of inconsistent terms between the predicted extreme value sequence and the true extreme value sequence;

[0126] Finally, combine and Calculate the weighted loss function through p:

[0127]

[0128] According to the result of the extreme value loss, the mixture of experts model automatically updates the internal network parameters. When the number of updates meets the set requirements, the result output by the mixture of experts model is the seasonal term prediction result.

[0129] Step 5) Add the trend term prediction result obtained in step 3) and the seasonal term prediction result obtained in step 4) to obtain the final time series prediction result

[0130]

[0131] This summation process ensures that the model simultaneously considers both the long-term trend and short-term seasonal fluctuations in the time series, thus generating more accurate predictions.

[0132] Research materials The research materials used the ETTh1, ETTh2, ETTm1, and ETTm2 datasets. The mean squared error (MSE) and mean absolute error (MAE) evaluation metrics were used to evaluate the model performance. Table 1 shows the comparison of the results of the method proposed in the present invention with other methods.

[0133] Table 1

[0134]

[0135] Although the present invention has been described above in conjunction with the accompanying drawings, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can make many improvements and variations without departing from the gist of the present invention, and these all fall within the protection scope of the present invention.

Claims

1. A time series prediction method based on extreme value loss and hybrid expert structure, characterized in that: The method comprises the following steps: Step 1) Data preparation: pre-process the acquired time series data and divide it into training set, validation set and test set according to the predetermined ratio; and process the missing values ​​and outliers of the acquired historical data; Step 2) sequence decomposition, decomposing the time series data preprocessed in step 1), using moving average to calculate trend item data; subtracting trend item data from preprocessed time series data to obtain seasonal item data; Step 3) Predicting trend items: Using the RLinear model, the trend item data is predicted through the normalization layer, the multi-layer perceptron and the linear projection layer to obtain the trend item prediction result; Step 4) Predict seasonal items: including: 4-1) Using a multi-scale decomposition method to decompose seasonal item data into different time modes, and mixing different time modes to obtain enhanced seasonal item data; 4-2) Inputting the enhanced seasonal item data into the mixed expert model; 4-3) Each expert in the hybrid expert model will output their own prediction results; at the same time, the enhanced seasonal item data is input into the Top-K gating mechanism to obtain the weight assigned to each expert, and the result output by each expert is multiplied by the weight to obtain the prediction result of the hybrid expert model; 4-4) Calculating the extreme value loss based on the prediction result of the hybrid expert model obtained in step 4-3) and the enhanced seasonal item data obtained in step 4-1), and automatically updating the network parameters inside the hybrid expert model through the extreme value loss; after multiple updates, until the number of iterations meets the set requirements, the seasonal item prediction result is obtained; Step 5) adds the trend item prediction result obtained in step 3) and the season item prediction result obtained in step 4) to obtain the final time series prediction result.

2. The time series prediction method according to claim 1, characterized in that: The specific contents of the step 2) are as follows: Step 2-1) Decompose the time series data preprocessed in step 1) and use the moving average to calculate the trend item data: In formula (1), x trend (t) represents the t-th trend item, K is the window size of the moving average, and x(ti) represents the ti-th time series data; Step 2-2) Subtract the trend term data from the preprocessed time series data to obtain the seasonal term data to represent periodic fluctuations. The formula is: x seasonal (t)=x(t)-x trend (t) (2) In formula (2), x seasonal (t) represents the t-th seasonal term, and x(t) represents the t-th time series data.

3. The time series prediction method according to claim 1, characterized in that: In step 3), the RLinear model includes a normalization layer, a multilayer perceptron, and a linear projection layer. The specific content of step 3) is as follows: Step 3-1) normalize the trend item data through a normalization layer; Step 3-2) Input the normalized trend item data into the multi-layer perceptron to extract the features of the time series; Step 3-3) Use a linear projection layer to project the features extracted by the multilayer perceptron; Step 3-4) Denormalize the projection result of step 3-3) to obtain the trend item prediction result: In formula (3), X trend Represents the trend item, Represents the trend item prediction result, RevIN norm Represents the normalization process, RevIN denorm represents the denormalization process, and W represents the weight matrix of the linear projection layer.

4. The time series prediction method according to claim 1, characterized in that: The specific contents of step 4-1) are as follows: Step 4-1-1) Downsample the seasonal data to extract the first layer of time pattern: τ1=AvgPooling(x seasonal ) (4) In formula (4), τ1 represents the first layer time pattern, AvgPooling represents the average pooling operation; Step 4-1-2) For each subsequent layer of temporal patterns, further average pooling operation is performed on the previous layer of temporal patterns to extract: τ i =AvgPooling(τ i-1 ),i=2,3,…,h (5) In formula (5), h represents the total number of layers of downsampling operations; Step 4-1-3) Use a multi-layer perceptron to fuse the temporal patterns layer by layer into u i : u i =τ i +MLP(τ i+1 ),i=1,2,…,h-1 (6) In formula (6), u i Represents the i-th time series information after mixing; Finally, the enhanced seasonal item data u1 is obtained, abbreviated as u.

5. The time series prediction method according to claim 1, characterized in that: The specific contents of step 4-3) are as follows: 4-3-1) Input the enhanced seasonal item data obtained in step 4-1) into the Top-K gating mechanism to obtain the score of each expert. After normalizing the score through the Softmax function, select the largest k scores and normalize them again: S=Softmax(TopK(Softmax(u),k)) (7) In formula (7), k is the number of selected experts, and S is the assigned weight of each expert; 4-3-2) Each expert in the hybrid expert model receives the enhanced seasonal item data obtained in step 4-1) as input, and inputs it into a deep neural network with an activation layer and an output layer. The output result obtained by each expert is: E i (u)=Linear output (GELU(Linear hidden (in))) (8) In formula (8), and r is the dimension of the low-rank matrix; 4-3-3) The output results of the selected k expert models are multiplied by the weights to obtain the prediction results of the hybrid expert model:

6. The time series prediction method according to claim 1, characterized in that: The specific contents of step 4-4) are as follows: Step 4-4-1) Calculate the prediction value of the hybrid expert model With the true value The loss between them is used to get the prediction loss In formula (10), H represents the predicted time step, is the mean square error loss; 4-4-2) Divide the data into segments according to the upper quartile and the lower quartile. The data points that are greater than the upper quartile, less than the upper quartile and greater than the lower quartile, and less than the lower quartile are called the left extreme value, the middle value, and the right extreme value respectively; the left extreme value is assigned 1, the middle value is assigned 0, and the right extreme value is assigned -1; the results after the assignment of the true value and the predicted value are recorded as e t and Right now: In formula (11) and formula (12), θ high and θ low are the upper and lower quartiles, respectively; 4-4-3) Calculate extreme value loss The extreme value loss is the loss between the predicted extreme value sequence and the true extreme value sequence: 4-4-4) Calculate the difference between the predicted extreme value sequence and the actual extreme value sequence, and record the proportion of non-zero items in the difference as p, that is: In formula (14), 1 is an indicator function, which is used to calculate the number of inconsistent items between the predicted extreme value sequence and the actual extreme value sequence; 4-4-5) Combined prediction loss and extreme value loss Calculate the final loss function based on the non-zero proportion p 4-4-6) Based on the results of extreme value loss, the hybrid expert model automatically updates the internal network parameters. When the number of updates meets the set requirements, the output result of the hybrid expert model is the seasonal item prediction result.

Citation Information

Cited By

  • Multi-sensor data processing method suitable for high-altitude meteorological detection

    CN121276654A