Time series prediction model incremental updating method and device, equipment and medium

CN122778028APending Publication Date: 2026-09-18THE HONG KONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510318568.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0004]然而,在数据更新快、数据多的情况下,当面临大量增量数据时,传统基于增量学习的时间序列预测模型计算开销大,更新速度有待提高

Benefits of technology

[0049] By comparing the predicted loss value obtained from the current model's predictions based on the current samples in the incremental dataset with an update threshold, it is determined whether to use the current samples for incremental updates. This introduces the concept of triggering incremental updates based on an update threshold. This filters the data in the incremental dataset, avoiding the need for incremental updates on all data, minimizing unnecessary computational overhead, improving computational efficiency, and thus increasing the model's update speed. Furthermore, by calculating the update threshold based on the old model and the training dataset, a cleverly set trigger threshold can ensure model accuracy while reducing unnecessary computational overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122778028A_ABST
    Figure CN122778028A_ABST
Patent Text Reader

Abstract

The application provides a time series prediction model incremental updating method and device, equipment and medium. The method comprises the following steps: obtaining an existing time series prediction old model, a training data set and an incremental data set; calculating an updating threshold based on the old model and the training data set; taking the old model as an initialized current new model, performing an incremental updating step on the current new model: sequentially selecting a sample in the incremental data set as a current sample, predicting by using the current new model for the current sample to calculate a prediction loss value; if the prediction loss value is greater than or equal to the updating threshold, performing incremental updating on the current new model by using the current sample to obtain an updated new model as the current new model of the next incremental updating; repeating the incremental updating step until all samples in the incremental data set are traversed to obtain an incremental updating model of the old model. By using the above scheme, the model updating speed can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data prediction technology, and in particular to an incremental update method, apparatus, device and medium for time series prediction models. Background Technology

[0002] Time series forecasting models are commonly used to predict data based on time series data and have a wide range of applications, such as in the transportation and finance sectors. Training time series forecasting models using network structures, such as LSTM (Long Short-Term Memory), GRU (Gated Recurrent Unit), and Transformer network structures, requires a significant amount of time.

[0003] Traditional full-data update methods struggle to meet the demands of rapidly adapting to large amounts of new data. Incremental learning has become a key technique for improving the performance of time series forecasting models. Through incremental learning, models can dynamically acquire knowledge from new data and adjust themselves in a timely manner to better adapt to dynamic changes. This approach increases the training speed of models, enabling them to capture data fluctuations over time without sacrificing too much of their existing performance, thus providing users with more reliable decision support.

[0004] However, when dealing with rapid data updates and large amounts of data, traditional time series prediction models based on incremental learning suffer from high computational costs and need to improve update speed when faced with a large amount of incremental data. Summary of the Invention

[0005] This application provides an incremental update method, apparatus, device, and medium for time series forecasting models to solve the aforementioned technical problems in the prior art.

[0006] According to a first aspect of this application, an incremental update method for a time series forecasting model is provided, comprising:

[0007] Obtain existing time series prediction models, training datasets, and incremental datasets;

[0008] Calculate the update threshold based on the old model and the training dataset;

[0009] Using the old model as the initialization point, the following incremental update steps are performed on the current new model:

[0010] One sample from the incremental dataset is selected as the current sample in sequence. For the current sample, the new model is used to make a prediction to calculate the prediction loss value.

[0011] Compare the predicted loss value with the updated threshold;

[0012] If the predicted loss value is greater than or equal to the update threshold, the current sample is used to incrementally update the current new model to obtain the updated new model, which is then used as the current new model for the next incremental update.

[0013] Repeat the above incremental update steps until all samples in the incremental dataset have been traversed. Use the updated model obtained from the last incremental update step as the incremental update model of the old model.

[0014] In some embodiments, the step of using the current new model to make predictions for the current sample to calculate the prediction loss value includes:

[0015] For the current sample, forward propagation is performed using the current new model to obtain the new model's predicted value;

[0016] The predicted loss value is obtained by calculating the degree of difference between the predicted value of the new model and the true value corresponding to the current sample.

[0017] In some embodiments, incrementally updating the current new model using the current samples includes:

[0018] The old model is used to perform forward propagation on the current sample to obtain the old model's predicted value;

[0019] Calculate the distillation loss value based on the old model prediction value and the new model prediction value;

[0020] Calculate the distillation loss function based on the predicted loss value and the distillation loss value;

[0021] Backpropagation is performed on the current new model based on the distillation loss function to update the model parameters of the current new model.

[0022] In some embodiments, the distillation loss function is calculated using the following formula:

[0023] loss = loss new +p*loss dist ;

[0024] Wherein, the loss new Let loss be the predicted loss value. dist Let p be the distillation loss value, and p be the distillation coefficient, where p ∈ [0.05, 0.2].

[0025] In some embodiments, calculating the update threshold based on the old model and the training dataset includes:

[0026] The old model is used to make predictions based on the training dataset to calculate the loss value;

[0027] Calculate the mean and standard deviation of all loss values;

[0028] The updated threshold is calculated based on the average value and the standard deviation.

[0029] In some embodiments, calculating the update threshold based on the average value and the standard deviation includes calculating the update threshold using the following formula:

[0030] K = μ + α·σ;

[0031] Where K is the update threshold, μ is the average value, σ is the standard deviation, and α is the hyperparameter, and α∈[0.3, 0.7].

[0032] In some embodiments, the method further includes: performing model evaluation on the incremental update model.

[0033] In some embodiments, the model evaluation of the incremental update model includes:

[0034] The union of the training dataset and the incremental dataset is taken as the full dataset, and the old model is trained using the full dataset to obtain the full update model.

[0035] Obtain a test dataset, and for the test dataset, make predictions using the incremental update model and the full update model respectively to obtain the first test result data of the incremental update model and the second test result data of the full update model;

[0036] The first evaluation result of the incremental update model is obtained by comparing the second test result data with the first test result data.

[0037] In some embodiments, the model evaluation of the incremental update model further includes:

[0038] The training time of the incremental update model and the training time of the full update model are statistically analyzed.

[0039] The ratio of the training time of the full update model to the training time of the incremental update model is calculated to obtain the second evaluation result of the incremental update model.

[0040] According to a second aspect of this application, an incremental update apparatus for a time series forecasting model is provided, the apparatus comprising:

[0041] The data acquisition module is used to acquire existing time series prediction models, training datasets, and incremental datasets.

[0042] The threshold calculation module is used to calculate the updated threshold based on the old model and the training dataset;

[0043] The incremental update module is used to perform an incremental update step on the current new model, which is initialized with the old model. The incremental update step includes:

[0044] One sample from the incremental dataset is selected sequentially as the current sample. For the current sample, the current new model is used to make a prediction to calculate the prediction loss value. The prediction loss value is compared with the update threshold. If the prediction loss value is greater than or equal to the update threshold, the current sample is used to incrementally update the current new model to obtain the updated new model, which is used as the current new model for the next incremental update.

[0045] The incremental update module repeatedly executes the incremental update step until all samples in the incremental dataset have been traversed, and the updated model obtained from the last incremental update step is used as the incremental update model of the old model.

[0046] According to a third aspect of this application, an electronic device is provided, comprising: a processor and a memory storing computer program instructions; the processor, when executing the computer program instructions, implements any of the above-described incremental update methods for time series forecasting models.

[0047] According to a fourth aspect of this application, a computer-readable storage medium is provided, characterized in that the computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement any of the above-described incremental update methods for time series prediction models.

[0048] In summary, the incremental update method, apparatus, equipment, and medium for time series forecasting models provided in this application have at least the following beneficial effects:

[0049] By comparing the predicted loss value obtained from the current model's predictions based on the current samples in the incremental dataset with an update threshold, it is determined whether to use the current samples for incremental updates. This introduces the concept of triggering incremental updates based on an update threshold. This filters the data in the incremental dataset, avoiding the need for incremental updates on all data, minimizing unnecessary computational overhead, improving computational efficiency, and thus increasing the model's update speed. Furthermore, by calculating the update threshold based on the old model and the training dataset, a cleverly set trigger threshold can ensure model accuracy while reducing unnecessary computational overhead. Attached Figure Description

[0050] To more clearly illustrate the technical solutions in the specific embodiments of this application, the accompanying drawings will be briefly introduced below in conjunction with the accompanying drawings. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings or solutions can be obtained based on these drawings without creative effort.

[0051] Figure 1 This is a flowchart of an incremental update method for a time series prediction model in one embodiment of this application;

[0052] Figure 2 This is a flowchart of the incremental update step in one embodiment of this application;

[0053] Figure 3 This is a flowchart of an incremental update method for a time series prediction model in another embodiment of this application;

[0054] Figure 4 This is a structural diagram of an incremental update device for a time series prediction model in one embodiment of this application;

[0055] Figure 5 This is a structural diagram of an electronic device provided as an embodiment of the present application. Detailed Implementation

[0056] To make the above and other features and advantages of this application clearer, the application is further described below with reference to the accompanying drawings. It should be understood that the specific embodiments given herein are for the purpose of explanation to those skilled in the art, and are exemplary only, not restrictive.

[0057] In the following description, numerous specific details are set forth to provide a thorough understanding of this application. However, it will be apparent to those skilled in the art that the specific details are not required to practice this application. In other instances, well-known steps or operations have not been described in detail to avoid obscuring this application.

[0058] The incremental update method for time series prediction models provided in this application can be executed by the incremental update device for time series prediction models provided in this application, which can be configured in an electronic device.

[0059] In this application: the old model for time series forecasts before the incremental update is defined as M. old Its training dataset is defined as D old ={(X old ,Y old )},in X old It is the training input data, N old It is the sample size, T oldY is the time length of each sample in the training dataset. old It is the true value to be predicted for the sample in the training dataset. Similarly, define D. new It is an incremental dataset, D val It is the validation set, D test This is the test dataset. New model M new Based on M old Perform incremental updates, specifically, M new Use the incremental training framework to read D new Model learning and parameter updates are performed. Furthermore, a full update model M is defined. F M F In the full dataset D F ={D old ∪D new Training was started from scratch on}. All models were eventually trained on D. test An evaluation will be conducted.

[0060] For example, taking traffic flow prediction using a time series forecasting model as an example, X old It can be the training input sequence of traffic flow, and the true value Y of the sample. old The input sequence contains the actual traffic flow for the second day at the last time step.

[0061] For example, taking stock price prediction using a time series forecasting model, X old The training input sequence is a sequence of stock prices, for example, stock prices sorted by time, and the true value Y of the sample is... old The input sequence contains the closing price of the stock on the second day following the last time step.

[0062] refer to Figure 1 and Figure 2 This application provides an incremental update method for a time series forecasting model, which includes the following steps S110 to S170.

[0063] S110: Obtain existing time series prediction models, training datasets, and incremental datasets.

[0064] The old model M for time series forecasting old It has already acquired some knowledge by training on historical data. The training dataset is a set of old data over a certain period of time, while the incremental dataset is a set of new data over another period of time.

[0065] For example, to predict traffic flow, the training dataset can be historical traffic flow sorted by time.

[0066] For example, when performing stock price prediction, the training dataset can be historical stock prices sorted by time. For instance, the old data is the closing price time series of all A-share stocks from January 1, 2008 to December 31, 2017, and the new data can be the closing price time series of all A-share stocks from January 1, 2018 to December 31, 2018.

[0067] S130: Calculate the update threshold based on the old model and training dataset.

[0068] Specifically, without updating the model parameters, the old model is trained on the training dataset, and the update threshold is calculated based on the loss value of the training iterations.

[0069] S150: The current new model is initialized with the old model, and incremental update steps are performed on the current new model.

[0070] Among them, reference Figure 2 The incremental update steps include the following steps S151 to S155.

[0071] S151: Select one sample from the incremental dataset as the current sample in turn, and use the new model to make a prediction for the current sample in order to calculate the prediction loss value.

[0072] The data in the incremental dataset corresponds to the samples used for model updates. The sample size can be determined by the training parameters set by the user. For example, if the batch_size is set to 1 in the training parameters, it means that the sample size is 1, and the number of samples drawn in one training session is 1. Specifically, for a sample in the incremental dataset, the current new model is used to make a prediction without updating the model parameters. The prediction loss value corresponding to this sample is calculated based on the prediction result and used as the loss value for the current new model to make predictions on this sample.

[0073] S153: Compare the predicted loss value with the update threshold.

[0074] S155: If the predicted loss value is greater than or equal to the update threshold, the current sample is used to perform an incremental update on the current new model to obtain the updated new model, which will be used as the current new model for the next incremental update.

[0075] The predicted loss value corresponding to the current sample is compared with the update threshold. If the predicted loss value is greater than or equal to the update threshold, the update trigger condition is met, and the current sample is used to incrementally update the current new model. Specifically, if the predicted loss value is less than the update threshold, it is considered that the model already has knowledge corresponding to the new data, and the current sample is not used for incremental updates, thus filtering the samples, that is, filtering the data in the incremental dataset according to the update threshold.

[0076] S170: Repeat the above incremental update steps until all samples in the incremental dataset have been traversed. Use the updated model obtained from the last incremental update step as the incremental update model of the old model.

[0077] Specifically, for the first sample in the incremental dataset, the predicted loss value is calculated using the same new model as the old model. If the predicted loss value of the first sample is less than the update threshold, the first sample is not used for incremental updates, and the same new model as the old model is used to calculate the predicted loss value of the second sample. If the predicted loss value of the first sample is greater than or equal to the update threshold, the first sample is used for incremental updates to obtain the updated new model. Then, the updated new model is used to calculate the predicted loss value of the second sample, and the update trigger condition is checked again. If the condition is met, an incremental update is performed; otherwise, no update is performed. In this way, the update trigger condition is checked for each sample in the incremental dataset, and the data in the incremental dataset is selectively updated.

[0078] Traditional non-selective incremental update methods require updating all data when faced with large amounts of incremental data. The incremental update method for time series prediction models described above differs from traditional non-selective incremental update methods. It compares the predicted loss value obtained by the new model based on the current samples in the incremental dataset with an update threshold to determine whether to use the current samples for incremental updates. This introduces the concept of triggering incremental updates based on an update threshold, filtering the data in the incremental dataset and avoiding the need to update all data, thus minimizing unnecessary computational overhead and improving computational efficiency, thereby increasing the model update speed. Furthermore, the selection of the triggering condition involves determining when to initiate incremental updates, which can be based on various factors such as time intervals, changes in data volume, or changes in model performance. This application cleverly sets the trigger threshold by calculating the update threshold based on the old model and training dataset, ensuring model accuracy while reducing unnecessary computational overhead. Based on this, this application can improve the update speed of time series prediction models, control the loss of model accuracy compared to full updates, and achieve accurate and efficient incremental updates for time series prediction models. When applied to predictions with large amounts of updated data, it can better adapt to dynamic changes.

[0079] In some embodiments, step S130 includes steps (a1) to (a3).

[0080] Step (a1): Use the old model to make predictions based on the training dataset to calculate the loss value.

[0081] Specifically, using the old model M old In the training dataset D oldThe training proceeds, but the gradients are not updated, keeping the model parameters unchanged. The batch size is set to 1, meaning one sample is used per training iteration. After multiple iterations of training, the loss value for each iteration is obtained.

[0082] Step (a2): Calculate the mean and standard deviation of all loss values.

[0083] Step (a3): Calculate the updated threshold based on the mean and standard deviation.

[0084] In this application, the original training dataset D is used. old Calculate the update threshold without using the incremental dataset D as new data to be learned. new To avoid data leaks.

[0085] In some embodiments, step (a3) ​​calculates the update threshold using the following formula (1):

[0086] K = μ + α·σ; Formula (1)

[0087] Where K is the update threshold, μ is the mean, σ is the standard deviation, and α is the hyperparameter, and α∈[0.3, 0.7].

[0088] Specifically, α is preferably set to 0.5. Experiments and normality tests revealed that using the old model M... old Perform non-training tests on the old training dataset (keeping the model parameters unchanged). After training reaches a stable state, M... old In D old The loss value approximately follows a Gaussian distribution loss~N(μ,σ). The mean μ and standard deviation σ of the Gaussian distribution are calculated, and the update threshold is calculated accordingly.

[0089] In some embodiments, in step S151, a prediction is made using the current new model for the current sample to calculate the prediction loss value, including the following steps (b1) to (b2).

[0090] Step (b1): For the current sample, use the current new model to perform forward propagation to obtain the new model prediction value.

[0091] Forward propagation using the current new model for prediction will not update the model parameters of the current new model.

[0092] Step (b2): Calculate the degree of difference between the new model's predicted value and the actual value corresponding to the current sample to obtain the predicted loss value.

[0093] Here, the degree of difference is data that measures the difference between two values. For example, the degree of difference can be MSE (mean squared error). For instance, for a current sample, using the current new model M...new The new model prediction value P is obtained by performing forward propagation. new And calculate the new model prediction value P. new The mean squared error (MSE) between the true value of this sample and the actual value is used as the prediction loss. new It is understandable that the degree of difference can also be represented by other data, such as KL divergence (relative entropy). By employing steps (b1) and (b2), the loss value predicted by the current new model on the sample can be accurately obtained.

[0094] In some embodiments, step S155, the step of incrementally updating the current new model using the current samples, includes the following steps (c1) to (c4).

[0095] Step (c1): Use the old model to perform forward propagation on the current sample to obtain the old model's predicted value.

[0096] For example, if the current sample is A, and the predicted loss value calculated by the new model for sample A is greater than or equal to the update threshold, then the old model M is used. old Forward propagation of this sample A yields the old model's predicted value P. old .

[0097] Step (c2): Calculate the distillation loss value based on the old model prediction and the new model prediction.

[0098] Specifically, step (c2) includes: calculating the degree of difference between the old model's predicted values ​​and the new model's predicted values ​​to obtain the distillation loss value. The degree of difference can be MSE, and the distillation loss value is [missing value]. dist =MSE(P new ,P old It is understandable that KL divergence can be used instead of MSE to measure the difference between two distributions.

[0099] Step (c3): Calculate the distillation loss function based on the predicted loss value and the distillation loss value.

[0100] Specifically, the distillation loss function includes two terms: the predicted loss value and the distillation loss value. The predicted loss value of the current new model is used to learn incremental data, while the distillation term is used to limit changes in the model's parameters, so that the new model retains as much knowledge as possible from the old model.

[0101] Specifically, the distillation loss function is calculated using the following formula (2):

[0102] loss = loss new +p*loss dist ; Formula (2)

[0103] Where, loss newTo predict the loss value, loss dist denoted as , where is the distillation loss value. p is the distillation coefficient, p∈[0.05,0.2], preferably set to 0.1.

[0104] Step (c4): Backpropagate the current new model based on the distillation loss function to update the model parameters of the current new model.

[0105] Backpropagation of the current new model updates the gradients and model parameters. Specifically, this can be achieved by performing incremental updates based on gradient descent.

[0106] This application uses a distillation loss function during incremental model updates, enabling a balance between preserving existing knowledge and learning new knowledge when learning incremental data by adding a distillation loss value. Specifically, assuming the old model is M... old The model has already been trained on historical data and has acquired some knowledge. The model updated on the incremental dataset is denoted as M. new When new incremental data is introduced, use that data to evaluate M. new To update and preserve old knowledge, M new In the prediction results P of incremental data new With M old In the prediction results P of incremental data old The smaller the MSE or KL divergence, the better. The introduction of the distillation loss function allows the new model to learn new data while retaining as much knowledge as possible from the old model. This is achieved by minimizing the MSE or KL divergence, ensuring that the new model's prediction distribution on new data is similar to that of the old model.

[0107] The model uses a distillation loss function during incremental updates. This strategy performs well in applications such as stock price prediction and traffic flow prediction because it allows for flexible adaptation to new data features while retaining the understanding of historical trends.

[0108] In some embodiments, the incremental update method for time series forecasting models further includes an evaluation step: evaluating the incremental update model.

[0109] Specifically, refer to Figure 3 The evaluation steps may include steps S191 to S195.

[0110] S191: Take the union of the training dataset and the incremental dataset as the full dataset, and use the full dataset to train the old model to obtain the full updated model.

[0111] The full dataset includes all data from the training dataset and the incremental dataset. The old model for time series prediction is updated on the full dataset to obtain the full updated model.

[0112] S193: Obtain the test dataset. For the test dataset, perform predictions using the incremental update model and the full update model respectively, and obtain the first test result data of the incremental update model and the second test result data of the full update model.

[0113] The test dataset is used to test the model. For example, the first test result data includes the first predicted values ​​obtained by testing the data in the test dataset using the incremental update model, and the second test result data includes the second predicted values ​​obtained by using the full update model.

[0114] For example, the preset value could be traffic flow, or the predicted value could be stock price.

[0115] S195: By comparing the second test result data with the first test result data, the first evaluation result of the incremental update model is obtained.

[0116] For example, the first evaluation result may include the information coefficient difference and the ranking information coefficient difference. Specifically, step S195 includes: calculating the first information coefficient and the first ranking information coefficient based on the first test result data and the true values ​​corresponding to the test dataset; calculating the second information coefficient and the second ranking information coefficient based on the second test result data and the true values ​​corresponding to the test dataset; calculating the difference between the first information coefficient and the second information coefficient to obtain the information coefficient difference; and calculating the difference between the first ranking information coefficient and the second ranking information coefficient to obtain the ranking information coefficient difference.

[0117] The formulas for calculating the Information Coefficient (IC) and the Rank Information Coefficient (RankIC) are as follows:

[0118]

[0119] Among them, Y test and rank(Y) test The test datasets D and D are respectively. test The true values ​​of all data and the numerical values ​​obtained based on the data are sorted. and The test dataset D is shown below. test The predicted values ​​of all data and the numerical values ​​obtained based on the predicted values ​​are sorted, with corr representing the correlation calculation.

[0120] For example, in one embodiment, Y test To test the true values ​​of traffic flow at multiple times in the dataset, rank(Y) test This corresponds to the sorting of the actual traffic flow values; To test the predicted traffic flow values ​​in the dataset, The predicted traffic flow values ​​are sorted. In another embodiment, Y test To test the true values ​​of stock prices at multiple times in the dataset, rank(Y) test This corresponds to the sorting of the true values ​​of stock prices; To test the predicted stock prices in the dataset, The predicted stock prices are sorted.

[0121] The formulas for calculating the information coefficient difference and the ranking information coefficient difference are as follows:

[0122] IC′=IC new -IC F

[0123] RankIC′=RankIC new -RankIC F

[0124] Among them, IC new and RankIC new It is an incremental update model M new Based on test dataset D test The obtained first information coefficient and first ranking coefficient, IC F and RankIC F It is a full update model M F Based on test dataset D test The obtained second information coefficient and second ranking coefficient, IC′ is the information coefficient difference, and RankIC′ is the ranking information coefficient difference.

[0125] By calculating the incremental update model and the full update model on the test dataset D test The difference between IC and RankIC can be used as an evaluation metric to accurately assess the accuracy loss of the incremental update model compared to the full update model.

[0126] In some embodiments, reference Figure 3 The evaluation process also includes steps S197 to S199.

[0127] S197: Training time for the statistical incremental update model and the training time for the full update model.

[0128] S199: Calculate the ratio of the training time of the full update model to the training time of the incremental update model to obtain the second evaluation result of the incremental update model.

[0129] For example, the incremental update model M obtained through incremental updates new In incremental dataset D newThe training duration is T. new Full update of model M F In D F ={D old ∪D new The training time on} is T. F The formula for calculating the second evaluation result is as follows:

[0130]

[0131] Here, TimeRatio represents the value of the second evaluation result. By calculating the second evaluation result, the improvement in training speed between incremental updates and full updates can be accurately assessed.

[0132] Furthermore, the incremental update method of the time series forecasting model in this application is described in detail through a case study as follows:

[0133] This implementation uses 300 stocks from the A1 and A2 stock exchanges. Data from January 1, 2012 to December 31, 2015 is used as the training dataset D. old Data from January 1, 2016 to December 31, 2016 is used as the incremental dataset D. new The data from January 1, 2017 to December 31, 2017 will be used as the validation dataset D. val The data from January 1, 2018 to December 31, 2019 will be used as the test dataset D. test .

[0134] First, the data is preprocessed by standardizing the daily stock prices according to the following formula:

[0135]

[0136] Where closingprice represents the closing price of the stock on that day, and the subscript t represents day t.

[0137] Then, read an old model M that has already been trained on the training dataset. old .

[0138] Next, the old model is incrementally updated. The entire process of incremental training of the old model is timed to obtain T. new .

[0139] Specifically, it includes the following:

[0140] (1) Learning about incremental update triggers

[0141] Use M old In D oldRetrain, but do not update the gradient. Set batch_size to 1, calculate the average value μ and standard deviation σ of the loss values ​​for all iterations, and calculate the update threshold K = μ + α·σ, where α is set to 0.5.

[0142] (2) Incremental update of model parameters

[0143] Perform incremental updates on the model, also setting the batch_size to 1. For D new For each sample, the current new model M is first used. new (The initial new model is M) old The current new model (which is the same as the previous updated new model) is then forward-propagated to obtain the new model prediction value P. new The loss value is calculated as the MSE between the new model's predicted values ​​and the true values ​​of the samples. new If loss new If it is less than K, then M is not performed. new Gradient calculation (backpropagation) and gradient update. If the loss value is greater than or equal to K, then M is used. old Forward propagation of this sample yields the old model's predicted value P. old And calculate the loss. dist =MSE(P new ,P old The final loss value of the distillation loss function is obtained as loss = loss new +p*loss dist p is set to 0.1. Based on the final loss value, M is... new Update the model to serve as the current new model for training the next sample.

[0144] Finally, the model is evaluated using incremental updates of model M. new For D test Test and obtain IC new and RankIC new Metrics. For comparison, a full-update model M also needs to be trained. F And calculate the training time T for a full update of the model. F And finally in D test ICs being tested F and RankIC F Indicators. Ultimately, compare T. new and T F and IC new RankIC new and IC F RankIC F According to tests, T new Only T FAbout 2 / 3, while IC new RankIC new Compared to IC F RankIC F The accuracy loss is controlled within 5%.

[0145] This application provides an incremental update device for a time series forecasting model, referencing... Figure 4 The device includes a data acquisition module 410, a threshold calculation module 430, and an incremental update module 450.

[0146] The data acquisition module 410 is used to acquire existing time series prediction old models, training datasets, and incremental datasets.

[0147] The threshold calculation module 430 is used to calculate the updated threshold based on the old model and the training dataset.

[0148] The incremental update module 450 is used to initialize the current new model with the old model and perform incremental update steps on the current new model.

[0149] The incremental update step includes: sequentially selecting one sample from the incremental dataset as the current sample; using the current new model to make a prediction for the current sample to calculate the predicted loss value; comparing the predicted loss value with the update threshold; if the predicted loss value is greater than or equal to the update threshold, then using the current sample to perform an incremental update on the current new model to obtain the updated new model, which will be used as the current new model for the next incremental update.

[0150] The incremental update module 450 repeatedly executes the incremental update step until all samples in the incremental dataset have been traversed, and the updated new model obtained from the last incremental update step is used as the incremental update model of the old model.

[0151] In some embodiments, the time series forecasting model incremental update apparatus further includes a model evaluation module for evaluating the incremental update model.

[0152] Specifically, the data acquisition module 410 is also used to acquire the test dataset. The model evaluation module is used to: take the union of the training dataset and the incremental dataset as the full dataset, and use the full dataset to train the old model to obtain the full-update model; for the test dataset, make predictions using the incremental update model and the full-update model respectively to obtain the first test result data of the incremental update model and the second test result data of the full-update model; and obtain the first evaluation result of the incremental update model by comparing the second test result data with the first test result data.

[0153] The model evaluation module can also be used to: calculate the training time of the incremental update model and the training time of the full update model; calculate the ratio of the training time of the full update model to the training time of the incremental update model to obtain the second evaluation result of the incremental update model.

[0154] It should be understood that the specific features, operations, and details described above with respect to the method of this application can also be similarly applied to the apparatus of this application, or vice versa. Furthermore, each step of the method of this application described above can be performed by a corresponding component or unit of the apparatus or system of this application.

[0155] It should be understood that the various modules / units of the device of this application can be implemented wholly or partially through software, hardware, firmware, or a combination thereof. Each module / unit can be embedded in the processor of the electronic device in hardware or firmware form or independent of the processor, or it can be stored in the memory of the electronic device in software form for the processor to call to execute the operation of each module / unit. Each module / unit can be implemented as an independent component or module, or two or more modules / units can be implemented as a single component or module.

[0156] like Figure 5 As shown, this application provides an electronic device 500, which includes a processor 501 and a memory 502 storing computer program instructions. The processor 501 executes the computer program instructions to implement the steps of the above-described incremental update method for time series prediction models. This electronic device 500 can be broadly categorized as a server, terminal, or any other electronic device with the necessary computing and / or processing capabilities.

[0157] In one embodiment, the electronic device 500 may include a processor, memory, network interface, communication interface, etc., connected via a system bus. The processor of the electronic device 500 can be used to provide necessary computing, processing, and / or control capabilities. The memory of the electronic device 500 may include non-volatile storage media and internal memory. The non-volatile storage media may store an operating system, computer programs, etc. The internal memory can provide an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface and communication interface of the electronic device 500 can be used to connect and communicate with external devices via a network. When the computer program is executed by the processor, it performs the steps of the method of this application.

[0158] In addition, this application provides a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the above-described incremental update method for the time series prediction model.

[0159] Those skilled in the art will understand that the method steps of this application can be performed by a computer program instructing related hardware, such as electronic device 500 or a processor. The computer program can be stored in a non-transitory computer-readable storage medium, and its execution causes the steps of this application to be performed. Depending on the context, any reference herein to memory, storage, or other media may include non-volatile or volatile memory. Examples of non-volatile memory include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid-state drive, etc. Examples of volatile memory include random access memory (RAM), external cache memory, etc.

[0160] The technical features described above can be combined arbitrarily. Although not all possible combinations of these technical features are described, any combination of these technical features should be considered to be covered by this specification, provided that such combination does not contain contradictions.

[0161] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. An incremental update method for a time series forecasting model, characterized in that, include: Obtain existing time series prediction models, training datasets, and incremental datasets; Calculate the update threshold based on the old model and the training dataset; Using the old model as the initialization point, the following incremental update steps are performed on the current new model: One sample from the incremental dataset is selected as the current sample in sequence. For the current sample, the new model is used to make a prediction to calculate the prediction loss value. Compare the predicted loss value with the updated threshold; If the predicted loss value is greater than or equal to the update threshold, the current sample is used to incrementally update the current new model to obtain the updated new model, which is then used as the current new model for the next incremental update. Repeat the above incremental update steps until all samples in the incremental dataset have been traversed. Use the updated model obtained from the last incremental update step as the incremental update model of the old model.

2. The method according to claim 1, characterized in that, The step of using the new model to make predictions for the current sample and calculating the prediction loss value includes: For the current sample, forward propagation is performed using the current new model to obtain the new model's predicted value; The predicted loss value is obtained by calculating the degree of difference between the predicted value of the new model and the true value corresponding to the current sample.

3. The method according to claim 2, characterized in that, The incremental update of the current new model using the current samples includes: The old model is used to perform forward propagation on the current sample to obtain the old model's predicted value; Calculate the distillation loss value based on the old model prediction value and the new model prediction value; Calculate the distillation loss function based on the predicted loss value and the distillation loss value; Backpropagation is performed on the current new model based on the distillation loss function to update the model parameters of the current new model.

4. The method according to claim 3, characterized in that, The formula for calculating the distillation loss function is as follows: loss=loss new +p*loss dist ; Wherein, the loss new Let loss be the predicted loss value. dist Let p be the distillation loss value, and p be the distillation coefficient, where p ∈ [0.05, 0.2].

5. The method according to claim 1, characterized in that, The calculation of the update threshold based on the old model and the training dataset includes: The old model is used to make predictions based on the training dataset to calculate the loss value; Calculate the mean and standard deviation of all loss values; The updated threshold is calculated based on the average value and the standard deviation.

6. The method according to claim 5, characterized in that, The step of calculating the update threshold based on the average value and the standard deviation includes calculating the update threshold using the following formula: K = μ + α·σ; Where K is the update threshold, μ is the average value, σ is the standard deviation, and α is the hyperparameter, and α∈[0.3, 0.7].

7. The method according to any one of claims 1-6, characterized in that, The method further includes: performing model evaluation on the incremental update model.

8. The method according to claim 7, characterized in that, The step of evaluating the incremental update model includes: The union of the training dataset and the incremental dataset is taken as the full dataset, and the old model is trained using the full dataset to obtain the full update model. Obtain a test dataset, and for the test dataset, make predictions using the incremental update model and the full update model respectively to obtain the first test result data of the incremental update model and the second test result data of the full update model; The first evaluation result of the incremental update model is obtained by comparing the second test result data with the first test result data.

9. The method according to claim 8, characterized in that, The process of evaluating the incremental update model further includes: The training time of the incremental update model and the training time of the full update model are statistically analyzed. The ratio of the training time of the full update model to the training time of the incremental update model is calculated to obtain the second evaluation result of the incremental update model.

10. An incremental update device for a time series forecasting model, characterized in that, The device includes: The data acquisition module is used to acquire existing time series prediction models, training datasets, and incremental datasets. The threshold calculation module is used to calculate the updated threshold based on the old model and the training dataset; The incremental update module is used to perform an incremental update step on the current new model, which is initialized with the old model. The incremental update step includes: One sample from the incremental dataset is selected sequentially as the current sample. For the current sample, the current new model is used to make a prediction to calculate the prediction loss value. The prediction loss value is compared with the update threshold. If the prediction loss value is greater than or equal to the update threshold, the current sample is used to incrementally update the current new model to obtain the updated new model, which is used as the current new model for the next incremental update. The incremental update module repeatedly executes the incremental update step until all samples in the incremental dataset have been traversed, and the updated model obtained from the last incremental update step is used as the incremental update model of the old model.

11. An electronic device, characterized in that, include: Processor and memory storing computer program instructions; When the processor executes the computer program instructions, it implements the steps of the method according to any one of claims 1-9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 9.