Semi-real-time predictive model updating method based on parallel verification of double models

CN122114687BActive Publication Date: 2026-09-22GUIZHOU YUNDUAN HUIHONG TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610578888.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-29
Publication Date
2026-09-22
Estimated Expiration
2046-04-29

AI Technical Summary

Technical Problem

[0004]为了解决上述技术问题,提供基于双模型并行验证的半实时预测模型更新方法,以解决现有的问题

Benefits of technology

本申请构建半实时双模型,其有益效果在于通过模型A进行实时预测,保证预测的连续性,而模型B在后台进行探索性训练,这种架构既利用了深度学习模型的强大拟合能力,又通过双模型机制平衡了系统的稳定性和灵活性;分别获取预测负荷数据和预测验证数据,其有益效果在于通过并行预测,同步获得了同一批数据在两个模型下的输出结果,以便后续评估两个模型的预测性能差异;计算表征模型A与模型B全局预测性能差异的整体偏差量,其有益效果在于反映了两个模型的预测性能差异性,并以此作为判断依据,评估是否要对模型A进行参数更新;对该周期进行多尺度划分,其有益效果在于通过多尺度时间窗口划分,能够同时捕捉电力负荷数据的长时趋势和短时突变特征;计算表征负荷波动剧烈性的第一评估值,其有益效果在于量化了负荷波动的剧烈程度,使得模型在训练时能够关注那些变化剧烈、难以预测的关键时段,保证了模型对物理世界变化的敏感性和跟随性,初步评估该时间窗口的数据对模型训练的重要性;确定第二评估值,其有益效果在于通过对比局部误差与全局误差的差异,识别出模型未能学习到哪些时间窗口的数据特征,以进一步评估该时间窗口的数据对模型训练的重要性;得到各时间窗口的重要性权重,其有益效果在于通过单个时间窗口下的数据波动以及模型预测偏差两个维度,综合评估该时间窗口的数据特征的重要性,以为后续的加权训练提供了精准的指引,使得模型训练更加聚焦于那些既波动剧烈又预测不准的关键数据,从而高效提升预测精度;选择并执行不同的训练更新策略,对模型B进行增量训练并选择性更新模型A,其中,重要性权重在增量训练过程中参与损失计算,其有益效果在于当整体偏差量较大时,利用加权的样本数据对模型B进行针对性增量训练并将其模型参数同步给A,保证预测精度,在整体偏差量较小时,仅微调模型B而不更新模型A的参数,既保持了模型B对数据的敏感度,又维持了模型A的稳定性,防止因过度更新导致的模型震荡或过拟合,同时,通过重要性权重在进行增量训练时调整损失函数,引导模型B能够关注高权重样本的数据特征,以快速响应负荷模式突变的敏捷性,提升了复杂工况下对电力负荷的预测精度和模型自适应能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122114687B_ABST
    Figure CN122114687B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of load prediction, in particular to a semi-real-time prediction model updating method based on double-model parallel verification, which comprises the following steps: constructing a semi-real-time double model, including model A for predicting power load and model B for background incremental training; performing parallel prediction on load data of each monitoring period to obtain predicted load data and prediction verification data; calculating a first evaluation value representing load fluctuation intensity based on real load data in each monitoring period; calculating an overall deviation amount; determining a second evaluation value to obtain an importance weight; selecting and executing different training updating strategies to perform incremental training on model B and selectively update model A, wherein the importance weight participates in loss calculation during the incremental training process. The application improves the prediction accuracy of power load and the model self-adaptive ability under complex working conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of load forecasting technology, specifically to a semi-real-time forecasting model update method based on dual-model parallel verification. Background Technology

[0002] With the development of new power systems, power load is affected by multiple factors such as fluctuations in new energy output, extreme weather, and electricity pricing policies, exhibiting complex characteristics such as strong nonlinearity, time-varying nature, and randomness. Power load forecasting is a key technology for ensuring the safe and stable operation of the power grid and realizing economic dispatch and demand-side management.

[0003] Traditional load forecasting models typically employ offline training and periodic updates based on historical load data from the power grid to predict power load. However, due to the high randomness of power grid operation and large load fluctuations, this offline training method struggles to cope with sudden load fluctuations, failing to capture local feature changes in a timely manner, leading to model performance degradation and decreased prediction accuracy. Meanwhile, while full model updates based on real-time data can improve model adaptability, they consume enormous computational resources, incur high operating costs, and the training process may cause prediction service interruptions. This makes it difficult to meet the engineering requirements of real-time power system dispatching, and fails to reduce computational costs and improve the model's response speed to sudden fluctuations while ensuring prediction accuracy. Summary of the Invention

[0004] To address the aforementioned technical issues, a semi-real-time prediction model update method based on dual-model parallel verification is provided to resolve existing problems.

[0005] The solution to the technical problem in this application is to provide a semi-real-time prediction model update method based on dual-model parallel verification, including the following steps: A semi-real-time dual model is constructed, including Model A for predicting power load and Model B for incremental training in the background. Model A and Model B are deep learning models with the same structure. The dual model is used to perform parallel prediction of load data in each monitoring period to obtain predicted load data and prediction verification data respectively. After each monitoring cycle ends, based on the actual load data, the following steps are performed: Analyze the differences between the predicted load data and the predicted verification data and the actual load data throughout the entire monitoring period, and calculate the overall deviation that characterizes the difference in global prediction performance between Model A and Model B. The entire monitoring period is divided into multiple scales. For each time window under each time scale, the fluctuation characteristics of the load data are analyzed, and the first assessment value characterizing the severity of load fluctuation is calculated. The deviations of the predicted load data and the predicted validation data from the actual load data under a single time window are evaluated to reflect the significance of local prediction performance differences relative to the global performance. A second evaluation value is determined, and the importance weight of each time window is obtained by combining the first evaluation value. Using the overall bias, different training and update strategies are selected and implemented to incrementally train model B and selectively update model A. The importance weights participate in the loss calculation during the incremental training process.

[0006] Preferably, the acquisition of predicted load data and predicted verification data includes: using model A to predict the load data for each monitoring period to obtain predicted load data; and using model B to predict the load data for each monitoring period to output predicted verification data.

[0007] Preferably, the calculation process for the overall deviation is as follows: For the actual load data under each monitoring period, the prediction error between the actual load data and the predicted load data is calculated as the first total error; the prediction error between the actual load data and the predicted verification data is calculated as the second total error. The overall deviation is the difference between the first total error and the second total error.

[0008] Preferably, the calculation process for the first evaluation value is as follows: For each time scale of each monitoring cycle, the difference between the load data of two adjacent moments within each time window is taken as the relative difference. The range of all relative differences within each time window is calculated, and the difference between the range of each time window and its adjacent time windows is taken as the fluctuation difference of each time window. The first assessment value is positively correlated with both the range and the fluctuation difference.

[0009] Preferably, determining the second evaluation value includes: For each monitoring period, the prediction error between the predicted load data and the actual load data within each time window at each time scale is denoted as the first local error; the prediction error between the predicted verification data and the actual load data within each time window at each time scale is denoted as the second local error; the difference between the first local error and the second local error is calculated as the local deviation. If the local deviation is greater than or equal to the overall deviation, the second evaluation value for each time window is the maximum of the first local error and the second local error; otherwise, the second evaluation value is the minimum of the first local error and the second local error.

[0010] Preferably, the importance weight is positively correlated with both the first evaluation value and the second evaluation value.

[0011] Preferably, the selection and execution of different training update strategies includes: if the overall deviation is greater than or equal to a preset threshold, then a model update strategy is executed; otherwise, a conventional maintenance strategy is executed.

[0012] Preferably, a regular incremental set and a weighted incremental set are constructed respectively. The model update strategy is as follows: Model B is incrementally trained through the weighted incremental set, and the model parameters of Model B after incremental training are synchronized to Model A to update the model parameters of Model A. The regular maintenance strategy is as follows: Model B is incrementally trained through the regular incremental set, but the model parameters of Model A are not updated.

[0013] Preferably, the process of obtaining the regular incremental set is as follows: the load data of a single monitoring period is defined as a regular sample, and the regular samples corresponding to each monitoring period and the previous multiple monitoring periods are used to form a regular incremental set; wherein, the importance weight of the regular sample is a preset value.

[0014] Preferably, the process of obtaining the weighted incremental set is as follows: defining the load data within each time window as an independent sample; for each monitoring period, constructing a weighted incremental set by including all independent samples at all time scales and multiple previous regular samples.

[0015] This application has at least the following beneficial effects: This application constructs a semi-real-time dual-model architecture. The advantage lies in using Model A for real-time prediction, ensuring prediction continuity, while Model B undergoes exploratory training in the background. This architecture leverages the powerful fitting capabilities of deep learning models and balances system stability and flexibility through the dual-model mechanism. Separately acquiring prediction load data and prediction validation data allows for parallel prediction, simultaneously obtaining the output results of the same batch of data under both models, facilitating subsequent evaluation of the predictive performance differences between the two models. Calculating the overall deviation representing the global predictive performance difference between Model A and Model B reflects the predictive performance differences between the two models and provides a basis for further evaluation. As a basis for judgment, it is assessed whether to update the parameters of model A; the period is divided into multiple scales, which has the benefit of capturing both long-term trends and short-term abrupt changes in power load data through multi-scale time window division; the first evaluation value characterizing the severity of load fluctuations is calculated, which has the benefit of quantifying the severity of load fluctuations, enabling the model to focus on critical periods of drastic and unpredictable changes during training, ensuring the model's sensitivity and following of changes in the physical world, and initially assessing the importance of the data in this time window for model training; the second evaluation value is determined, which has the benefit of identifying the difference between local and global errors. The model fails to learn the data features of which time windows, allowing for further evaluation of the importance of data in those time windows for model training. Importance weights for each time window are obtained, which provides a comprehensive assessment of the importance of data features within a single time window through two dimensions: data volatility and model prediction bias. This offers precise guidance for subsequent weighted training, enabling the model to focus more on key data that is both highly volatile and inaccurately predicted, thereby efficiently improving prediction accuracy. Different training and update strategies are selected and implemented, incrementally training model B and selectively updating model A. The importance weights participate in the loss during incremental training. The advantages of loss calculation are that when the overall deviation is large, it uses weighted sample data to perform targeted incremental training on model B and synchronizes its model parameters to model A to ensure prediction accuracy. When the overall deviation is small, it only fine-tunes model B without updating the parameters of model A, which maintains the sensitivity of model B to data and the stability of model A, preventing model oscillation or overfitting caused by excessive updates. At the same time, by adjusting the loss function during incremental training through importance weights, it guides model B to focus on the data characteristics of high-weight samples, so as to quickly respond to the agility of load pattern changes and improve the prediction accuracy and model adaptability of power load under complex operating conditions. Attached Figure Description

[0016] The semi-real-time prediction model update method based on dual-model parallel verification of this application will be further described in detail below with reference to the accompanying drawings.

[0017] Figure 1 A flowchart illustrating the steps of a semi-real-time prediction model update method based on dual-model parallel verification provided in an embodiment of this application; Figure 2 A flowchart illustrating the steps of a method for obtaining a first evaluation value provided in an embodiment of this application. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description of the semi-real-time prediction model update method based on dual-model parallel verification proposed in this application, in conjunction with the accompanying drawings and implementation examples, is provided. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit the scope of this application.

[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0020] Please see Figure 1 The diagram illustrates a flowchart of a semi-real-time prediction model update method based on dual-model parallel verification according to an embodiment of this application. The method includes the following steps: Step 1: Construct a semi-real-time dual model, including Model A for predicting power load and Model B for incremental training in the background. Model A and Model B are deep learning models with the same structure. Use the dual model to perform parallel prediction of load data for each monitoring period to obtain predicted load data and prediction verification data respectively.

[0021] As power systems become increasingly complex, accurate load forecasting is crucial for stable grid operation and optimized dispatch. LSTM, as a deep learning model capable of capturing long-term dependencies in time series data, has achieved remarkable results in time series performance tasks.

[0022] Based on the above analysis, a semi-real-time dual model for power load forecasting is constructed. The dual model includes model A and model B, wherein model A is used for online forecasting of power load, and model B is used for incremental training in the background. In this embodiment, both Model A and Model B are constructed using Long Short-Term Memory (LSTM) network models with the same structure. The optimizer is the Adam optimizer, and the loss function is the MES loss. It should be noted that the LSTM model is a well-known technology and will not be described in detail here.

[0023] The load data of the power grid during historical operation is collected, and the load data is filled with missing values ​​and normalized. Thus, multiple consecutive time periods are taken as a monitoring period. The load data under multiple monitoring periods are used to form a training set, and model A and model B in the semi-real-time dual model are trained respectively. In this embodiment, historical load data of the power grid over the past year, including voltage, current, and active power, are collected. The monitoring period is 30 minutes. As for other implementation methods, the implementer can set the time according to the power load forecasting requirements of the actual implementation scenario. The maximum-minimum normalization method is used for normalization, and the KNN algorithm is used for missing value filling. The maximum-minimum normalization method and the KNN algorithm are well-known technologies and will not be described in detail here.

[0024] Regarding the first Load data for each monitoring cycle Input it into the semi-real-time dual model, and use model A to analyze the first... The load data of the monitoring cycle is used to predict and obtain the first monitoring cycle. Forecast load data for each monitoring period Using Model B to analyze the first... The load data of the monitoring cycle is used to predict and output the first cycle. Predictive validation data for each monitoring period ; It should be noted that the parameters of model A remain frozen within a single monitoring period, and only online prediction tasks are performed. Model B continuously receives historical incremental data in the background for updates. When the two models perform parallel predictions, model B has a more up-to-date parameter state than model A.

[0025] At this point, we have obtained the predicted load data and the predicted verification data.

[0026] Step 2: Analyze the differences between the predicted load data and the predicted verification data and the actual load data throughout the entire monitoring period, and calculate the overall deviation that characterizes the difference in global prediction performance between Model A and Model B; based on the actual load data in each monitoring period, divide the period into multiple scales, and for each time window under each time scale, analyze the fluctuation characteristics of the load data, and calculate the first evaluation value characterizing the severity of load fluctuation.

[0027] In the After the first monitoring cycle ends, for the first Real load data under each monitoring cycle To evaluate the differences in the predictive performance of the models, the overall bias is calculated, specifically as follows: Computational model A in the th Forecast load data for each monitoring period Compared with actual load data The prediction error between them is taken as the first total error; Computational model B in the 1st Predictive validation data for each monitoring period Compared with actual load data The prediction error between them is taken as the second total error; In this embodiment, the prediction error is measured by calculating the mean absolute percentage error. The calculation of the mean absolute percentage error is a well-known technique and will not be described in detail here.

[0028] Calculate the difference between the first total error and the second total error, which is the overall deviation. If the overall deviation is greater than or equal to the preset threshold, incremental training is performed on model B, and the model parameters of model B after incremental training are synchronized to model A to update the model parameters of model A. Otherwise, incremental training is performed on model B, but the model parameters of model A are not updated. In this embodiment, the preset threshold value is 2%. In other implementation methods, the implementer can set it according to the actual situation.

[0029] Using the first Real load data under each monitoring cycle Incremental training is performed on model B, and the specific process is as follows.

[0030] First, the integration of new energy sources into the grid has led to increasingly frequent abnormal fluctuations in the power system. Traditional power load forecasting models based on fixed parameters suffer from performance degradation when encountering significant power system fluctuations. Updating these models based on real-time load data requires substantial computing resources, resulting in high operating costs. In contrast, semi-real-time updated models, trained on a large amount of historical load data, use a small amount of real-time load data for incremental background training. Furthermore, they do not update the model parameters unless absolutely necessary. Updates are only triggered when significant power system fluctuations cause a decline in actual forecasting performance. This approach ensures the accuracy of the forecasting model while reducing the computing resources required for its operation, without affecting real-time prediction.

[0031] When the model parameter update strategy is triggered, it indicates that the load data of the power system has undergone significant abnormal changes, and the trend features learned by Model A are insufficient to adapt to the abnormal load fluctuations at this time. Therefore, incremental training of Model B is required based on real-time load data. However, if all data are trained with equal importance during incremental training of Model B, the occasional abnormal fluctuation features will be masked by a large number of other steady-state features, making it difficult for the model to learn them fully. Even after incremental training, it will still be difficult to grasp the distribution and change characteristics of the load data, resulting in the model's predictive performance not being effectively improved.

[0032] Fluctuations in power grid load data are mostly sporadic spikes and anomalies, such as sudden voltage drops, sudden current surges, harmonic exceedances, and sudden events like new energy grid connection. These fluctuations do not exhibit a fixed trend in time sequence and can occur at any moment. In power load forecasting, these sporadic anomalies typically occur over short periods relative to a typical monitoring cycle, and their characteristics are easily masked by large amounts of load data, affecting the model's learning performance. Therefore, the entire monitoring cycle is divided into time windows of different time scales to analyze the fluctuations in load data within short periods and calculate a first evaluation value. The flowchart of the method for obtaining the first evaluation value provided in this application embodiment is shown below. Figure 2 As shown, it specifically includes: Based on multiple time scales, all moments within each monitoring period are divided into multiple scales to obtain multiple time windows under each time scale. In this embodiment, each monitoring cycle is divided into 30 consecutive time windows with a time scale of 1 minute; and each monitoring cycle is divided into 10 consecutive time windows with a time scale of 3 minutes. This embodiment does not impose any special restrictions on this.

[0033] For each monitoring period, the difference between two adjacent moments within each time window at each time scale is taken as the relative difference; In this embodiment, the difference between the load data at each time point within each time window and the previous time point is taken as the relative difference.

[0034] For each time scale, the range of all relative differences within each time window is calculated, and the difference in the range between each time window and its adjacent time windows is taken as the fluctuation difference of each time window. It should be noted that the calculation of the range is a well-known technique and will not be elaborated upon here.

[0035] In this embodiment, the calculation process of fluctuation difference is as follows: calculate the absolute value of the difference between the range of each time window and the previous time window, as the first difference; calculate the absolute value of the difference between the range of each time window and the next time window, as the second difference; and take the average of the first difference and the second difference as the fluctuation difference. It should be noted that for the first time window, the fluctuation difference is the second difference, and for the last time window, the fluctuation difference is the first difference.

[0036] It should be noted that the relative difference reflects the intensity of load data fluctuations between two moments. In a normally operating power grid, load data changes between adjacent moments are relatively stable within the second-level data range. However, when a voltage drop or current surge occurs, this stable operating state of the power grid is disrupted. Therefore, the larger the range, the more drastic the fluctuations in the load data of the power grid within that time window, and the higher the importance of the data features within that time window. This is key to improving the model's ability to learn fluctuation features and make accurate predictions. The larger the fluctuation difference, the more significant the difference in operating state between that time window and adjacent time windows. There is only one significant fluctuation change within that time window, and its duration is relatively short compared to the entire monitoring cycle. The model has difficulty learning the features of data with a small amount of data, so the data within that time window is more important.

[0037] The first evaluation value of each time window at each time scale is positively correlated with the range and the fluctuation difference. It should be noted that a positive correlation means that the dependent variable increases as the independent variable increases, and decreases as the independent variable decreases.

[0038] In this embodiment, the calculation process of the first evaluation value is as follows: based on the preset first weight and the preset second weight, the range and fluctuation difference are weighted and summed to obtain the first evaluation value, wherein the sum of the preset first weight and the preset second weight is 1. In this embodiment, the preset first weight is 0.6 and the preset second weight is 0.4. As for other implementation methods, the implementer can set them according to the actual situation.

[0039] It should be noted that the larger the first evaluation value, the greater the abnormal change in the load data within that time window and the shorter the duration. Such abnormal changes within a short period of time are relatively important for the model to learn abnormal features during training.

[0040] Thus, the first evaluation value for each time window at each time scale is obtained.

[0041] Step 3: Evaluate the deviation of the predicted load data and the predicted verification data from the actual load data within a single time window, reflecting the significance of the local prediction performance difference relative to the global difference, and determine the second evaluation value.

[0042] Furthermore, the first evaluation value reflects the degree of abnormal fluctuation in real-time load data, representing the importance of a data segment from the perspective of the power grid's operating conditions. However, this feature may represent a common fluctuation state in the power grid, and its variation characteristics have already been learned by the model through the training set. That is, it can predict the load data under abnormal fluctuations of this operating condition very well in the prediction results. Therefore, in order to avoid the model learning too much about familiar operating condition characteristics, which would prevent the model from effectively learning other less common abnormal operating condition characteristics, a second evaluation value is calculated by analyzing the difference between the actual load data and the model's prediction results. Specifically: For each monitoring period, the prediction error between the predicted load data and the actual load data within each time window at each time scale is denoted as the first local error; the prediction error between the predicted verification data and the actual load data within each time window at each time scale is denoted as the second local error. In this embodiment, the prediction error is measured by calculating the mean absolute percentage error. The calculation of the mean absolute percentage error is a well-known technique and will not be described in detail here.

[0043] Calculate the difference between the first local error and the second local error, and use it as the local deviation. If the local deviation is greater than or equal to the overall deviation, then the second evaluation value for each time window at each time scale is the maximum of the first local error and the second local error; otherwise, the second evaluation value is the minimum of the first local error and the second local error. In this embodiment, the calculation process for the second evaluation value is as follows:

[0044] in, For the first The first time scale The second evaluation value for each time window, For the first The first time scale The first local error within a time window For the first The first time scale The second local error within a time window For the first The first time scale Local deviation within a time window For the first The overall deviation over a monitoring period This represents the function that takes the maximum value. This represents the function that takes the minimum value; where, ; , For the first The first total error of the monitoring cycle For the first The second total error of the monitoring cycle; It should be noted that when the local deviation is greater than or equal to the overall deviation, it indicates that the prediction error of that time window is significantly greater than the overall error. This reflects that the load data in that time window exhibits non-stationary, drastic fluctuations or novel distribution patterns, which may be features that the model has never learned before, making effective load prediction impossible. In this case, the data in that time window is of high importance. Therefore, the second evaluation value is selected as the maximum of the first and second local errors to increase the weight of that time window during training, forcing the model to focus on learning this anomalous feature. Conversely, if the local deviation is less than the overall error, it indicates that the load data in that time window belongs to normal operating conditions, which the model can predict well. In this case, the data in that time window is of low importance. Therefore, the second evaluation value is selected as a smaller value to avoid the model over-focusing on normal data and causing overfitting.

[0045] Thus, the second evaluation value for each time window at each time scale is obtained.

[0046] Step 4: Based on the first and second evaluation values, obtain the importance weights for each time window; using the overall deviation, select and execute different training and update strategies to incrementally train model B and selectively update model A, wherein the importance weights participate in the loss calculation during the incremental training process.

[0047] Furthermore, based on the first and second evaluation values, the importance weights are determined as follows: The importance weight of each time window at each time scale is positively correlated with both the first and second evaluation values. In this embodiment, the importance weight is calculated as follows: the first evaluation value and the second evaluation value are normalized respectively, the mean of the normalized first evaluation value and the normalized second evaluation value is calculated, and the sum of the mean and the value of 1 is used as the importance weight of each time window under each time scale; secondly, the maximum and minimum value normalization method is used to normalize the first evaluation value and the second evaluation value of all time windows under the same time scale. The maximum and minimum value normalization method is a well-known technique and will not be described in detail here.

[0048] It should be noted that the greater the importance weight, the more important the data features in that time window are, and the more likely they are features that the model has never learned before.

[0049] The load data of a single monitoring period is defined as a regular sample. The regular samples corresponding to each monitoring period and the previous multiple monitoring periods are combined to form a regular incremental set. The importance weight of the regular sample is a preset value. In this embodiment, the preset value is 1. In other implementation methods, the implementer can set it according to the actual situation. Secondly, the first... The load data from one monitoring cycle and the previous six monitoring cycles constitute a regular incremental set, which can be set by the implementer according to the actual situation as another implementation method.

[0050] Define the load data within each time window as an independent sample; For each monitoring period, a weighted incremental set is formed by all independent samples at all time scales and multiple regular samples prior to them; In this embodiment, the number of regular samples in the weighted incremental set is 6. As for other implementation methods, the implementer can set it according to the actual situation.

[0051] It should be noted that the importance weight of independent samples in the weighted incremental set is greater than the importance weight of regular samples.

[0052] If the overall deviation is greater than or equal to a preset threshold, a model update strategy is executed; otherwise, a regular maintenance strategy is executed. The preset threshold is 2%. The model update strategy is as follows: Incrementally train model B using a weighted incremental set, and then synchronize the model parameters of model B after incremental training to model A to update the model parameters of model A. The standard maintenance strategy is to incrementally train model B using a standard incremental set, but without updating the model parameters of model A; during the training process of model B, each sample participates in the loss calculation according to its importance weight. It should be noted that during the training process of Model B, the loss weight of each sample is its importance weight. By increasing the proportion of data features in the more important time window in the total loss, Model B can focus on learning the data features in that time window.

[0053] It should be understood that, although Figure 1 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 1 At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.

[0054] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0055] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application. Therefore, any simple modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of this application, without departing from the content of the technical solution of this application, shall fall within the protection scope of the technical solution of this application.

Claims

1. A semi-real-time prediction model update method based on dual-model parallel validation, characterized in that, The method includes the following steps: A semi-real-time dual model is constructed, including Model A for predicting power load and Model B for incremental training in the background. Model A and Model B are deep learning models with the same structure but independently maintained parameters. The dual model is used to perform parallel prediction of load data in each monitoring period to obtain predicted load data and prediction verification data respectively. After each monitoring cycle ends, based on the actual load data, the following steps are performed: Analyze the differences between the predicted load data and the predicted verification data and the actual load data throughout the entire monitoring period, and calculate the overall deviation that characterizes the difference in global prediction performance between Model A and Model B. The entire monitoring period is divided into multiple scales. For each time window under each time scale, the fluctuation characteristics of the load data are analyzed, and the first assessment value characterizing the severity of load fluctuation is calculated. The deviations of the predicted load data and the predicted validation data from the actual load data under a single time window are evaluated to reflect the significance of the local prediction performance difference relative to the global prediction performance difference. A second evaluation value is determined, and the importance weight of each time window is obtained by combining the first evaluation value. Using the overall bias, different training and update strategies are selected and executed to incrementally train model B and selectively update model A. The importance weights participate in the loss calculation during the incremental training process. Specifically, for each monitoring period, the prediction error between the predicted load data and the actual load data within each time window at each time scale is denoted as the first local error; the prediction error between the predicted verification data and the actual load data within each time window at each time scale is denoted as the second local error; the difference between the first local error and the second local error is calculated as the local deviation; if the local deviation is greater than or equal to the overall deviation, the second evaluation value for each time window is the maximum value between the first local error and the second local error; otherwise, the second evaluation value is the minimum value between the first local error and the second local error. The selection and execution of different training update strategies include: if the overall deviation is greater than or equal to a preset threshold, then the model update strategy is executed; otherwise, the regular maintenance strategy is executed. Specifically, a regular incremental set and a weighted incremental set are constructed respectively. The model update strategy is as follows: Model B is incrementally trained using the weighted incremental set, and the model parameters of Model B after incremental training are synchronized to Model A to update the model parameters of Model A. The regular maintenance strategy is as follows: Model B is incrementally trained using the regular incremental set, but the model parameters of Model A are not updated. Among them, the load data of a single monitoring period is defined as a regular sample, and the regular samples corresponding to each monitoring period and the previous multiple monitoring periods are used to form a regular incremental set. The importance weight of the regular sample is a preset value. The load data within each time window is defined as an independent sample; for each monitoring period, all independent samples at all time scales, along with multiple previous regular samples, are combined to form a weighted incremental set.

2. The semi-real-time prediction model update method based on dual-model parallel verification as described in claim 1, characterized in that, The acquisition of predicted load data and predicted verification data includes: using model A to predict the load data for each monitoring period and acquiring predicted load data; using model B to predict the load data for each monitoring period and outputting predicted verification data.

3. The semi-real-time prediction model update method based on dual-model parallel verification as described in claim 1, characterized in that, The calculation process for the overall deviation is as follows: For the actual load data under each monitoring period, the prediction error between the actual load data and the predicted load data is calculated as the first total error; the prediction error between the actual load data and the predicted verification data is calculated as the second total error. The overall deviation is the difference between the first total error and the second total error.

4. The semi-real-time prediction model update method based on dual-model parallel verification as described in claim 1, characterized in that, The calculation process for the first evaluation value is as follows: For each time scale of each monitoring cycle, the difference between the load data of two adjacent moments within each time window is taken as the relative difference. The range of all relative differences within each time window is calculated, and the difference between the range of each time window and its adjacent time windows is taken as the fluctuation difference of each time window. The first assessment value is positively correlated with both the range and the fluctuation difference.

5. The semi-real-time prediction model update method based on dual-model parallel verification as described in claim 1, characterized in that, The importance weight is positively correlated with both the first and second evaluation values.

Citation Information

Patent Citations

  • Electricity market user load prediction method, system, equipment and medium

    CN119940665A

  • Electricity consumption prediction method

    CN121258730A