A power grid time-varying load ultra-short-term prediction method and system
Patent Information
- Application Number
- CN202611096435.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-23
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-07-23
AI Technical Summary
但电网负荷在进行超短期预测时,大多依赖人工选定相关影响因素,难以随用电结构的变化自动调整所考察的指标
[0016] The present invention provides a method and system for ultra-short-term forecasting of time-varying loads in power grids. By decomposing historical load data, it separates the base load component and the meteorologically sensitive load component, allowing for differentiated treatment of load components with different characteristics. Based on this, it calculates the time-delay mutual information between the meteorologically sensitive load component and various meteorological factors, automatically determining the effective lag duration of each meteorological factor and performing time-shift alignment. This accurately captures the lag effect of meteorological conditions on load, avoiding the blindness of manually setting lag durations, thus providing more precise input features for the forecasting model. Furthermore, the time-shift aligned meteorological factors, base load component, and non-meteorological factor data are used together as input features to construct a training sample set and pre-train the model. This enables the model to comprehensively utilize multi-source heterogeneous information, considering both the inherent trend of load changes and the dynamic influence of external related factors, effectively improving the baseline accuracy of the initial forecasting model.
Smart Images

Figure CN122620429B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power forecasting technology, and more specifically, to a method and system for ultra-short-term forecasting of time-varying loads in power grids. Background Technology
[0002] Currently, ultra-short-term load forecasting is a crucial foundation for real-time dispatching and security control of power systems. Its forecasting results directly serve core aspects such as automatic generation control, dynamic economic dispatch, and early warning of safety and stability. In real-world operating scenarios, grid load is influenced by multiple factors, including meteorological conditions, event disturbances, and electricity price signals, exhibiting strong time-varying and nonlinear fluctuation characteristics. Therefore, a systematic design is needed to enhance the dynamic tracking capability and adaptive update mechanism of load forecasting methods, enabling them to achieve high-precision, fast-response ultra-short-term power forecasting in complex and ever-changing electricity consumption environments.
[0003] In related technologies, data-driven forecasting methods are widely used in load forecasting due to their ability to uncover hidden patterns in historical data. These methods predict future power output by analyzing load sequences and their related factors. However, when making ultra-short-term load forecasts, most methods rely on manually selecting relevant influencing factors, making it difficult to automatically adjust the indicators considered in response to changes in electricity consumption patterns. For example, when faced with new situations such as temporary sporting events, staggered production, or sudden power outages, these methods often fail to promptly identify new related factors. As a result, the selected indicators may contain redundant information or omit important information, making it difficult to adapt to diverse changes in operating conditions. Summary of the Invention
[0004] The problem addressed by this invention is how to improve the accuracy of forecasting ultra-short-term loads of the power grid.
[0005] To address the above problems, this invention provides a method and system for ultra-short-term forecasting of time-varying loads in power grids.
[0006] In a first aspect, the ultra-short-term forecasting method for time-varying loads of the power grid of the present invention includes: Historical load data and corresponding related factor data are acquired, and the historical load data is decomposed into components to obtain the basic load component and the meteorological sensitive load component. The related factor data includes meteorological factor data and non-meteorological factor data. Based on the time delay mutual information calculation of the meteorological sensitive load component and each meteorological factor data, the effective lag time corresponding to each meteorological factor data is determined. Based on the effective lag time, the corresponding meteorological factor data is time-shifted and aligned. Each time-shifted meteorological factor data, the base load component, and the non-meteorological factor data are used as input features. The total load in the historical load data is used as the output target. A training sample set is constructed, and the initial prediction model is pre-trained using the training sample set to obtain the trained prediction model. In the rolling forecasting process based on the forecasting model, the relevant factor data of the current forecasting period is input into the forecasting model to obtain the load forecast value of the current forecasting period and determine the forecasting error. The cumulative forecasting error sequence of multiple consecutive forecasting periods is used as a feedback signal to determine the correlation coefficient between each non-meteorological factor data and the forecasting error sequence. The feature set of the initial forecasting model is updated according to the correlation coefficient. The updated feature set is used to fine-tune the parameters of the prediction model, and the correlation factor data for the next prediction period is input into the fine-tuned prediction model to output the load prediction value for the next prediction period.
[0007] Optionally, the step of acquiring historical load data and corresponding related factor data, and performing component decomposition on the historical load data to obtain the basic load component and the meteorologically sensitive load component, includes: The historical load data is decomposed into time series to extract the trend component and periodic component of the load. The trend component and the periodic component are superimposed to obtain the basic load component. Based on the historical load data and the basic load components, the remaining load sequence is obtained; Based on the correlation coefficient between the remaining load sequence and each meteorological factor data, meteorological factor data with a correlation coefficient greater than a preset threshold are selected as sensitive meteorological factors; The remaining load sequence is regressed and fitted to the sensitive meteorological factors to obtain the meteorological sensitive load components.
[0008] Optionally, the step of calculating the time delay mutual information based on the meteorological sensitive load component and each meteorological factor data to determine the effective lag time corresponding to each meteorological factor data includes: The sequence composed of each meteorological factor data and the meteorological sensitive load component is used as the sequence to be analyzed. A preset number of candidate lag durations are set, and the meteorological factor data is time-shifted according to the candidate lag durations to obtain the time-shifted sequence corresponding to each candidate lag duration; For each time-shifted sequence, the mutual information value between the time-shifted sequence and the meteorological sensitive load component is determined, and a mutual information sequence is obtained based on the mutual information value corresponding to each candidate lag duration; The candidate lag duration corresponding to the maximum value of the mutual information value in the mutual information sequence is determined as the effective lag duration corresponding to the meteorological factor data.
[0009] Optionally, for each time-shifted sequence, determining the mutual information value between the time-shifted sequence and the meteorologically sensitive load component, and obtaining a mutual information sequence based on the mutual information value corresponding to each candidate lag duration, includes: For each time-shifted sequence, joint probability density estimation and marginal probability density estimation are performed on the time-shifted sequence and the meteorological sensitive load component, and the mutual information value corresponding to the time-shifted sequence is determined based on the results of the joint probability density estimation and the results of the marginal probability density estimation. The mutual information values of each meteorological factor data under all candidate lag times are arranged in order of lag time to obtain the initial mutual information sequence corresponding to the meteorological factor data, and the initial mutual information sequence is smoothed to obtain a smoothed mutual information sequence. In the smooth mutual information sequence, the search proceeds from the zero lag position in the direction of increasing lag duration to determine the lag duration corresponding to the first peak, and it is determined whether there is a significant decrease after the first peak. If there is a significant decrease, the lag duration corresponding to the first peak is determined as the effective lag duration. The condition for determining a significant decrease is that the mutual information value at the k-th lag point after the first peak is lower than the product of the mutual information value of the first peak and a preset proportional coefficient, where k is a preset search step size.
[0010] Optionally, the step of time-shifting the corresponding meteorological factor data according to the effective lag duration, using each time-shifted meteorological factor data, the base load component, and the non-meteorological factor data as input features, and using the total load in the historical load data as the output target, constructs a training sample set, and uses the training sample set to pre-train the initial prediction model to obtain the trained prediction model, includes: For each historical moment, based on the effective lag time corresponding to each meteorological factor data, the original meteorological factor data of the earlier moment that is separated from the historical moment by the effective lag time is extracted, and used as the time-shifted aligned meteorological factor data of the historical moment; The time-shifted aligned meteorological factor data, the base load component of the historical time, and the non-meteorological factor data are combined to form the input feature vector of the historical time, and the total load of the historical time is used as the output target to form an input-output sample pair; The input-output sample pairs from all historical moments are aggregated to construct the training sample set; The initial prediction model is trained using the training sample set. The parameters of the initial prediction model are iteratively updated by minimizing the error between the load prediction value output by the initial prediction model and the total load until the preset training conditions are met, thus obtaining the trained prediction model.
[0011] Optionally, the step of training the initial prediction model using the training sample set, and iteratively updating the parameters of the initial prediction model by minimizing the error between the load prediction value output by the initial prediction model and the total load, until a preset training condition is met, to obtain the trained prediction model, includes: Each of the input feature vectors in the training sample set is input into the initial prediction model to obtain the corresponding load prediction value; Based on the predicted load value and the corresponding total load, the prediction residual is determined, and a loss function is constructed using the sum of squares of the prediction residuals; The parameters of the initial prediction model are backpropagated using the loss function to determine the gradient of each parameter, and the parameters are iteratively updated based on the gradient. After each iteration update, the loss value of the loss function under the current parameters is determined. When the change in the loss value between two adjacent iterations is less than the preset convergence threshold, it is determined that the preset training condition is met, and the prediction model under the current parameters is taken as the prediction model that has been trained.
[0012] Optionally, in the rolling forecasting process based on the forecasting model, the relevant factor data for the current forecast period are input into the forecasting model to obtain the load forecast value for the current forecast period and determine the forecast error. The cumulative forecast error sequence over multiple consecutive forecast periods is used as a feedback signal to determine the correlation coefficient between each non-meteorological factor data and the forecast error sequence. The feature set of the initial forecasting model is updated based on the correlation coefficient, including: In the rolling forecasting process, the predicted load value for each forecasting period is subtracted from the corresponding actual total load to obtain the forecasting error for that forecasting period. The forecasting errors of multiple consecutive forecasting periods are then combined in chronological order to form a forecasting error sequence. When the cumulative error of the prediction error sequence exceeds a preset threshold, the prediction error sequence is used as a feedback signal to trigger a feedback evaluation of the feature set of the initial prediction model. In the feedback evaluation, the correlation coefficient between each non-meteorological factor data and the prediction error sequence is determined, and non-meteorological factor data with a correlation coefficient greater than or equal to a preset addition threshold are selected as new associated factors. Non-meteorological factor data with a correlation coefficient less than a preset removal threshold are identified from the used non-meteorological factor data as associated factors to be removed. The newly added related factors are added to the feature set, and the related factors to be removed are removed from the feature set to obtain the updated feature set.
[0013] Optionally, when the cumulative error of the prediction error sequence exceeds a preset threshold, the prediction error sequence is used as a feedback signal to trigger a feedback evaluation of the feature set of the initial prediction model, including: The mean, variance, and recent trend slope of the prediction error sequence are determined, and the mean, variance, and recent trend slope are combined into an actual multidimensional error feature vector. The actual multidimensional error feature vector is compared with the preset multidimensional threshold vector dimension by dimension. When the values of at least two dimensions in the actual multidimensional error feature vector exceed the corresponding dimension thresholds in the preset multidimensional threshold vector, the cumulative error is determined to exceed the preset threshold. After determining that the cumulative error exceeds the preset threshold, the prediction error sequence is marked as the feedback signal; The evaluation depth level of the feedback evaluation is determined based on the excess magnitude of each dimension in the multidimensional error feature vector, wherein the evaluation depth level is used to control the degree of screening of subsequent correlation coefficients. Using the evaluation depth level and the prediction error sequence as input, a feedback evaluation of the feature set of the initial prediction model is triggered.
[0014] Optionally, the step of fine-tuning the parameters of the prediction model using the updated feature set, inputting the correlation factor data for the next prediction period into the fine-tuned prediction model, and outputting the load forecast value for the next prediction period includes: Based on the updated feature set, extract the values of the corresponding features from the historical load data to construct a fine-tuning sample set; Using the current parameters of the prediction model as initial parameters, the prediction model is incrementally trained using the fine-tuning sample set to obtain the prediction model with fine-tuned parameters. Obtain the correlation factor data for the next forecast period, and filter the corresponding feature values from the correlation factor data according to the updated feature set, input the fine-tuned forecast model with the parameters, and output the load forecast value for the next forecast period.
[0015] In a second aspect, the present invention provides a power grid time-varying load ultra-short-term forecasting system, comprising: The data decomposition unit is used to acquire historical load data and corresponding related factor data, and to decompose the historical load data into components to obtain the basic load component and the meteorological sensitive load component. The related factor data includes meteorological factor data and non-meteorological factor data. The time delay analysis unit is used to calculate the time delay mutual information based on the meteorological sensitive load component and each meteorological factor data respectively, and to determine the effective lag time corresponding to each meteorological factor data respectively. The model training unit is used to perform time-shift alignment on the corresponding meteorological factor data according to the effective lag time, take each of the time-shift aligned meteorological factor data, the base load component, and the non-meteorological factor data as input features, take the total load in the historical load data as the output target, construct a training sample set, and use the training sample set to pre-train the initial prediction model to obtain the trained prediction model. The prediction and feedback unit is used to input the relevant factor data of the current prediction period into the prediction model during the rolling prediction process based on the prediction model, obtain the load prediction value of the current prediction period and determine the prediction error, use the prediction error sequence accumulated over multiple consecutive prediction periods as a feedback signal, determine the correlation coefficient between each non-meteorological factor data and the prediction error sequence, and update the feature set of the initial prediction model according to the correlation coefficient. The fine-tuning prediction unit is used to fine-tune the parameters of the prediction model using the updated feature set, input the correlation factor data of the next prediction period into the parameter-fine-tuned prediction model, and output the load prediction value of the next prediction period.
[0016] The present invention provides a method and system for ultra-short-term forecasting of time-varying loads in power grids. By decomposing historical load data, it separates the base load component and the meteorologically sensitive load component, allowing for differentiated treatment of load components with different characteristics. Based on this, it calculates the time-delay mutual information between the meteorologically sensitive load component and various meteorological factors, automatically determining the effective lag duration of each meteorological factor and performing time-shift alignment. This accurately captures the lag effect of meteorological conditions on load, avoiding the blindness of manually setting lag durations, thus providing more precise input features for the forecasting model. Furthermore, the time-shift aligned meteorological factors, base load component, and non-meteorological factor data are used together as input features to construct a training sample set and pre-train the model. This enables the model to comprehensively utilize multi-source heterogeneous information, considering both the inherent trend of load changes and the dynamic influence of external related factors, effectively improving the baseline accuracy of the initial forecasting model.
[0017] Meanwhile, during the rolling forecasting process, this invention uses the cumulative forecast error sequence from multiple consecutive forecasting periods as a feedback signal to dynamically calculate the correlation coefficient between each non-meteorological factor data and the forecast error sequence, and updates the feature set of the initial forecasting model accordingly. This adaptive feature update mechanism enables the forecasting model to automatically discover new correlation factors introduced by new situations such as temporary events, staggered production, or sudden power rationing, while eliminating redundant or invalid features. This solves the problem in existing technologies where manually selected influencing factors are difficult to automatically adjust with changes in power consumption structure, significantly enhancing the adaptability and robustness of the forecasting method in complex and variable environments. Finally, the updated feature set is used to fine-tune the parameters of the forecasting model, achieving online adaptive optimization of the model with minimal computational cost. The correlation factors for the next forecasting period are input into the fine-tuned model to output the predicted value. The automatic update architecture ensures that the forecasting model always remains synchronized with the time-varying characteristics of the power grid operation, continuously providing high-precision ultra-short-term load forecasts.
[0018] In summary, this invention constructs high-quality input features for the model by separating load components, automatically identifying the effects of meteorological lags, and fusing multi-source data. Based on this, an error feedback mechanism dynamically senses changes in the operating environment, enabling the model to autonomously identify emerging correlation factors in new scenarios and promptly eliminate redundant features, maintaining the real-time effectiveness of the input features. Combined with a lightweight online model fine-tuning strategy, it can continuously provide high-precision ultra-short-term load forecasts under complex and ever-changing power grid operating conditions, providing reliable decision-making support for real-time dispatching, dynamic economic dispatching, and safety and stability early warning. Attached Figure Description
[0019] Figure 1 This is a flowchart illustrating the ultra-short-term forecasting method for time-varying loads in power grids according to an embodiment of the present invention. Figure 2 This is a schematic diagram of the structure of the ultra-short-term forecasting system for time-varying loads of the power grid according to an embodiment of the present invention. Detailed Implementation
[0020] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the accompanying drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.
[0021] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0022] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to"; the term "based on" means "at least partially based on"; the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments"; and the term "optionally" means "optional embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first," "second," etc., mentioned in this invention are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.
[0023] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0024] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties. The collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0025] Combination Figure 1 As shown in the figure, an embodiment of the present invention provides a method for ultra-short-term forecasting of time-varying load in a power grid, comprising: Historical load data and corresponding related factor data are acquired, and the historical load data is decomposed into components to obtain the basic load component and the meteorological sensitive load component. The related factor data includes meteorological factor data and non-meteorological factor data.
[0026] Specifically, historical load data for at least one year, with sampling intervals not exceeding fifteen minutes, is collected from the historical database of the power grid dispatch automation system. Simultaneously, meteorological data such as temperature, humidity, wind speed, and sunshine duration at the same resolution are obtained from numerical weather prediction service providers or the power meteorological network. Date type, time-of-use pricing, and special event markers such as temporary events, staggered production, and sudden power outages are extracted from the power market trading system, dispatch operation management system, and event logs. All data are aligned on a unified time axis, missing values are interpolated to complete the data, and abnormal peaks are statistically identified and smoothed out, forming a well-organized raw dataset. The preprocessed historical load sequence is then subjected to variational mode decomposition. After a preset number of decompositions, iterative decomposition is performed into several intrinsic mode components. The correlation coefficient between each component and typical meteorological factors such as temperature is calculated. Components with correlations higher than a first threshold are identified as meteorologically sensitive load components and superimposed into a comprehensive sequence. The remaining components are reconstructed as basic load components, thus completing the separation of load components.
[0027] Based on the time delay mutual information calculation of the meteorological sensitive load component and each meteorological factor data, the effective lag time corresponding to each meteorological factor data is determined.
[0028] Specifically, for each meteorological factor, a maximum lag time window covering a reasonable range of impact delays is preset, and a series of discrete lag durations are generated with the data sampling interval as the step size. For each lag duration, the meteorological sensitive load component sequence is shifted forward by this duration, and the mutual information value is calculated together with the original meteorological factor sequence. During the calculation, the two sets of sequences are first mapped to a two-dimensional phase space, and the joint probability density and marginal probability density are estimated using the nearest neighbor point statistical method. Then, the mutual information value under this lag condition is obtained according to the definition of mutual information. After traversing all lag durations, the curve of mutual information changing with lag time is obtained. The lag duration corresponding to the first significant peak point on the curve is found, which is the effective lag duration of the meteorological factor. If no significant peak point appears within the window, the lag point corresponding to the fastest increase in mutual information value is taken as the effective lag duration. The effective lag duration is determined independently for each meteorological factor in this way.
[0029] Based on the effective lag time, the corresponding meteorological factor data is time-shifted and aligned. Each time-shifted meteorological factor data, the base load component, and the non-meteorological factor data are used as input features. The total load in the historical load data is used as the output target to construct a training sample set. The initial prediction model is pre-trained using the training sample set to obtain the trained prediction model.
[0030] Specifically, based on the effective lag time of each meteorological factor, the original time series is shifted forward by the corresponding number of steps to align the timestamps of the shifted meteorological data with the actual time of impact on the load. All time-shifted and aligned meteorological factor data, the base load component series, and non-meteorological factor data such as date type, electricity price, and event markers are concatenated column-wise to form a multi-dimensional input feature matrix. Simultaneously, the measured total load value at the same time is used as the output target, forming a one-to-one corresponding training sample set. The sample set is divided into a training set and a validation set in chronological order. A recurrent neural network containing two long short-term memory hidden layers and one fully connected output layer is selected as the initial prediction model. Using the training set data, iterative training is performed using mean squared error as the loss function and an adaptive learning rate optimizer. The error is evaluated on the validation set in each round. When the validation error no longer decreases for several consecutive rounds, training is terminated and the optimal model parameters are saved, resulting in the trained baseline prediction model.
[0031] In the rolling forecasting process based on the forecasting model, the relevant factor data of the current forecasting period are input into the forecasting model to obtain the load forecast value of the current forecasting period and determine the forecasting error. The cumulative forecasting error sequence of multiple consecutive forecasting periods is used as a feedback signal to determine the correlation coefficient between each non-meteorological factor data and the forecasting error sequence. The feature set of the initial forecasting model is updated according to the correlation coefficient.
[0032] Specifically, the forecasting model enters a rolling forecasting mode. For each forecast period, it acquires current weather, electricity price, and event marker data from a real-time data interface. After time-shift alignment and feature stitching according to a predetermined process, this data is input into the forecasting model to obtain the load forecast for the current period. Once the measured load value for the current period is obtained, the error between the forecast and the measured value is calculated and pushed into a fixed-length first-in-first-out sliding window, ensuring that the window always contains the forecast error sequence of the most recent consecutive forecast periods. Using this error sequence as a feedback signal, the observation sequence of each non-meteorological factor within the same recent time period is extracted one by one. The correlation coefficient method is used to calculate the correlation coefficient between the non-meteorological factor sequence and the forecast error sequence, and its statistical significance is evaluated, thereby quantifying the strength of the correlation between various non-meteorological factors and recent forecast deviations.
[0033] For example, suppose a city's power grid has deployed the prediction model of this invention and entered online rolling prediction mode, performing load prediction every fifteen minutes. The non-meteorological factors used in the initial feature set include: date type (weekday, weekend, holiday), time-of-use electricity prices (peak, flat, and valley periods corresponding to different price values), and a "peak-shifting production" event marker manually entered by the dispatching and operation management system. One Friday, the city temporarily announces a large-scale nighttime football match to be held at the central stadium that evening, expected to last until late. Because this match information is a temporary decision, there is no corresponding "temporary match" event marker in the initial feature set. Therefore, in the first few prediction cycles before and during the match, the model input features fail to reflect this newly added electricity consumption impact factor. After the match begins, the electricity load of commerce, lighting, and transportation around the stadium increases significantly, but the model still predicts according to the characteristics of a regular weekday, resulting in the output load prediction value being lower than the actual load value. The system continuously calculates the prediction error in each prediction cycle and pushes the error into a first-in-first-out sliding window of length ninety-six (corresponding to all prediction cycles within a day). As the competition progressed, positive deviations appeared for several consecutive periods, and the prediction error sequence within the window gradually showed a trend of continuous positive and increasing magnitude.
[0034] Approximately one hour into the event, the root mean square error of the accumulated prediction error sequence within the sliding window exceeded the preset alarm threshold, immediately triggering a feedback evaluation of the feature set. The system extracted the actual observation sequences of various non-meteorological factors from memory over the past ninety-six periods and calculated the Pearson correlation coefficient between each sequence and the current prediction error sequence. The results showed that the event marker "temporary event," previously not included in the features, had a correlation coefficient as high as 0.58 with the prediction error sequence due to the incremental electricity consumption effect at multiple monitoring points around the venue. Furthermore, the p-value was far less than 0.05, significantly exceeding the addition threshold of 0.3, and was automatically identified and marked as a newly added associated factor. Simultaneously, the event marker "off-peak production," originally in the feature set, had an observation sequence that remained at zero values due to the lack of recent off-peak production scheduling instructions. Its correlation coefficient with the prediction error sequence was only 0.01, below the removal threshold of 0.05, and was marked as an associated factor to be removed. The correlation coefficients between the two non-meteorological factors, time-of-use electricity price and date type, and the prediction error series were 0.21 and 0.33, respectively. The former did not reach the threshold for addition, while the latter remained within the retention range, maintaining its original state.
[0035] Based on the feedback and evaluation results, the feature set is automatically updated: "Temporary Events" is added as a new feature dimension to the feature set, and a one-hot encoding rule and corresponding input feature vector position are assigned to it; simultaneously, the "Off-Peak Production" event marker is removed from the feature set, reducing the dimension of the input feature vector. After the update, the system uses historical data from the last seven days to construct a fine-tuning sample set, and only performs incremental fine-tuning on the model output side-layer parameters with a small learning rate, enabling the model to quickly absorb the information of the new features and adapt to changes in feature dimensions. After fine-tuning, when the next prediction cycle arrives, the event marker value of "Temporary Events" can be read as active from the meteorological, electricity price, and event marker data obtained from the real-time interface. This marker value is concatenated with other features and input into the fine-tuned prediction model. The load prediction value output by the model successfully captures the additional electricity demand brought about by the event, and the prediction error quickly falls back to the normal range.
[0036] The updated feature set is used to fine-tune the parameters of the prediction model, and the correlation factor data for the next prediction period is input into the fine-tuned prediction model to output the load prediction value for the next prediction period.
[0037] Specifically, a high correlation threshold for including new features and a low correlation threshold for removing redundant features are set, and the correlation coefficient calculation results of all non-meteorological factors are iterated. Non-meteorological factors with absolute correlation coefficients greater than the high threshold and statistically significant are added if they were not previously included in the model's input features; non-meteorological factors with absolute correlation coefficients lower than the low threshold in multiple consecutive evaluation periods are removed from the existing feature set. It should also be noted that the LSTM prediction model learns the general temporal characteristics and long-term dependency patterns of load changes through a large amount of historical data during the pre-training phase, and this knowledge is embedded in the network weights. If all network parameters are significantly updated during the fine-tuning phase, the new data may overwrite the learned old knowledge, causing the model to lose its memory of general patterns while adapting to new scenarios, resulting in a decrease in prediction accuracy instead of an increase. To avoid the aforementioned problems, this embodiment takes the following measures during fine-tuning to balance old and new data: First, the weights of the main network are frozen, meaning the underlying network parameters that have converged during the pre-training phase remain unchanged, ensuring that the general temporal features learned by the model from historical data are not overwritten by gradient updates from new data. Second, only a small number of parameters on the input and output sides are fine-tuned, enabling the model to adapt to new feature dimensions or recent changes in data distribution, while the model's overall basic cognitive ability to handle load changes remains unaffected. Third, a learning rate much lower than that used in the pre-training phase is used for parameter updates, ensuring that the parameter adjustments are constrained to a small neighborhood of the optimal solution in the pre-training phase. Finally, fine-tuning training is terminated after only a few rounds, preventing the model from iterating repeatedly on limited recent samples and becoming overly biased towards new data. Through these methods, after fine-tuning, the model can absorb new information brought about by recent changes in the operating environment while fully retaining the general predictive capabilities established during the pre-training phase, thus achieving an effective balance between old and new data. Based on this updated feature set, the weights of the main network of the prediction model are frozen, and only a few rounds of fine-tuning training are performed on the parameters of the input layer and related feature extraction layers, enabling the model to quickly adapt to changes in feature dimensions. After fine-tuning, the correlation factor data for the next prediction period is organized according to the updated feature set and input into the fine-tuned prediction model, outputting the load prediction value for the next period. Error calculation, correlation assessment, feature updating, and parameter fine-tuning are repeated in each prediction period to achieve continuous adaptive tracking of the prediction model to the time-varying characteristics of the power grid.
[0038] In one embodiment of the present invention, real-time data interaction is achieved with the power grid dispatch automation system, numerical weather prediction system, power market trading system and dispatch operation management system through a data interface, forming a fully automated closed loop from data collection, feature construction, model training to online rolling forecasting.
[0039] Under normal operating conditions, the system continuously acquires historical load data, meteorological factor data, and non-meteorological factor data updated every fifteen minutes. Variational mode decomposition is used to separate the load data into a basic load component and a meteorologically sensitive load component. The meteorologically sensitive load component performs time-delay mutual information calculations with each meteorological factor, such as temperature, humidity, wind speed, and solar irradiance, automatically identifying the lag characteristics of different meteorological factors' impact on the load. For example, the lag effect of summer temperature changes on air conditioning load is typically one to two hours, while the lag response of humidity changes to industrial load may manifest several hours later. The system can independently determine the optimal lag time for each meteorological factor without requiring manual experience-based settings.
[0040] When the power grid operating environment changes significantly due to events such as temporary sporting events, staggered production, or sudden power outages, the cumulative error index of the prediction error sequence triggers a feedback evaluation mechanism. The system automatically calculates the correlation coefficients between each non-meteorological factor and the recent prediction deviation, dynamically identifies newly added influencing factors, and eliminates redundant factors that have failed. Subsequently, it uses recent historical data to fine-tune the model parameters, enabling the prediction model to quickly adapt to the current operating conditions. The entire process, from event occurrence to feature update and model adaptation, can be completed automatically within several prediction cycles, without interrupting the continuous operation of rolling predictions.
[0041] For example, during the peak summer season, this embodiment can accurately capture the lagged cumulative effect of continuously rising temperatures on air conditioning load, while automatically identifying load drops caused by temporary power rationing orders, and adjusting model input features in a timely manner to maintain prediction accuracy within the acceptable deviation range for dispatching operations. In special scenarios such as power supply for large-scale events and responses to extreme weather, the aforementioned adaptive mechanism continuously provides reliable load forecast data for real-time grid dispatching and safe and stable control, effectively supporting the refined operation and management of the power grid in complex and ever-changing environments.
[0042] This embodiment of the power grid time-varying load ultra-short-term forecasting method decomposes historical load data into components, separating the base load component and the meteorologically sensitive load component, allowing for differentiated treatment of load components with different characteristics. Based on this, time-delay mutual information calculation is performed on the meteorologically sensitive load component and various meteorological factors. The effective lag duration of each meteorological factor is automatically determined and time-shifted, accurately capturing the lag effect of meteorological conditions on load, avoiding the blindness of manually setting lag durations, thus providing more accurate input features for the forecasting model. Furthermore, the time-shifted aligned meteorological factors, base load component, and non-meteorological factor data are used together as input features to construct a training sample set and pre-train the model. This allows the model to comprehensively utilize multi-source heterogeneous information, considering both the inherent trend of load changes and the dynamic influence of external related factors, effectively improving the baseline accuracy of the initial forecasting model.
[0043] Meanwhile, during the rolling forecasting process, this invention uses the cumulative forecast error sequence from multiple consecutive forecasting periods as a feedback signal to dynamically calculate the correlation coefficient between each non-meteorological factor data and the forecast error sequence, and updates the feature set of the initial forecasting model accordingly. This adaptive feature update mechanism enables the forecasting model to automatically discover new correlation factors introduced by new situations such as temporary events, staggered production, or sudden power rationing, while eliminating redundant or invalid features. This solves the problem in existing technologies where manually selected influencing factors are difficult to automatically adjust with changes in power consumption structure, significantly enhancing the adaptability and robustness of the forecasting method in complex and variable environments. Finally, the updated feature set is used to fine-tune the parameters of the forecasting model, achieving online adaptive optimization of the model with minimal computational cost. The correlation factors for the next forecasting period are input into the fine-tuned model to output the predicted value. The automatic update architecture ensures that the forecasting model always remains synchronized with the time-varying characteristics of the power grid operation, continuously providing high-precision ultra-short-term load forecasts.
[0044] In summary, this embodiment constructs high-quality input features for the model by separating load components, automatically identifying the effects of meteorological lags, and fusing multi-source data. Based on this, an error feedback mechanism dynamically senses changes in the operating environment, enabling the model to autonomously identify emerging correlation factors in new scenarios and promptly eliminate redundant features, maintaining the real-time effectiveness of the input features. Combined with a lightweight online model fine-tuning strategy, it can continuously provide high-precision ultra-short-term load forecasts under complex and ever-changing power grid operating conditions, providing reliable decision-making support for real-time dispatching, dynamic economic dispatching, and safety and stability early warning.
[0045] Optionally, the step of acquiring historical load data and corresponding related factor data, and performing component decomposition on the historical load data to obtain the basic load component and the meteorologically sensitive load component, includes: The historical load data is decomposed into time series to extract the trend component and periodic component of the load. The trend component and the periodic component are superimposed to obtain the basic load component. Based on the historical load data and the basic load components, the remaining load sequence is obtained; Based on the correlation coefficient between the remaining load sequence and each meteorological factor data, meteorological factor data with a correlation coefficient greater than a preset threshold are selected as sensitive meteorological factors; The remaining load sequence is regressed and fitted to the sensitive meteorological factors to obtain the meteorological sensitive load components.
[0046] Specifically, the preprocessed historical load sequence is decomposed into a time series using a seasonal and trend decomposition algorithm based on local weighted regression. The period parameter is set as the number of sampling points within a day; for example, when the sampling interval is fifteen minutes, the period parameter is set to 96. Trend components reflecting long-term load growth or decline, and periodic components reflecting fixed daily or weekly fluctuations, are extracted. The values of the trend components and periodic components at each sampling time are added point-by-point to synthesize the basic load component sequence. The original historical load sequence is then subtracted point-by-point from the basic load component sequence obtained in the previous step, with the basic load value subtracted from the original load value, resulting in a residual load sequence after removing long-term trends and fixed periodic patterns. This residual load sequence contains load variation information caused by meteorological fluctuations, random events, and other non-periodic factors. The Pearson correlation coefficient between the remaining load sequence and each collected meteorological factor data is calculated one by one. For example, the correlation coefficients between the remaining load sequence and the temperature sequence, humidity sequence, and wind speed sequence are calculated. Meteorological factors whose absolute correlation coefficient values exceed a preset threshold are screened out. This preset threshold is generally set above zero and ensures that at least two meteorological factors are screened out. The screened meteorological factors are marked as sensitive meteorological factors. For example, in the historical summer data of a certain region, the absolute values of the correlation coefficients between the remaining load sequence and six meteorological factors—temperature, perceived temperature, relative humidity, wind speed, precipitation, and sunshine duration—are 0.52, 0.48, 0.35, 0.24, 0.10, and 0.28, respectively. If the preset threshold is set to 0.25, the meteorological factors with a correlation coefficient greater than 0.25 are temperature, perceived temperature, relative humidity, and sunshine duration, totaling four, which meets the requirement of selecting at least two meteorological factors. If the threshold is set to 0.40, only temperature and perceived temperature are selected, still meeting the requirement of at least two factors, but relative humidity and sunshine duration, which also affect the load, may be missed. If the threshold is set to 0.20, five factors can be selected: temperature, perceived temperature, relative humidity, wind speed, and sunshine duration. However, too many factors may introduce precipitation (correlation coefficient 0.10), which has a weak correlation with the load, thus increasing the redundancy of subsequent modeling. Based on the above comparison, those skilled in the art can select an appropriate threshold under the guidance of this principle, according to the distribution of correlation coefficients between each meteorological factor and the remaining load sequence in the actual data, preferably within the range of 0.2 to 0.35. Taking the data in this example as an example, when the threshold is set to 0.25, it ensures that a sufficient number of sensitive meteorological factors are included in the subsequent analysis, while excluding weakly correlated factors, which is a relatively reasonable value.Using the selected sensitive meteorological factors as the set of independent variables and the remaining load sequence as the dependent variable, a multiple linear regression model or a support vector regression model is constructed. The least squares method or the sequence minimum optimization algorithm is used to fit and solve the model parameters. After the regression fitting is completed, the observed values of the sensitive meteorological factors are input into the fitted regression model, and the predicted value sequence output by the model is the meteorological sensitive load component.
[0047] In this optional embodiment, by performing time-series decomposition on historical load data, extracting trend and periodic components, and superimposing them to obtain the base load component, the relatively stable and predictable inherent components of the load can be separated first, allowing the base load component to accurately reflect the long-term trend and periodic pattern of electricity demand. Subtracting the base load component from the original load yields the remaining load sequence, concentrating the impact of meteorological fluctuations and random disturbances into the remaining sequence, providing a pure analytical object free from trend interference for subsequent identification of meteorologically sensitive factors. The correlation coefficients between the remaining load sequence and each meteorological factor data are calculated one by one, and meteorological factors with correlation coefficients greater than a preset threshold are selected as sensitive meteorological factors. This automatically identifies meteorological variables that truly have a significant impact on the load based on the statistical correlation of the data itself, avoiding the inclusion of irrelevant or weakly correlated meteorological factors in subsequent modeling and reducing feature redundancy. The remaining load sequence is regressed and fitted to the selected sensitive meteorological factors, and the resulting fitted sequence serves as the meteorologically sensitive load component. This allows the meteorologically sensitive component to accurately reflect the marginal contribution of changes in meteorological conditions to the load through quantitative regression relationships, with clear physical meaning and strong interpretability. Overall, the basic load component and the meteorologically sensitive load component are strictly distinguished in terms of their causes and do not overlap in terms of their values. This provides a reliable foundation for subsequent time-delay mutual information analysis of the meteorologically sensitive component and for building high-quality input features for the prediction model, directly improving the model's ability to capture the mechanism of load changes and its prediction accuracy.
[0048] Optionally, the step of calculating the time delay mutual information based on the meteorological sensitive load component and each meteorological factor data to determine the effective lag time corresponding to each meteorological factor data includes: The sequence composed of each meteorological factor data and the meteorological sensitive load component is used as the sequence to be analyzed. A preset number of candidate lag durations are set, and the meteorological factor data is time-shifted according to the candidate lag durations to obtain the time-shifted sequence corresponding to each candidate lag duration; For each time-shifted sequence, the mutual information value between the time-shifted sequence and the meteorological sensitive load component is determined, and a mutual information sequence is obtained based on the mutual information value corresponding to each candidate lag duration; The candidate lag duration corresponding to the maximum value of the mutual information value in the mutual information sequence is determined as the effective lag duration corresponding to the meteorological factor data.
[0049] Specifically, the decomposed comprehensive meteorological sensitive load component sequence, along with all original meteorological factor sequences obtained from the numerical weather prediction service provider interface or the power meteorological network, including but not limited to temperature sequences, perceived temperature sequences, relative humidity sequences, wind speed sequences, wind direction sequences, cloud cover sequences, solar irradiance sequences, and precipitation sequences, are used together as paired sequences to be analyzed, preparing for the calculation of time delay mutual information for each pair of sequences. For the meteorological factor of temperature, a maximum lag time window is set. The duration of this window is determined based on the physical delay characteristics of the meteorological factor's impact on the load, generally set to 24 hours or 48 hours. Using the sampling time interval of historical load data as the basic step size, a preset number of candidate lag times are evenly divided within this window. For example, with a sampling interval of 15 minutes and a maximum window of 24 hours, a total of 96 candidate lag times are generated. For each of these 96 candidate lag times, a time-shift operation is sequentially performed on the original temperature sequence, that is, the entire temperature sequence is shifted forward along the time axis by the number of sampling points corresponding to that candidate lag time, resulting in 96 time-shifted temperature sequences corresponding to different candidate lag times.
[0050] For each candidate lag duration, the time-shifted temperature sequence is aligned with the meteorological sensitive load component sequence by timestamp, and the overlapping period where both have valid values is extracted. The time-shifted temperature sequence and the meteorological sensitive load component sequence within the overlapping period are mapped into a two-dimensional phase space. The mutual information value between the two is calculated using a k-nearest neighbor-based mutual information estimation algorithm, where k is generally taken as four to six. During the calculation, the k nearest neighbors of each sample point in the two-dimensional phase space are first found. The joint probability density function and the marginal probability density function are estimated by statistically analyzing the distribution of sample points within each neighborhood. Then, the mutual information value is approximated by discrete summation according to the integral definition formula of mutual information. The calculated mutual information value is recorded as the mutual information value corresponding to the candidate lag duration. After traversing all ninety-six candidate lag durations, a mutual information sequence of length ninety-six is obtained. This sequence reflects the changes in the statistical dependence between temperature and meteorological sensitive load under different lag times. Analyzing the mutual information sequence obtained above, the maximum value of the mutual information value is found in the sequence, and the candidate lag time corresponding to the maximum value point is determined as the effective lag time of the meteorological factor of temperature. If there are multiple local peaks with similar values in the mutual information sequence, the candidate lag time corresponding to the first significant peak point is selected as the effective lag time, prioritizing the capture of the earliest lag effect. Following the same processing procedure, the mutual information sequence between each other meteorological factor data such as humidity, wind speed, and solar irradiance and the meteorological sensitive load component is calculated independently, and the corresponding effective lag time is determined for each. Finally, a unique effective lag time parameter for each meteorological factor is obtained. These parameters are used for time-shift alignment of meteorological factors.
[0051] In this optional embodiment, each meteorological factor data is paired with a meteorologically sensitive load component as a sequence to be analyzed. Multiple candidate lag durations covering a reasonable range of impact delays are set for each meteorological factor. By traversing the candidate lag durations to generate corresponding time-shifted sequences, the correlation strength between meteorological factors and meteorologically sensitive load components at different lag times can be comprehensively scanned with fine temporal granularity. Mutual information values are calculated for each time-shifted sequence and the meteorologically sensitive load component. The ability of mutual information to capture nonlinear statistical dependencies between variables is utilized, which, compared to linear correlation coefficients, more fully reflects the complex nonlinear delay coupling characteristics that may exist between meteorological factors and meteorologically sensitive loads, avoiding omissions or misjudgments of effective lag relationships due to linear assumptions. The candidate lag duration corresponding to the maximum mutual information value in the mutual information sequence is determined as the effective lag duration of the meteorological factor. This data-driven approach can automatically lock the optimal lag time without relying on empirical rules or repeated manual trial and error to determine the lag parameters of each meteorological factor. This improves the accuracy and efficiency of differentiating the lag characteristics of different meteorological factors and ensures that the determined lag duration is statistically the moment when the meteorological factor is most closely related to the load response. This ensures that the meteorological characteristics input into the prediction model correspond precisely to load changes in time, laying the foundation for improving the model's prediction accuracy.
[0052] Optionally, for each time-shifted sequence, determining the mutual information value between the time-shifted sequence and the meteorologically sensitive load component, and obtaining a mutual information sequence based on the mutual information value corresponding to each candidate lag duration, includes: For each time-shifted sequence, joint probability density estimation and marginal probability density estimation are performed on the time-shifted sequence and the meteorological sensitive load component, and the mutual information value corresponding to the time-shifted sequence is determined based on the results of the joint probability density estimation and the results of the marginal probability density estimation. The mutual information values of each meteorological factor data under all candidate lag times are arranged in order of lag time to obtain the initial mutual information sequence corresponding to the meteorological factor data, and the initial mutual information sequence is smoothed to obtain a smoothed mutual information sequence. In the smooth mutual information sequence, the search proceeds from the zero lag position in the direction of increasing lag duration to determine the lag duration corresponding to the first peak, and it is determined whether there is a significant decrease after the first peak. If there is a significant decrease, the lag duration corresponding to the first peak is determined as the effective lag duration. The condition for determining a significant decrease is that the mutual information value at the k-th lag point after the first peak is lower than the product of the mutual information value of the first peak and a preset proportional coefficient, where k is a preset search step size.
[0053] Specifically, for each time-shifted sequence generated under a candidate lag duration, it is aligned with the meteorological sensitive load component sequence by timestamp and the overlapping effective time period is extracted. Then, a joint probability density estimation and marginal probability density estimation are performed using a kernel density estimation method. The kernel function is a Gaussian kernel, and the bandwidth is automatically selected through cross-validation or empirical rules. The sample points of the two sequences are projected onto a two-dimensional plane, and the contribution of all sample points is superimposed on each grid point using the kernel function to obtain a two-dimensional joint probability density function. Then, the joint probability density function is integrated and summed along its respective dimension to obtain the respective one-dimensional marginal probability density function. After that, the values of the joint probability density function and the two marginal probability density functions at each sample point are substituted into the mutual information calculation formula. By traversing all sample points, accumulating point by point and taking the average, the mutual information value between the time-shifted sequence and the meteorological sensitive load component is obtained. For the meteorological factor of temperature, such as the aforementioned ninety-six candidate lag times, the ninety-six mutual information values calculated one by one under the ninety-six candidate lag times are arranged in ascending order of candidate lag times to form an initial mutual information sequence reflecting the change of mutual information with lag time. The initial mutual information sequence is smoothed using a local weighted scatter smoothing method, with the smoothing window width ranging from one-tenth to one-fifth of the sequence length, to eliminate local jitter in the sequence caused by sampling noise or short-term random fluctuations, making the overall trend and peak structure of mutual information change with lag time clearer and more discernible, thus obtaining a smoothed mutual information sequence.
[0054] In this embodiment, for each candidate lag duration, the time-shifted sequence generated is aligned one by one with the meteorological sensitive load component sequence according to the timestamp. The overlapping time period where both have complete data records on the time axis is extracted to obtain a set of aligned sample pairs. Each sample pair contains the time-shifted sequence value and the meteorological sensitive load component value at the same moment. All sample points in this set of sample pairs are plotted on a two-dimensional plane, with each sample point corresponding to a coordinate position on the plane. The kernel density estimation method is used to perform probability density analysis on this set of two-dimensional sample points. Specifically, a Gaussian kernel is selected as the kernel function, and a Gaussian distribution is generated with each sample point as the center. The range of the Gaussian distribution spreading outward is controlled by the bandwidth parameter. The Gaussian distributions generated by all sample points are superimposed point by point on the plane. After superposition, a joint probability density distribution is formed on the entire two-dimensional plane. The density value at any position is equal to the sum of the contributions of all sample points at that position. The bandwidth parameter is then automatically determined using cross-validation or empirical rules. Cross-validation involves dividing the sample data into training and validation parts, evaluating the estimation error for different bandwidth values, and selecting the bandwidth value that minimizes the estimation error. Empirical rules estimate a suitable bandwidth value based on the standard deviation and sample size of the sample data, eliminating the need for repeated manual trial and error. After obtaining the joint probability density distribution on the two-dimensional plane, the joint probability density distribution is integrated and summed in each variable direction. Specifically, the joint probability density distribution is summed along the direction of the meteorological sensitive load component to obtain the marginal probability density distribution of the time-shifted sequence variable; the joint probability density distribution is also summed along the direction of the time-shifted sequence to obtain the marginal probability density distribution of the meteorological sensitive load component variable. These two marginal probability density distributions reflect the independent distribution of the two variables, respectively. After obtaining the joint probability density distribution and the two marginal probability density distributions, for each sample point within the overlapping time period, the joint probability density value at that point and the corresponding probability density values on the two marginal probability density distributions are extracted. These three density values are calculated and accumulated point-by-point according to the definition of mutual information, and the average value of all sample points is taken to obtain the mutual information value between the time-shifted sequence and the meteorological sensitive load component under the candidate lag duration. This mutual information value quantitatively reflects the statistical dependence between the meteorological factor and the meteorological sensitive load component under the lag duration. By calculating the corresponding mutual information value for each candidate lag duration, an initial mutual information sequence reflecting the change of mutual information with lag duration can be obtained, which is used for subsequent smoothing processing and effective lag duration determination.
[0055] On a smooth mutual information sequence, starting from the position where the lag time is zero, scan point by point along the direction where the lag time gradually increases to find the first local peak point in the sequence. The condition for determining a local peak point is that the mutual information value of the point is greater than the mutual information values of its left and right adjacent points. Once the first peak point is found, immediately record the lag time position and mutual information value of the peak point. After determining the first peak, starting from that peak position, skip a preset search step size in the direction of increasing lag duration. For example, the step size is three sampling intervals, i.e., three lag points. When the sampling interval is fifteen minutes, this step size corresponds to forty-five minutes. Check whether the smoothed mutual information value at the k-th lag point is lower than the product of the mutual information value of the first peak and a preset scaling factor. This scaling factor is usually between 0.7 and 0.8, preferably 0.75. If this condition is met, it is determined that there is a significant decrease after the first peak, indicating that the peak reflects the complete transition characteristics of the meteorological factor's influence on the load from establishment to decay, rather than a false spike caused by noise. At this time, the candidate lag duration corresponding to the first peak is determined as the effective lag duration of the meteorological factor data. If the significant decrease condition is not met, continue searching for the next peak after that peak point and repeat the above determination until a peak that meets the condition is found.
[0056] In this optional embodiment, the kernel density estimation method is used to jointly estimate the marginal probability density of the time-shifted sequence and the meteorologically sensitive load component and calculate the mutual information value. This can accurately quantify the nonlinear dependence between the two and avoid the limitations of the linear correlation coefficient. The mutual information values under each candidate lag time are arranged sequentially and smoothed to effectively filter out sampling noise and random jitter, making the regular trend of mutual information change with lag time more prominent. The search for the first peak value starting from zero lag and increasing in the direction of increase is used for verification, combined with the significant decrease condition that the mutual information value at a specific step after the peak value is lower than the product of the peak value and a preset proportional coefficient. This can automatically capture the first optimal response moment when meteorological factors have a substantial impact on the load, while eliminating false peak interference caused by data fluctuations. This ensures that the determined effective lag time is statistically reliable and physically reasonable, providing a solid foundation for accurate time-shift alignment of subsequent meteorological data.
[0057] Optionally, the step of time-shifting the corresponding meteorological factor data according to the effective lag duration, using each time-shifted meteorological factor data, the base load component, and the non-meteorological factor data as input features, and using the total load in the historical load data as the output target, constructs a training sample set, and uses the training sample set to pre-train the initial prediction model to obtain the trained prediction model, includes: For each historical moment, based on the effective lag time corresponding to each meteorological factor data, the original meteorological factor data of the earlier moment that is separated from the historical moment by the effective lag time is extracted, and used as the time-shifted aligned meteorological factor data of the historical moment; The time-shifted aligned meteorological factor data, the base load component of the historical time, and the non-meteorological factor data are combined to form the input feature vector of the historical time, and the total load of the historical time is used as the output target to form an input-output sample pair; The input-output sample pairs from all historical moments are aggregated to construct the training sample set; The initial prediction model is trained using the training sample set. The parameters of the initial prediction model are iteratively updated by minimizing the error between the load prediction value output by the initial prediction model and the total load until the preset training conditions are met, thus obtaining the trained prediction model.
[0058] Specifically, for each historical moment, the timestamp of that historical moment is obtained, and the effective lag time corresponding to various meteorological factor data is called one by one. Taking temperature as an example, if the effective lag time of temperature is two hours, then the timestamp of that historical moment is used as the basis to go back two hours, and the measured temperature value of that earlier moment is extracted from the original temperature sequence that has not undergone time shift. This measured temperature value is assigned to the current historical moment as the temperature feature value after time shift alignment. For other meteorological factors such as humidity, wind speed, and solar irradiance, they are extracted in the same way as when their effective lag time is the same. Each historical moment obtains a set of time-shift aligned meteorological factor data that accurately corresponds to the actual impact time.
[0059] At each historical moment, the values of all meteorological factor data after time-shift alignment, the base load component value obtained at that historical moment, and the corresponding non-meteorological factor data, including date type codes, time-of-use electricity price values, and unique thermal codes or binary flag bits for various special events, are horizontally concatenated in a preset fixed order to form a multi-dimensional numerical vector, which serves as the input feature vector for that historical moment. Simultaneously, the measured total load value for that historical moment is read from the original historical load data and used as the output target label corresponding to the input feature vector, forming a complete input-output sample pair. Furthermore, following the timestamp order from earliest to latest, all available historical moments are traversed, and the input-output sample pairs generated for each historical moment are collected sequentially. Boundary moment samples that cannot form a complete input feature vector due to missing data from earlier moments caused by meteorological factor time-shift alignment are removed. The remaining sample pairs are summarized in chronological order into a complete training sample set containing all valid historical moments. The number of rows in this training sample set equals the total number of valid historical moments, and the number of columns is the input feature vector dimension plus one output target dimension.
[0060] In this embodiment, a Long Short-Term Memory (LSTM) recurrent neural network is selected as the initial prediction model. This network includes an input layer, two LSM hidden layers, and a fully connected output layer. The number of neurons in the input layer is equal to the dimension of the input feature vector. Each LSM hidden layer contains 128 memory units, and a dropout layer with a dropout probability of 0.2 is added after each LSM hidden layer to prevent overfitting. The fully connected output layer consists of one neuron and outputs the predicted load value. The constructed training sample set is divided into the first 80% as the training set and the last 20% as the validation set according to time sequence. The input feature vectors in the training set are input into the initial prediction model in batches. After calculating the predicted load value layer by layer through forward propagation, the mean squared error loss function is used to calculate the error between the predicted load value and the corresponding output target label, i.e., the measured total load value. An adaptive moment estimation optimizer is used with an initial learning rate of 0.001. The gradient of the loss function with respect to the weights and bias parameters of each layer in the network is calculated through the backpropagation algorithm. The parameters are iteratively updated based on the gradient. After each round of traversal of the entire training set, i.e., the validation error is calculated on the validation set, the early stopping mechanism is triggered when the validation error has not decreased in fifteen consecutive rounds of iteration, training is stopped, and the model parameters of the round with the smallest error on the validation set are saved as the prediction model after training is completed.
[0061] In this optional embodiment, by extracting the original meteorological data of the earlier time with an effective lag time interval for each historical moment as the time-shifted aligned feature value, a precise causal correspondence between meteorological factors and load response in time series is achieved, avoiding distortion of the relationship between input features and output target caused by lag mismatch. The time-shifted aligned meteorological factors, base load components, and non-meteorological factor data are combined into an input feature vector, and sample pairs are constructed one by one with the total load at the same historical moment as the output target. This ensures that each sample fully carries the synchronous information of all known influencing factors at that moment, and the mapping relationship between input and output is clear, facilitating model learning. All effective sample pairs are collected in chronological order to construct a training sample set, and forward propagation is used to calculate the predicted value. The model parameters are iteratively updated in reverse with the goal of minimizing the mean square error until the early stopping condition is met. This allows the model to fully learn historical patterns while effectively suppressing overfitting, ultimately obtaining a prediction model with high prediction accuracy and strong generalization ability under baseline conditions.
[0062] Optionally, the step of training the initial prediction model using the training sample set, and iteratively updating the parameters of the initial prediction model by minimizing the error between the load prediction value output by the initial prediction model and the total load, until a preset training condition is met, to obtain the trained prediction model, includes: Each of the input feature vectors in the training sample set is input into the initial prediction model to obtain the corresponding load prediction value; Based on the predicted load value and the corresponding total load, the prediction residual is determined, and a loss function is constructed using the sum of squares of the prediction residuals; The parameters of the initial prediction model are backpropagated using the loss function to determine the gradient of each parameter, and the parameters are iteratively updated based on the gradient. After each iteration update, the loss value of the loss function under the current parameters is determined. When the change in the loss value between two adjacent iterations is less than the preset convergence threshold, it is determined that the preset training condition is met, and the prediction model under the current parameters is taken as the prediction model that has been trained.
[0063] Specifically, the input feature vectors corresponding to each historical moment in the training sample set are extracted one by one and fed into the input layer of the initial prediction model in chronological order. The input layer passes the received multidimensional vectors to the first long short-term memory hidden layer as is. The forget gate, input gate, and output gate inside the first long short-term memory hidden layer calculate the forgetting ratio, the increment of new information, and the output control signal according to the input vector at the current moment and the hidden state at the previous moment, respectively, update the memory cell state at the current moment, and generate the hidden state vector of the layer. The hidden state vector of the first long short-term memory hidden layer is used as the input of the second long short-term memory hidden layer. The second long short-term memory hidden layer further extracts higher-level temporal features in the same way. Finally, the fully connected output layer maps the hidden state of the last moment of the second long short-term memory hidden layer into a scalar value, which is the load prediction value corresponding to the input feature vector.
[0064] For each training sample, the predicted load value output by the model is subtracted point by point from the output target, i.e., the measured total load value at that historical moment, in the same sample to obtain the prediction residual for that sample. After squaring the prediction residual element by element, the arithmetic mean of the squared residuals of all samples in a batch of training samples is calculated. This mean is used as the loss function value of the current batch. The loss function adopts the mean squared error form, and the batch size is set to sixty-four samples. After obtaining the loss function value for a batch, the loss function value is backpropagated using an automatic differentiation mechanism. Starting from the output layer, the partial derivatives of the loss function value with respect to the weights of the fully connected layer, the biases of the fully connected layer, the weight matrix and bias vector of each gate unit in the second long short-term memory hidden layer, and the weight matrix and bias vector of each gate unit in the first long short-term memory hidden layer are calculated layer by layer according to the chain rule, so as to obtain the gradient value of each parameter in each layer. Then, an adaptive moment estimation optimizer is used to calculate the adaptive learning rate for each parameter using the first moment estimation and second moment estimation of the gradient. The product of the adaptive learning rate and the corresponding gradient is used as the parameter update amount. The update amount is subtracted from the current parameter value to complete the update of all model parameters after training in this batch in this iteration. After traversing all batches in the training set and updating all parameters in each round, a complete forward propagation is performed on the entire validation set. The mean squared error loss value of all samples in the validation set is calculated and recorded. The validation loss value of this round is compared with the validation loss value of the previous round, and the absolute value of the difference between the two is calculated as the change in the loss value. When the change in this value is less than the preset convergence threshold for three consecutive rounds, the preset training condition is considered met. The preset convergence threshold is set to [value missing]. to Preferably, the preset convergence threshold is 5× At this point, stop iterative training and save the prediction model under the current parameters as the completed prediction model.
[0065] In this optional embodiment, the input feature vectors from the training sample set are fed into the initial prediction model one by one. After the long short-term memory hidden layer extracts the temporal dependencies layer by layer and the fully connected output layer maps them, the load prediction value for each historical moment is obtained, realizing end-to-end forward computation from multi-source input features to load output. The mean squared error loss function is constructed using the sum of squared prediction residuals between the predicted value and the measured total load. This quantifies the prediction bias of the model on each sample into a unified scalar optimization objective, making the update direction of the model parameters clearly point to minimizing the overall prediction error. The gradient of the parameters of each layer of the model is obtained by backpropagation using the loss function, and the parameters are iteratively updated using an adaptive moment estimator based on the gradient. This can automatically adjust the learning step size of each parameter according to the first and second moments of the gradient, ensuring the convergence speed while avoiding training oscillations caused by excessive parameter update amplitude. After each iteration, the change in the loss value is calculated. When the change in the loss value of adjacent iterations is less than the preset convergence threshold, the training is considered complete. Training is terminated with strict quantitative criteria, which can ensure that the model parameters are close enough to the optimal solution and prevent invalid iterations and overfitting. The final prediction model has stable and excellent load prediction accuracy under the baseline working conditions.
[0066] Optionally, in the rolling forecasting process based on the forecasting model, the relevant factor data for the current forecast period are input into the forecasting model to obtain the load forecast value for the current forecast period and determine the forecast error. The cumulative forecast error sequence over multiple consecutive forecast periods is used as a feedback signal to determine the correlation coefficient between each non-meteorological factor data and the forecast error sequence. The feature set of the initial forecasting model is updated based on the correlation coefficient, including: In the rolling forecasting process, the predicted load value for each forecasting period is subtracted from the corresponding actual total load to obtain the forecasting error for that forecasting period. The forecasting errors of multiple consecutive forecasting periods are then combined in chronological order to form a forecasting error sequence. When the cumulative error of the prediction error sequence exceeds a preset threshold, the prediction error sequence is used as a feedback signal to trigger a feedback evaluation of the feature set of the initial prediction model. In the feedback evaluation, the correlation coefficient between each non-meteorological factor data and the prediction error sequence is determined, and non-meteorological factor data with a correlation coefficient greater than or equal to a preset addition threshold are selected as new associated factors. Non-meteorological factor data with a correlation coefficient less than a preset removal threshold are identified from the used non-meteorological factor data as associated factors to be removed. The newly added related factors are added to the feature set, and the related factors to be removed are removed from the feature set to obtain the updated feature set.
[0067] Specifically, after the prediction model goes online and enters rolling prediction mode, each prediction cycle obtains the relevant factor data of the current cycle from the real-time data interface. After the above data preprocessing process, the meteorological factors are time-shifted and aligned, the basic load components are extracted, and the feature vectors are concatenated. The concatenated feature vector of the current cycle is sent to the trained prediction model. The model calculates and outputs the load prediction value of the current prediction cycle through forward propagation. After the power grid dispatch automation system collects the measured total load value of the cycle, the measured total load value is subtracted from the load prediction value to obtain the prediction error of the prediction cycle. The prediction error retains the sign to reflect the direction of the deviation. A fixed-length circular queue is maintained in the system memory as a sliding window. The window length is the number of all prediction cycles in a day. Each newly generated prediction error is pushed into the tail of the queue in chronological order. When the queue is full, the earliest error value at the head of the queue is automatically discarded, so that the queue always contains the prediction errors of the most recent consecutive prediction cycles. All error values in the queue are taken out in chronological order and combined to form a prediction error sequence.
[0068] Simultaneously, the cumulative error level of the prediction error sequence is monitored in real time. The root mean square error (RMSE) is used as a measure of cumulative error. The RMSE value at the current moment is obtained by squaring all error values in the prediction error sequence, taking the arithmetic mean, and then taking the square root. This RMSE value is compared with the system's preset error alarm threshold. This alarm threshold is set according to the power grid dispatch's business requirements for prediction accuracy, and is usually two to three times the historical benchmark prediction accuracy. When the RMSE value exceeds the preset threshold, it is determined that the input features of the current prediction model can no longer fully reflect the actual changes in the recent power grid operating environment, and a feedback evaluation process for the feature set of the initial prediction model is immediately triggered. If the RMSE value does not exceed the preset threshold, rolling prediction continues without triggering feedback evaluation.
[0069] After triggering the feedback assessment, the actual observation sequence of each non-meteorological factor data within the same continuous time period covered by the prediction error sequence is extracted from the database or memory. Non-meteorological factor data includes date-type coded sequences, time-of-use electricity price sequences, and labeled sequences for various special events such as temporary sporting events, staggered production, and sudden power rationing. Each non-meteorological factor observation sequence and prediction error sequence is treated as two sets of variables, and the Pearson correlation coefficient between them is calculated. Simultaneously, the significance test p-value of this correlation coefficient is calculated. After traversing all non-meteorological factors and completing the correlation coefficient calculation, the absolute value of the correlation coefficient is compared with the coefficient... The system compares the newly added data with a preset threshold, which can be between 0.3 and 0.4, with a preferred value of 0.35. Non-meteorological factor data with an absolute correlation coefficient greater than or equal to the newly added threshold and a p-value less than 0.05 are selected and marked as newly added related factors. At the same time, the system compares the absolute correlation coefficients of non-meteorological factor data already existing in the current feature set with a preset removal threshold, which can be between 0.05 and 0.1, with a preferred value of 0.08. Non-meteorological factor data with an absolute correlation coefficient less than the removal threshold are selected and marked as related factors to be removed.
[0070] For example, suppose a regional power grid operates normally on summer weekdays. The initial feature set typically includes the following non-meteorological factors: date type (weekday / weekend), time-of-use pricing, and a "peak-shifting" event marker. On a certain day, a large-scale nighttime event is temporarily held in the region; this event information was not initially included in the feature set. During the event, the prediction model's forecasts for several consecutive periods begin to deviate significantly from the actual load, and the root mean square error of the prediction error sequence gradually increases, eventually exceeding a preset alarm threshold and triggering a feedback evaluation. The observation sequences and prediction error sequences of all non-meteorological factors within recent consecutive time periods (e.g., the last 96 periods, i.e., one day) are extracted from memory, and the correlation coefficients are calculated for each. The calculation results show that the correlation coefficient between the temporary event marker sequence and the prediction error sequence reaches 0.62, with a p-value of 0.001, which is much higher than the addition threshold of 0.3, indicating a significant statistical association between the event and the prediction deviation. The system marks this as a new associated factor. The correlation coefficient of the time-of-use electricity price sequence is 0.18, with a p-value of 0.04. Although there is some correlation, it is lower than the addition threshold and is not included. The peak-shifting production marker sequence has consistently been zero recently (no peak-shifting production events have occurred), and its correlation coefficient with the prediction error sequence is only 0.02, far lower than the removal threshold of 0.05. The system determines that this feature has lost its explanatory power for load changes under the current operating environment and marks it as an associated factor to be removed. The date type has a correlation coefficient of 0.35 and is retained. Therefore, temporary events are added to the feature set as new features, while peak-shifting production markers are removed from the feature set. The updated feature set can promptly reflect the new load disturbances brought about by the events and cleans up redundant features that are no longer effective under the current environment, ensuring that the model input features are always synchronized with the actual operating state of the power grid.
[0071] Based on the feedback evaluation results, an update operation is performed on the feature set of the initial prediction model. Each item in the list of newly added related factors is checked one by one. If the non-meteorological factor is not yet included in the input variable list of the current feature set, the variable name and coding rule corresponding to the non-meteorological factor are added to the feature set, and the corresponding feature dimension position is assigned to it when constructing the feature vector in the next prediction period. For each item in the list of related factors to be removed, the variable entry corresponding to the non-meteorological factor is found and deleted from the input variable list of the current feature set. The feature vector construction in subsequent prediction periods will no longer include this dimension, and the total number of dimensions of the input feature vector is reduced accordingly. This completes the update of the feature set, resulting in the updated feature set.
[0072] Optionally, when the cumulative error of the prediction error sequence exceeds a preset threshold, the prediction error sequence is used as a feedback signal to trigger a feedback evaluation of the feature set of the initial prediction model, including: The mean, variance, and recent trend slope of the prediction error sequence are determined, and the mean, variance, and recent trend slope are combined into an actual multidimensional error feature vector. The actual multidimensional error feature vector is compared with the preset multidimensional threshold vector dimension by dimension. When the values of at least two dimensions in the actual multidimensional error feature vector exceed the corresponding dimension thresholds in the preset multidimensional threshold vector, the cumulative error is determined to exceed the preset threshold. After determining that the cumulative error exceeds the preset threshold, the prediction error sequence is marked as the feedback signal; The evaluation depth level of the feedback evaluation is determined based on the excess magnitude of each dimension in the multidimensional error feature vector, wherein the evaluation depth level is used to control the degree of screening of subsequent correlation coefficients. Using the evaluation depth level and the prediction error sequence as input, a feedback evaluation of the feature set of the initial prediction model is triggered.
[0073] Specifically, the currently saved complete prediction error sequence is retrieved from the sliding window. This sequence contains prediction error values from the most recent N consecutive prediction periods. First, the arithmetic mean of the prediction error sequence is calculated by summing all error values in the sequence and dividing by the sequence length N to obtain the mean. Second, the variance of the prediction error sequence is calculated by first obtaining the deviation of each error value from the mean, then summing the squares of each deviation and dividing by N to obtain the variance. Then, the most recent M error values from the end of the prediction error sequence are extracted, where M is between 1 / 3 and 1 / 2 of N. These M error values are fitted with a univariate linear regression in chronological order, with the time sequence as the independent variable and the error value as the dependent variable. The slope of the regression line is estimated using the least squares method, and this slope is taken as the recent trend slope. A positive slope indicates that the prediction deviation continues to expand, while a negative slope indicates that the prediction deviation gradually converges. The calculated mean, variance, and recent trend slope are combined in sequence to form a three-dimensional actual multidimensional error feature vector. It should be noted that the recent trend slope is estimated using the least squares method to obtain the slope of the regression line obtained by fitting the most recent M error values from the end of the prediction error sequence in chronological order, with the time sequence number as the independent variable and the error value as the dependent variable. A positive slope indicates that the prediction deviation is continuously increasing (the prediction error is showing an increasing trend), while a negative slope indicates that the prediction deviation is gradually converging (the prediction error is showing a decreasing trend). This slope value is dynamically recalculated with each prediction period's error update and is not a pre-set fixed parameter. For example, assuming the current sliding window length is 96 periods and M is 48, the prediction error values of the most recent 48 periods are extracted, arranged by time sequence number 1 to 48, and a straight line is fitted using the least squares method with the sequence number as the x-axis and the error value as the y-axis. The slope of this line is the recent trend slope. If the slope is positive, it indicates that the overall prediction error has been increasing in the positive direction recently, and the model's predicted value is consistently low; if the slope is negative, it indicates that the prediction error is decreasing or converging in the negative direction, and the model's predicted value is gradually approaching the actual value. The slope value, together with the mean and variance, constitutes a three-dimensional error feature vector in the feedback evaluation trigger condition determination. The mean reflects the overall deviation of the error, the variance reflects the fluctuation range of the error, and the recent trend slope reflects the trend of error change. The three dimensions characterize the overall state of the prediction error from different perspectives.Furthermore, a preset multidimensional threshold vector is read from the system configuration parameters. This preset multidimensional threshold vector is also a three-dimensional vector. Its first dimension is the mean threshold, which is twice the mean of the historical normal prediction error; the second dimension is the variance threshold, which is three times the variance of the historical normal prediction error; and the third dimension is the trend slope threshold, which is 2.5 times the absolute value of the trend slope of the historical normal prediction error. The mean, variance, and recent trend slope in the actual multidimensional error feature vector are compared one by one with the corresponding thresholds in the preset multidimensional threshold vector. It is recorded whether each dimension exceeds its corresponding threshold. When the actual values of at least two of the three dimensions are greater than or equal to the preset threshold, it is determined that the cumulative error of the current prediction error sequence has exceeded the preset threshold, and the subsequent feedback evaluation process is triggered. If the number of dimensions exceeding the threshold is less than two, it is determined that the cumulative error has not yet exceeded the preset threshold, and the feedback evaluation is not triggered.
[0074] Once the cumulative error exceeds the preset threshold, the data copy of the entire prediction error sequence saved in the current sliding window is immediately timestamped and encapsulated into a feedback signal message. This message contains the start time, end time, sequence length, and all error values of the error sequence. The feedback signal message is then sent to the feedback evaluation processing queue, awaiting reading and processing by the subsequent correlation coefficient calculation module.
[0075] While generating feedback signals, the assessment depth level of this feedback assessment is determined based on the magnitude by which the values of each dimension in the actual multidimensional error feature vector exceed the preset threshold. Specifically, the following methods are used: calculate the magnitude by which the mean exceeds the threshold (actual mean divided by the mean threshold), the magnitude by which the variance exceeds the threshold (actual variance divided by the variance threshold), and the magnitude by which the trend slope exceeds the threshold (absolute value of the actual slope divided by the trend slope threshold). Take the maximum value among these three magnitudes. If the maximum value is less than 1.5, the assessment depth level is set to level 1; if the maximum value is between 1.5 and 2.5, the assessment depth level is set to level 2; if the maximum value is greater than 2.5, the assessment depth level is set to level 3. The higher the assessment depth level, the more serious the recent prediction deviation and the more comprehensive and in-depth the investigation of influencing factors is required.
[0076] The evaluation depth level is associated with and bound to the previously encapsulated feedback signal message, and sent together to the feedback evaluation module. The evaluation depth level is used to dynamically adjust the screening threshold in the subsequent correlation coefficient calculation. When the level is 1, the new threshold remains unchanged; when the level is 2, the new threshold is adjusted by 0.1 above or below the preset base value, and the removal threshold is adjusted by 0.2 above or below the preset base value to include more possible related factors for consideration; when the level is 3, the new threshold is adjusted by 0.2 above or below the preset base value, and the removal threshold is lowered by 0.05 to achieve the widest range of factor screening and ensure that no potential new influencing factors are missed. Using this evaluation depth level and the prediction error sequence as joint input, the feedback evaluation process for the feature set of the initial prediction model is initiated.
[0077] In this optional embodiment, a multidimensional error feature vector is constructed by calculating the mean, variance, and recent trend slope of the prediction error sequence. This comprehensively characterizes the degradation state of prediction performance from three dimensions: deviation magnitude, fluctuation degree, and persistence trend, avoiding potential omissions or misjudgments that may occur when evaluating based on a single indicator. The actual multidimensional error feature vector is compared dimension by dimension with a preset multidimensional threshold vector, and the cumulative error is determined to exceed the preset threshold only when at least two dimensions exceed the limit simultaneously. This significantly enhances the stability and anti-interference capability of the triggering mechanism and prevents unnecessary frequent evaluations caused by accidental error spikes. The evaluation depth level is determined based on the exceedance magnitude of each dimension, enabling the feedback evaluation to adaptively adjust the stringency standard of subsequent correlation coefficient screening according to the severity of prediction degradation. In cases of slight degradation, screening remains cautious to maintain model stability, while in cases of severe degradation, the threshold is relaxed to comprehensively investigate potential new factors. This achieves hierarchical, flexible, and adaptive feature update control that matches the degree of error, balancing the prediction system's rapid response to sudden changes with robust maintenance of stable operation.
[0078] Optionally, the step of fine-tuning the parameters of the prediction model using the updated feature set, inputting the correlation factor data for the next prediction period into the fine-tuned prediction model, and outputting the load forecast value for the next prediction period includes: Based on the updated feature set, extract the values of the corresponding features from the historical load data to construct a fine-tuning sample set; Using the current parameters of the prediction model as initial parameters, the prediction model is incrementally trained using the fine-tuning sample set to obtain the prediction model with fine-tuned parameters. Obtain the correlation factor data for the next forecast period, and filter the corresponding feature values from the correlation factor data according to the updated feature set, input the fine-tuned forecast model with the parameters, and output the load forecast value for the next forecast period.
[0079] Specifically, after updating the feature set, data within the most recent preset time window is re-extracted from the historical database. The length of the time window is set to be 7 to 30 days in the past to cover the recent complete operating cycle. Based on all the input variable names listed in the updated feature set, the historical numerical sequences of the corresponding variables within the time window are read one by one from the database, including time-shifted and aligned meteorological factor data, basic load components, and updated non-meteorological factor data. The read variable sequences are aligned along a unified time axis and concatenated column by column to form a fine-tuning input feature matrix. At the same time, the measured total load value within the same time window is extracted as the fine-tuning output target sequence. The fine-tuning input feature matrix and the fine-tuning output target sequence are matched row by row to form a fine-tuning sample set. The number of samples in the fine-tuning sample set is equal to the total number of valid sampling points within the time window.
[0080] All parameters of the trained prediction model, including the weight matrices and bias vectors of the two long short-term memory hidden layers and the weights and biases of the fully connected layer, are loaded into memory as initial parameters for fine-tuning. The model network structure remains unchanged. The fine-tuning sample set is divided into a training set and a validation set in chronological order, with the training set comprising the first 85%. An adaptive moment estimation optimizer is used, with the initial learning rate set to 1 / 10 of the initial learning rate during pre-training (0.0001). This optimization is applied only to the parameters of the fully connected layer on the output side and the last long short-term memory hidden layer. Gradient updates are performed on the parameters of the hidden layer. All parameters of the first long short-term memory hidden layer are frozen so that they do not participate in the current training round. The mean squared error is used as the loss function. The model is trained incrementally in small batches using the fine-tuning training set, with each batch containing 32 samples. Each iteration of the complete fine-tuning training set is considered one round. After each round of training, the validation loss is calculated on the fine-tuning validation set. When the fine-tuning validation loss no longer decreases within 5 consecutive rounds or the number of training rounds reaches 20 rounds, the incremental training is terminated. The model parameters with the minimum fine-tuning validation loss are saved and used to overwrite the original model parameters to obtain the prediction model after parameter fine-tuning. After completing parameter fine-tuning, the original related factor data for the next forecast period is obtained from the real-time data interface. This includes the forecast values of each meteorological element for the next period obtained from the numerical weather prediction interface, the time-of-use electricity price for the next period obtained from the power market system, and the special event marking information for the next period obtained from the dispatch and operation management system. According to the input variable list specified by the updated feature set, the corresponding data items are extracted one by one. The meteorological factor data is time-shifted and aligned according to their respective effective lag times. The categorical variables in the non-meteorological factor data are converted by one-hot encoding. All processed feature values are concatenated into a feature vector that is completely consistent with the dimension of the input layer of the fine-tuned model according to the fixed feature arrangement order during training. This feature vector is input into the prediction model after parameter fine-tuning. The model calculates layer by layer through forward propagation. The fully connected output layer outputs a scalar value, which is the load forecast value for the next forecast period. This load forecast value is sent to the power grid dispatch automation system for real-time dispatch and safety control.
[0081] In this optional embodiment, a fine-tuning sample set is constructed by extracting corresponding feature values from recent historical data based on the updated feature set. This ensures that the fine-tuned data strictly matches the updated feature dimensions, avoiding training failure caused by inconsistencies between old and new features. Incremental training is performed using the current parameters of the prediction model as initial parameters. Only the layer parameters closest to the output side are updated with a small learning rate while the underlying temporal feature extraction layer is frozen. This allows the model to quickly adapt to changes in the input structure caused by the addition or subtraction of features, while preserving the general load patterns and temporal sequences already learned by the model. This achieves rapid response to newly included related factors and prevents the loss of original predictive ability due to significant parameter adjustments. After obtaining the related factor data for the next prediction period, feature selection and time-shift alignment are strictly performed according to the updated feature set. The fine-tuned model is then input to output the load prediction value, ensuring that feature engineering and model state remain synchronized throughout the entire rolling prediction process. Once the feature set changes, the model can quickly adapt based on recent data and immediately produce the prediction value for the next period, ensuring the real-time continuity of online prediction.
[0082] Combination Figure 2 As shown, this embodiment of the invention also provides a power grid time-varying load ultra-short-term forecasting system, comprising: The data decomposition unit is used to acquire historical load data and corresponding related factor data, and to decompose the historical load data into components to obtain the basic load component and the meteorological sensitive load component. The related factor data includes meteorological factor data and non-meteorological factor data. The time delay analysis unit is used to calculate the time delay mutual information based on the meteorological sensitive load component and each meteorological factor data respectively, and to determine the effective lag time corresponding to each meteorological factor data respectively. The model training unit is used to perform time-shift alignment on the corresponding meteorological factor data according to the effective lag time, take each of the time-shift aligned meteorological factor data, the base load component, and the non-meteorological factor data as input features, take the total load in the historical load data as the output target, construct a training sample set, and use the training sample set to pre-train the initial prediction model to obtain the trained prediction model. The prediction and feedback unit is used to input the relevant factor data of the current prediction period into the prediction model during the rolling prediction process based on the prediction model, obtain the load prediction value of the current prediction period and determine the prediction error, use the prediction error sequence accumulated over multiple consecutive prediction periods as a feedback signal, determine the correlation coefficient between each non-meteorological factor data and the prediction error sequence, and update the feature set of the initial prediction model according to the correlation coefficient. The fine-tuning prediction unit is used to fine-tune the parameters of the prediction model using the updated feature set, input the correlation factor data of the next prediction period into the parameter-fine-tuned prediction model, and output the load prediction value of the next prediction period.
[0083] The advantages of the power grid time-varying load ultra-short-term prediction system of the present invention compared with the prior art are the same as the advantages of the above-mentioned power grid time-varying load ultra-short-term prediction method compared with the prior art, and will not be repeated here.
[0084] While the present invention has been disclosed above, its scope of protection is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention, and all such changes and modifications will fall within the scope of protection of the present invention.
Claims
1. A method for ultra-short-term forecasting of time-varying load in power grids, characterized in that, include: Historical load data and corresponding related factor data are acquired, and the historical load data is decomposed into components to obtain the basic load component and the meteorological sensitive load component. The related factor data includes meteorological factor data and non-meteorological factor data. Based on the time delay mutual information calculation of the meteorological sensitive load component and each meteorological factor data, the effective lag time corresponding to each meteorological factor data is determined. Based on the effective lag time, the corresponding meteorological factor data is time-shifted and aligned. Each time-shifted meteorological factor data, the base load component, and the non-meteorological factor data are used as input features. The total load in the historical load data is used as the output target. A training sample set is constructed, and the initial prediction model is pre-trained using the training sample set to obtain the trained prediction model. In the rolling forecasting process based on the forecasting model, the relevant factor data of the current forecasting period is input into the forecasting model to obtain the load forecast value of the current forecasting period and determine the forecasting error. The cumulative forecasting error sequence of multiple consecutive forecasting periods is used as a feedback signal to determine the correlation coefficient between each non-meteorological factor data and the forecasting error sequence. The feature set of the initial forecasting model is updated according to the correlation coefficient. The updated feature set is used to fine-tune the parameters of the prediction model, and the correlation factor data for the next prediction period is input into the fine-tuned prediction model to output the load prediction value for the next prediction period.
2. The method for ultra-short-term forecasting of time-varying load in power grids according to claim 1, characterized in that, The process of acquiring historical load data and corresponding related factor data, and performing component decomposition on the historical load data to obtain the basic load component and the meteorologically sensitive load component, includes: The historical load data is decomposed into time series to extract the trend component and periodic component of the load. The trend component and the periodic component are superimposed to obtain the basic load component. Based on the historical load data and the basic load components, the remaining load sequence is obtained; Based on the correlation coefficient between the remaining load sequence and each meteorological factor data, meteorological factor data with a correlation coefficient greater than a preset threshold are selected as sensitive meteorological factors; The remaining load sequence is regressed and fitted to the sensitive meteorological factors to obtain the meteorological sensitive load components.
3. The method for ultra-short-term forecasting of time-varying load in power grids according to claim 1, characterized in that, The step of calculating the time delay mutual information based on the meteorological sensitive load component and each meteorological factor data to determine the effective lag time corresponding to each meteorological factor data includes: The sequence composed of each meteorological factor data and the meteorological sensitive load component is used as the sequence to be analyzed. A preset number of candidate lag durations are set, and the meteorological factor data is time-shifted according to the candidate lag durations to obtain the time-shifted sequence corresponding to each candidate lag duration; For each time-shifted sequence, the mutual information value between the time-shifted sequence and the meteorological sensitive load component is determined, and a mutual information sequence is obtained based on the mutual information value corresponding to each candidate lag duration; The candidate lag duration corresponding to the maximum value of the mutual information value in the mutual information sequence is determined as the effective lag duration corresponding to the meteorological factor data.
4. The method for ultra-short-term forecasting of time-varying load in power grids according to claim 3, characterized in that, For each time-shifted sequence, determining the mutual information value between the time-shifted sequence and the meteorological sensitive load component, and obtaining a mutual information sequence based on the mutual information value corresponding to each candidate lag duration, includes: For each time-shifted sequence, joint probability density estimation and marginal probability density estimation are performed on the time-shifted sequence and the meteorological sensitive load component, and the mutual information value corresponding to the time-shifted sequence is determined based on the results of the joint probability density estimation and the results of the marginal probability density estimation. The mutual information values of each meteorological factor data under all candidate lag times are arranged in order of lag time to obtain the initial mutual information sequence corresponding to the meteorological factor data, and the initial mutual information sequence is smoothed to obtain a smoothed mutual information sequence. In the smooth mutual information sequence, the search proceeds from the zero lag position in the direction of increasing lag duration to determine the lag duration corresponding to the first peak, and it is determined whether there is a significant decrease after the first peak. If there is a significant decrease, the lag duration corresponding to the first peak is determined as the effective lag duration. The condition for determining a significant decrease is that the mutual information value at the k-th lag point after the first peak is lower than the product of the mutual information value of the first peak and a preset proportional coefficient, where k is a preset search step size.
5. The method for ultra-short-term forecasting of time-varying load in power grids according to claim 3, characterized in that, The process involves time-shifting the corresponding meteorological factor data according to the effective lag duration, using each time-shifted meteorological factor data, the base load component, and the non-meteorological factor data as input features, and using the total load in the historical load data as the output target to construct a training sample set. This training sample set is then used to pre-train the initial prediction model to obtain the trained prediction model. The process includes: For each historical moment, based on the effective lag time corresponding to each meteorological factor data, the original meteorological factor data of the earlier moment that is separated from the historical moment by the effective lag time is extracted, and used as the time-shifted aligned meteorological factor data of the historical moment; The time-shifted aligned meteorological factor data, the base load component of the historical time, and the non-meteorological factor data are combined to form the input feature vector of the historical time, and the total load of the historical time is used as the output target to form an input-output sample pair; The input-output sample pairs from all historical moments are aggregated to construct the training sample set; The initial prediction model is trained using the training sample set. The parameters of the initial prediction model are iteratively updated by minimizing the error between the load prediction value output by the initial prediction model and the total load until the preset training conditions are met, thus obtaining the trained prediction model.
6. The method for ultra-short-term forecasting of time-varying load in power grids according to claim 5, characterized in that, The process of training the initial prediction model using the training sample set, iteratively updating the parameters of the initial prediction model by minimizing the error between the load prediction value output by the initial prediction model and the total load, until a preset training condition is met, to obtain the trained prediction model, includes: Each of the input feature vectors in the training sample set is input into the initial prediction model to obtain the corresponding load prediction value; Based on the predicted load value and the corresponding total load, the prediction residual is determined, and a loss function is constructed using the sum of squares of the prediction residuals; The parameters of the initial prediction model are backpropagated using the loss function to determine the gradient of each parameter, and the parameters are iteratively updated based on the gradient. After each iteration update, the loss value of the loss function under the current parameters is determined. When the change in the loss value between two adjacent iterations is less than the preset convergence threshold, it is determined that the preset training condition is met, and the prediction model under the current parameters is taken as the prediction model that has been trained.
7. The method for ultra-short-term forecasting of time-varying load in power grids according to claim 1, characterized in that, In the rolling forecasting process based on the forecasting model, the relevant factor data for the current forecast period are input into the forecasting model to obtain the load forecast value for the current forecast period and determine the forecast error. The cumulative forecast error sequence over multiple consecutive forecast periods is used as a feedback signal to determine the correlation coefficient between each non-meteorological factor data and the forecast error sequence. The feature set of the initial forecasting model is updated based on the correlation coefficient, including: In the rolling forecasting process, the predicted load value for each forecasting period is subtracted from the corresponding actual total load to obtain the forecasting error for that forecasting period. The forecasting errors of multiple consecutive forecasting periods are then combined in chronological order to form a forecasting error sequence. When the cumulative error of the prediction error sequence exceeds a preset threshold, the prediction error sequence is used as a feedback signal to trigger a feedback evaluation of the feature set of the initial prediction model. In the feedback evaluation, the correlation coefficient between each non-meteorological factor data and the prediction error sequence is determined, and non-meteorological factor data with a correlation coefficient greater than or equal to a preset addition threshold are selected as new associated factors. Non-meteorological factor data with a correlation coefficient less than a preset removal threshold are identified from the used non-meteorological factor data as associated factors to be removed. The newly added related factors are added to the feature set, and the related factors to be removed are removed from the feature set to obtain the updated feature set.
8. The method for ultra-short-term forecasting of time-varying load in power grids according to claim 7, characterized in that, When the cumulative error of the prediction error sequence exceeds a preset threshold, the prediction error sequence is used as a feedback signal to trigger a feedback evaluation of the feature set of the initial prediction model, including: The mean, variance, and recent trend slope of the prediction error sequence are determined, and the mean, variance, and recent trend slope are combined into an actual multidimensional error feature vector. The actual multidimensional error feature vector is compared with the preset multidimensional threshold vector dimension by dimension. When the values of at least two dimensions in the actual multidimensional error feature vector exceed the corresponding dimension thresholds in the preset multidimensional threshold vector, the cumulative error is determined to exceed the preset threshold. After determining that the cumulative error exceeds the preset threshold, the prediction error sequence is marked as the feedback signal; The evaluation depth level of the feedback evaluation is determined based on the excess magnitude of each dimension in the multidimensional error feature vector, wherein the evaluation depth level is used to control the degree of screening of subsequent correlation coefficients. Using the evaluation depth level and the prediction error sequence as input, a feedback evaluation of the feature set of the initial prediction model is triggered.
9. The method for ultra-short-term forecasting of time-varying load in power grids according to claim 1, characterized in that, The step of fine-tuning the prediction model using the updated feature set, inputting the correlation factor data for the next prediction period into the fine-tuned prediction model, and outputting the load forecast value for the next prediction period includes: Based on the updated feature set, extract the values of the corresponding features from the historical load data to construct a fine-tuning sample set; Using the current parameters of the prediction model as initial parameters, the prediction model is incrementally trained using the fine-tuning sample set to obtain the prediction model with fine-tuned parameters. Obtain the correlation factor data for the next forecast period, and filter the corresponding feature values from the correlation factor data according to the updated feature set, input the forecast model after parameter fine-tuning, and output the load forecast value for the next forecast period.
10. A power grid time-varying load ultra-short-term forecasting system, characterized in that, include: The data decomposition unit is used to acquire historical load data and corresponding related factor data, and to decompose the historical load data into components to obtain the basic load component and the meteorological sensitive load component. The related factor data includes meteorological factor data and non-meteorological factor data. The time delay analysis unit is used to calculate the time delay mutual information based on the meteorological sensitive load component and each meteorological factor data respectively, and to determine the effective lag time corresponding to each meteorological factor data respectively. The model training unit is used to perform time-shift alignment on the corresponding meteorological factor data according to the effective lag time, take each of the time-shift aligned meteorological factor data, the base load component, and the non-meteorological factor data as input features, take the total load in the historical load data as the output target, construct a training sample set, and use the training sample set to pre-train the initial prediction model to obtain the trained prediction model. The prediction and feedback unit is used to input the relevant factor data of the current prediction period into the prediction model during the rolling prediction process based on the prediction model, obtain the load prediction value of the current prediction period and determine the prediction error, use the prediction error sequence accumulated over multiple consecutive prediction periods as a feedback signal, determine the correlation coefficient between each non-meteorological factor data and the prediction error sequence, and update the feature set of the initial prediction model according to the correlation coefficient. The fine-tuning prediction unit is used to fine-tune the parameters of the prediction model using the updated feature set, input the correlation factor data of the next prediction period into the parameter-fine-tuned prediction model, and output the load prediction value of the next prediction period.
Citation Information
Patent Citations
Short-term power load prediction method and system considering multi-dimensional feature extraction
CN119448225A
Meteorology sensitive load power estimation method and apparatus
US20190384879A1