A method and system for harmonizing loads for machine learning predictions at arbitrary time scales
By constructing a multimodal data feature library and a dimensional separation model, and combining decision tree and SHAP value analysis, the accuracy problem of load forecasting at any time scale was solved, achieving efficient and interpretable forecast results and supporting risk management in the electricity market.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUORUI NEW ENERGY (GUANGZHOU) CO LTD
- Filing Date
- 2025-06-25
- Publication Date
- 2026-04-14
AI Technical Summary
Existing load forecasting methods have shortcomings in terms of accuracy and adaptability to arbitrary time scales. In particular, they are difficult to meet the accuracy requirements of medium- and long-term and spot transactions in the scenario of electricity sales companies. Furthermore, traditional methods cannot effectively cover changes at all time scales.
A multimodal data feature library is constructed. Combining the physical logic of load change and prediction algorithms, a dimensional separation model and a decision tree model are used for training and optimization. The feature contribution is analyzed by SHAP value. Similar day matching and dimensional fusion are performed to generate load prediction results at any time scale. The results are verified by backtesting with actual data and incremental learning to form a closed-loop optimization.
It achieves accurate forecasting of centrally dispatched loads, provides efficient and interpretable forecasting results at any time scale, supports risk management and strategy formulation in electricity market transactions, reduces memory consumption, and is suitable for rapid deployment and iteration.
Smart Images

Figure CN120764752B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of electricity trading, and in particular to a machine learning prediction method and system for loads under unified dispatch at arbitrary time scales. Background Technology
[0002] Currently, there are various known methods for forecasting central dispatch load, with numerous application scenarios. These methods are primarily used on the grid and generation sides, while fewer methods are based on the scenario of electricity sales companies. As the main agents for electricity users, electricity sales companies are naturally sensitive to changes in electricity market prices. Central dispatch load, as demand, has a significant impact on the supply-demand relationship, and its fluctuations often directly cause fluctuations in electricity prices. Therefore, electricity sales companies are also highly sensitive to changes in central dispatch load. At the same time, electricity sales companies participate in electricity market transactions. Whether in the medium- and long-term market, such as annual long-term contracts, monthly medium- and long-term contracts, or even multi-day and day-ahead contracts, they need to have a qualitative and quantitative understanding of the macro-level central dispatch load and make corresponding decision-making judgments based on the available data to avoid potential risks.
[0003] Common forecasting methods are basically based on time-series electricity data, at most incorporating some basic weather indicators such as temperature and humidity, and then using traditional statistical regression, or directly using machine learning, artificial neural networks, etc. to predict data from data, performing end-to-end load forecasting. The forecast results and accuracy deviate significantly from the actual situation, and traditional forecasting methods cannot cover all time scales.
[0004] This invention, based on big data analysis of relevant change characteristics in real-world power sales company application scenarios, builds a richer multimodal data feature library, constructs a complete system for unified dispatch load feature engineering, and ultimately combines the physical logic of load changes and prediction algorithms to more accurately predict the unified dispatch load at any time scale. Summary of the Invention
[0005] This invention provides a machine learning prediction method and system for load regulation at arbitrary time scales, which addresses the accuracy requirements of boundary condition analysis in medium- and long-term and spot transactions.
[0006] On one hand, the present invention provides a machine learning prediction method for regulated loads at arbitrary time scales, comprising:
[0007] Historical load, meteorological, and economic data are acquired, a feature library is constructed, and key features are selected and preprocessed to obtain preprocessed data.
[0008] Based on the preprocessed data, a dimension separation model is constructed and trained and optimized to obtain a decision tree model;
[0009] Based on the decision tree model, the SHAP value of each feature is calculated, and the contribution of the global features is analyzed to obtain the analysis results.
[0010] Based on the analysis results, similar day matching and dimensional fusion are performed to obtain the fused load prediction results;
[0011] Based on the fused load forecast results, multi-time-scale hierarchical forecasting is performed to generate load curves at arbitrary scales.
[0012] Based on load curves of arbitrary scale, prediction deviations are verified through backtesting with actual data, data is monitored and incremental learning is triggered to form a closed-loop optimization.
[0013] Furthermore, historical load, meteorological, and economic data are acquired to construct a feature library. Key features are then selected and preprocessed to obtain preprocessed data, including:
[0014] Historical load, meteorological, and economic data at a 15-minute granularity were acquired as raw data.
[0015] Based on the raw data, a three-layer feature library is constructed, which includes basic features, composite features, and scene features;
[0016] The three-layer feature library is subjected to a Filter-Wrapper hybrid algorithm. First, the correlation coefficient between features and load is calculated for initial screening. Then, the feature subset is optimized by recursive feature elimination combined with the SVR model to obtain the key features to be screened.
[0017] Calculate the correlation coefficient between key features and load, as well as the contribution of the SHAP value, to obtain preprocessed data.
[0018] Furthermore, based on the preprocessed data, a dimension separation model is constructed and trained and optimized to obtain a decision tree model, including:
[0019] Based on the preprocessed data, a load forecasting formula is constructed by combining the concept of dimensional separation, and the basic load is decoupled from meteorological and economic fluctuations to obtain a dimensional separation model.
[0020] A sliding window strategy is used to incrementally learn the dimension separation model to obtain the trained dimension separation model.
[0021] The key parameters of the trained dimensional separation model are optimized through grid search and cross-validation to obtain the optimized trained dimensional separation model.
[0022] Based on the optimized training dimension separation model, a decision tree model is obtained by combining the Leaf-wise tree growth strategy and the GOSS sampling method.
[0023] Furthermore, based on the decision tree model, the SHAP value of each feature is calculated, and the contribution of the global features is analyzed to obtain the analysis results, including:
[0024] Based on the decision tree model, the SHAP value of each feature sample is calculated to obtain the contribution of each feature sample to the prediction result.
[0025] Based on the contribution of each feature sample to the prediction results, the overall influence of the features is analyzed, and the key driving factors are identified, which are the analysis results.
[0026] Based on key driving factors, the daily load factor and peak-valley difference indicators of the forecast results are compared with historical statistical intervals to evaluate the rationality of the forecast and obtain the evaluation results.
[0027] Based on the evaluation results, when the key driving factors deviate from the historical range by more than a preset threshold, abnormal prediction results are automatically marked. For abnormal results, feature weights are adjusted first in combination with human experience. If the deviation continues, the incremental training process of the model is triggered.
[0028] Furthermore, based on the analysis results, similar day matching and dimensional fusion are performed to obtain the fused load forecast results, including:
[0029] Based on the analysis results, the matching factors for the target time period are calculated and a matching factor sequence is constructed. The matching factors include N-dimensional vectors of meteorological and economic characteristics, cosine similarity, typical day coefficient, and time decay weight.
[0030] Based on the matching factor sequence, grey relational analysis is performed using the grey relational algorithm. By multiplying the daily load by the predicted load for that day, the hourly load for that day can be obtained. Similarly, the predicted daily and monthly loads for the current month and year can be obtained, resulting in the fused load forecast.
[0031] Furthermore, based on the fused load forecast results, multi-time-scale hierarchical forecasting is performed to generate load curves at arbitrary scales, including:
[0032] Based on the fused load forecast results, the total annual load is predicted and decomposed into 12 monthly load curves through year-month matching to form an annual forecast framework.
[0033] Based on the annual forecasting framework, the average monthly load for the target month is predicted, and the daily load distribution is generated by matching the monthly and daily time groups. The daily and hourly time groups are then overlaid to refine the data to the hour level to obtain the monthly forecast results.
[0034] Based on the monthly forecast results, a refined load curve with a 15-minute granularity is output to obtain the daily forecast results;
[0035] By freely combining different time scales using a dimensional matching algorithm, load forecast curves with user-specified accuracy can be generated.
[0036] Furthermore, based on load curves at arbitrary scales, prediction biases are verified through backtesting with actual data, data is monitored and incremental learning is triggered to form a closed-loop optimization, including:
[0037] The load forecast results at any scale are compared with the actual operating data through backtesting, and the deviation values of key indicators are calculated to obtain the verification results.
[0038] Based on the validation results, key indicators are monitored in real time. When the limits are exceeded continuously, the incremental learning process is automatically triggered to dynamically update the model parameters.
[0039] On the other hand, a machine learning prediction system for regulated loads at arbitrary time scales includes:
[0040] The acquisition module is used to acquire historical load, meteorological and economic data, build a feature library, and filter key features and perform preprocessing to obtain preprocessed data.
[0041] The processing module is used to construct a dimensional separation model based on preprocessed data and train and optimize it to obtain a decision tree model; based on the decision tree model, it calculates the SHAP value of each feature and analyzes the contribution of global features to obtain analysis results; based on the analysis results, it performs similar day matching and dimensional fusion to obtain fused load prediction results; based on the fused load prediction results, it performs multi-time-scale hierarchical prediction to generate load curves of arbitrary scales; based on the load curves of arbitrary scales, it verifies the prediction deviation through backtesting with actual data, monitors the data and triggers incremental learning to form a closed-loop optimization.
[0042] On the other hand, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement a machine learning prediction method for arbitrary time scales of the overall load as described above.
[0043] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a machine learning prediction method for arbitrary time scales of the unified load as described above.
[0044] On the other hand, the present invention also provides a computer program product, including a computer program that, when executed by a processor, implements a machine learning prediction method for arbitrary time scales of the unified load as described above.
[0045] The present invention provides a machine learning method and system for predicting the load of centralized regulation at arbitrary time scales. While predicting the load, it can identify many relevant factors and sensitivities affecting the load. This facilitates the identification of significant directional trends in the load when important boundary conditions change, such as abrupt climate changes or the La Niña / El Niño transition. It can also predict maximum and minimum values based on extreme values. With accurate boundary / confidence intervals, risk boundaries in trading are established, making trading behavior and strategies more controllable. Furthermore, the core prediction principle of "dimension separation" is relatively simple, the machine learning algorithm used is sufficiently efficient, and it consumes little memory, making it suitable for rapid deployment and iteration. Attached Figure Description
[0046] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0047] Figure 1 This is a flowchart illustrating the machine learning prediction method for the unified dispatch load at any time scale provided in this embodiment of the invention.
[0048] Figure 2 This is a schematic diagram of a machine learning prediction system for the unified load at arbitrary time scales provided in an embodiment of the present invention.
[0049] Figure 3 This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation
[0050] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0051] Figure 1 This is one of the flowcharts illustrating the machine learning prediction method for the unified load at any time scale provided in this embodiment of the invention.
[0052] like Figure 1 As shown in the figure, the machine learning prediction method for the unified dispatch load at any time scale provided in this embodiment of the invention mainly includes the following steps:
[0053] 11. Obtain historical load, meteorological and economic data, construct a feature library, and screen key features and preprocess them to obtain preprocessed data;
[0054] 12. Based on the preprocessed data, construct a dimension separation model and train and optimize it to obtain a decision tree model;
[0055] 13. Based on the decision tree model, calculate the SHAP value of each feature and analyze the global feature contribution to obtain the analysis results;
[0056] 14. Based on the analysis results, perform similar day matching and dimensional fusion to obtain the fused load forecast results;
[0057] 15. Based on the fused load forecast results, perform multi-time-scale hierarchical forecasting to generate load curves at arbitrary scales;
[0058] 16. Based on load curves of arbitrary scale, backtesting with actual data verifies prediction deviations, monitors data and triggers incremental learning to form closed-loop optimization.
[0059] In this embodiment of the invention, the complete physical logic behind the unified dispatch load is first outlined. It represents the "load" situation across Guangdong Province, reflecting the sum of the actual electricity demand of all the smallest electricity-consuming units in the Guangdong power system, including the total power of various electrical equipment such as industrial, commercial, and residential units. Due to the real-time nature of unified dispatch and the accuracy of publicly disclosed information down to every 15 minutes, forecasts typically only involve a specific key time-series load indicator for a day, month, or even year, such as minimum load, maximum load, or average load. Since it reflects "demand," the load, regardless of the time scale, is a comprehensive reflection of electricity consumption behavior at that current time scale, and electricity consumption behavior is influenced by numerous factors. Therefore, to meet the forecasting needs at any time scale, the central idea is "division of dimensions, with the quantity being decomposed into a basic superposition of fluctuations, and the dimensions determined by similar behaviors." This involves attempting to separate a basic load from the centrally dispatched load, and then considering single or multiple factors such as meteorology and economic conditions that cause fluctuations in electricity load. Each influencing factor (hereinafter referred to as a feature) can be accurately quantified by studying its correlation and sensitivity with the centrally dispatched load. In forecasting, only meteorological data and some expected economic data, or other comprehensive indicators, are needed to calculate the future impact of these different features on the centrally dispatched load. The importance of these features can also be explained by the SHAP value theory of game theory, thus possessing complete interpretability and facilitating subsequent model optimization and review. This provides substantial assistance in decision-making regarding the centrally dispatched load and subsequent transactions.
[0060] like Figure 1As shown in Figure 11, historical load, meteorological, and economic data are acquired, a feature library is constructed, and key features are selected and preprocessed to obtain preprocessed data, including:
[0061] 111. Obtain historical load, meteorological, and economic data at a 15-minute granularity as raw data;
[0062] 112. Based on the original data, construct a three-layer feature library containing basic features, composite features, and scene features;
[0063] 113. The Filter-Wrapper hybrid algorithm is used for the three-layer feature library. First, the correlation coefficient between features and load is calculated for initial screening. Then, the feature subset is optimized by recursive feature elimination combined with the SVR model to obtain the key features to be screened.
[0064] 114. Calculate the correlation coefficient between key features and load, as well as the contribution of the SHAP value, to obtain the preprocessed data.
[0065] In this embodiment of the invention, historical data on power system load is collected, including: load values at different time granularities (such as hours, days, etc.), data on load-related influencing factors, meteorological data including temperature, humidity, wind speed, wind force level, wind direction, sunshine, cloud cover, atmospheric pressure, rainfall, etc. in various cities, date types including weekdays, weekends, holidays, etc., economic activity indicators including industrial output value, business activity index, etc., population data, load ratio of various regions, etc.
[0066] Delete samples with a large number of missing values, and use statistical analysis or domain knowledge to identify, correct or remove outlier data points; outlier screening is based on the 3σ principle, where σ represents the standard deviation and μ represents the mean. If the absolute distance between the data and the mean is greater than 3 times the standard deviation, i.e. the part of [-∞,μ-3σ] and [μ+3σ,+∞], this part of the value is an outlier. Outliers are filled by using the average electricity consumption of the same type of date.
[0067] Based on domain knowledge and data characteristics, useful features are extracted. These features include basic and composite features of time series, meteorology, and economy. Basic features come from the data that can be obtained in the raw data preparation stage, while composite features come from statistics and experience summaries.
[0068] For time series data, time features (hours, days of the week, months, etc.) and lag features (load values at several past time points) were extracted. Appropriate time window statistics were performed on meteorological data, including but not limited to separate labeling of structured time series data, and data cleaning, stratification, and cross-referencing.
[0069] The specific calculation method for feature indicators is as follows:
[0070] Historical data for all days is categorized into typical days and typical time periods. Typical days are divided into typical Mondays, typical weekdays, typical weekends, and different holidays such as Spring Festival, New Year's Day, and Qingming Festival. At the same time, different load levels are assigned to these typical days (load weights are calculated for typical days, and the load weights of weekends / Saturdays compared to weekdays are also calculated and labeled accordingly after statistics). These are used as inputs for categorical feature variables. Typical time periods are divided into three types according to the peak, flat, and valley periods in Guangdong. In addition, data cross-reconstruction features are performed between days by multiplying the daily load of the past week and the daily peak-valley difference rate of the past week to reconstruct a new time series indicator feature.
[0071] The time-series characteristics are divided into three categories: daily load characteristic indicators, monthly load characteristic indicators, and annual load characteristic indicators. Daily load characteristics include, but are not limited to: daily average load, daily maximum load, daily minimum load, and daily load rate. Monthly load characteristics include, but are not limited to: monthly maximum load, monthly average daily load, monthly maximum peak-to-valley difference, monthly average daily load rate, and cumulative maximum load utilization hours. Annual load characteristics include, but are not limited to: annual maximum load, annual average load, seasonal imbalance coefficient, and annual maximum load utilization hours. Meteorological characteristics include basic data such as temperature, humidity, wind speed, wind force level, wind direction, sunshine, cloud cover, atmospheric pressure, and rainfall. Based on these basic data... In addition to the existing features, artificial additions were made to include: perceived temperature, urban heat effect coefficient, 2-day accumulated temperature effect, 3-day accumulated temperature effect, and 4-day accumulated temperature effect; economic features included GDP, CPI, energy intensity, income from the primary, secondary, and tertiary industries, export investment, and national economic income; some complex composite features underwent special processing within the system, distinguishing between typical weekdays, weekends, and statutory and adjusted holidays in terms of time series; in terms of data structure, it directly bins and crosses some indicators, such as multiplying temperature and humidity, and then converting them into classification feature tracking; at the same time, it also aggregates temperature according to region, such as the regional light effect of light and the regional temperature effect of temperature;
[0072] After constructing the above features, the feature library is complete. The feature library will statistically analyze the changing relationships of each feature, such as the trends of hourly temperature and regulated load by day, month, and year, and create scatter plots. This allows for the rapid determination of correlation and sensitivity. Using piecewise linear regression, it automatically divides the data into 1, 2, 3, or even more linear variation intervals, quantifying the impact of temperature on the composite change in each interval. This process applies to both simple and composite features. Correlation is based on a Filter-Wrapper hybrid feature selection algorithm to select the optimal feature. The Filter part calculates the correlation coefficient between the feature and the load.
[0073] ;
[0074] Among them, represents the th value of the feature, X represents the average value of the feature, represents the th value of the power, Y represents the average load of the unified regulation, and irrelevant features are removed according to the magnitude of the correlation coefficient;
[0075] Wrapper part: First, a dataset with n features is obtained through Filter. Second, SVR is used as the base model of the recursive feature elimination method RFE. Third, the number of finally retained features k (k < n) is given. Finally, in the first round, all features are trained, and each feature is ranked according to the MSE index, and the feature with the smallest score is removed. At this time, the number of features is reduced to n - 1, and the iteration continues with n - 1 features until the number of retained features is k, that is, the features finally entering the model training;
[0076] Finally, a multi-dimensional table of features and load can be counted, and the table is as follows:
[0077] ;
[0078] Uncertain features mainly rely on some scenarios for tracking, and more are the processing of text information. Here, a standard data storage format is designed to convert unstructured data into structured data, thus realizing the use of the additional inference ability of the large model to achieve multi-modal input:
[0079] The specific feature storage table is as follows:
[0080] ;
[0081] With the tracking of these features, it is possible to know which features can be input into the prediction model as features, thus realizing the pre-screening of the input features of the model and avoiding wasting too much computing resources and information loss in the repeated algorithm iteration process.
[0082] As Figure 1 shown in
[0083] 121. Based on the preprocessed data, combine the idea of dimensional separation to construct a load prediction formula, and decouple and model the basic load from meteorological and economic fluctuations to obtain a dimensional separation model;
[0084] 122. Adopt a sliding window strategy for incremental learning of the dimensional separation model to obtain the trained dimensional separation model;
[0085] 123. Optimize the key parameters of the trained dimension separation model through grid search and cross-validation to obtain the optimized trained dimension separation model;
[0086] 124. Based on the optimized training dimension separation model, combined with the Leaf-wise tree growth strategy and the GOSS sampling method, a decision tree model is obtained.
[0087] In this embodiment of the invention, in a practical application scenario, new power system load and related factor data are continuously collected. When a certain amount of new data is accumulated, incremental learning is prepared. The collected data is divided into an initial training set and a test set. The initial training set is used to build the basic version of the model, and the test set is used to initially evaluate the model performance. During the training process, the idea of "reinforcement learning / lifelong learning" is used. By setting a step size, the time window is continuously slid. Each time, the training set is imported for training, and the test set is used to test the generalization ability, thereby improving the prediction accuracy. For example, if the sliding step size is set to 90 days and the validation set scale is 30 days, then the first step will be to fit the model. The first step uses data from January 1, 2020 to June 30, 2020. The second step involves fitting data from July 1, 2020 to July 30, 2020, and then iterating through the decision tree on the same data. The third step involves fitting data from April 1, 2020 to September 30, 2020, and then iterating through the same data from October 1, 2020 to October 30, 2020, until the entire dataset is slid through. This step-by-step, iterative process can support datasets of any size; the only difference is the training time. There's no need to worry about future data expansion, as parallel computing is highly efficient while maintaining accuracy.
[0088] The initial settings for the LightGBM model include tree depth (max_depth), learning rate (learning_rate), number of trees (num_leaves), objective function (objective), and number of iterations (num_iterations). These parameters were initially set based on past experience and a preliminary understanding of the data, and further adjustments are needed to optimize the parameters.
[0089] Grid search combined with cross-validation (such as k-fold cross-validation) was used to fine-tune the parameters of the LightGBM model on the initial training set. During parameter tuning, the key evaluation metric was Mean Absolute Percentage Error (MAPE). This metric is sensitive to relative error and does not change with global scaling of the target variable, making it suitable for problems with large differences in the dimensions of the target variable. The parameter combination that minimizes this metric was selected. The formula for calculating MAPE is:
[0090] ;
[0091] in, For the first The true load value (actual observation value) of each sample. For the first The predicted loading value (model output value) of each sample, where nsamples is the total number of samples (the total number of data points participating in the evaluation). For a minimal constant (e.g.) This is used to avoid the denominator being zero, ensuring numerical stability;
[0092] Using the initial training set data and the determined parameter settings, train the LightGBM model. During training, observe the changes in the model's training and validation losses to determine if the model is overfitting or underfitting, and further adjust the parameters or data preprocessing methods as needed. Use LightGBM's native commands for incremental learning, specifying the previously trained model file (e.g., the 'init_model' parameter) and the path to the new data (e.g., the 'data' parameter). Simultaneously, integrate the preprocessed new data with the previous training data, or directly use the new data as incremental data. Parameter tuning and training: Use the grid search method again to fine-tune the model parameters to adapt to the characteristics and changes of the new data.
[0093] The specific method of grid search is as follows:
[0094] The hyperparameters that need to be optimized and their candidate value ranges are determined based on the nature of the data type of the unified dispatch load, the size of the dataset, and existing experimental experience.
[0095] After defining the LightGBM model, the grid search is initialized using the publicly available GridSearchCV method, specifying the model, parameter grid, cross-validation folds, and scoring criteria.
[0096] The `fit` method is called to fit the training set. The model will iterate through all parameter combinations and perform cross-validation on each parameter set.
[0097] After the grid search is complete, the optimal parameter combination is obtained through the best_params_ attribute. Then, the model is reinitialized using the optimal parameters, and the entire training set is trained.
[0098] During training, parameters such as the learning rate can be adjusted appropriately to enable the model to learn information from new data better while retaining existing knowledge. Through iterative training, the model gradually updates and optimizes its predictive ability to adapt to the changing patterns of power system load.
[0099] Subsequently, the updated model performance is continuously re-evaluated on the test set. Error metrics (MAPE and other indicators) between predicted and actual load values are calculated and compared with previous model performance to verify whether incremental learning has improved the model's prediction accuracy. This checks the model's generalization ability on new data, ensuring that the model has not overfitted to the new data and lost its grasp of the overall data patterns. Finally, a final model is output.
[0100] like Figure 1 As shown in Figure 13, based on the decision tree model, calculate the SHAP value of each feature and analyze the global feature contribution to obtain the analysis results, including:
[0101] 131. Based on the decision tree model, calculate the SHAP value of each feature sample to obtain the contribution of each feature sample to the prediction result;
[0102] 132. Based on the contribution of each feature sample to the prediction results, analyze the overall influence of the features and identify the key driving factors, i.e., the analysis results;
[0103] 133. Based on key driving factors, compare the predicted daily load factor and peak-valley difference with historical statistical intervals to evaluate the rationality of the prediction and obtain the evaluation results.
[0104] 134. Based on the evaluation results, when the key driving factors deviate from the historical range by more than the preset threshold, abnormal prediction results are automatically marked. For abnormal results, feature weights are adjusted first in combination with human experience. If the deviation continues, the incremental training process of the model is triggered.
[0105] In this embodiment of the invention, the SHAP library is installed in the Python environment, and the SHAP interpreter is imported: TreeExplainer (an interpreter specifically designed for tree-based models such as LightGBM) is imported into the code; then, SHAP values are calculated: using the trained LightGBM model and interpreter, SHAP values are calculated for samples in the training or test set. The SHAP value represents the contribution of each feature to the model's prediction result, that is, the impact of the value of that feature on the model's output considering all other features; the formula for calculating the SHAP value is based on the concept of Shapley value, a concept in game theory used to allocate payoffs in cooperative games; in the context of machine learning model interpretation, the SHAP value is used to allocate the contribution of the model's prediction to each feature; for a model with n features, the SHAP value of the i-th feature is... It can be calculated using the following formula:
[0106] ;
[0107] Where N is the set of all features, S is a subset of N excluding feature i, |S| is the number of features in set S, f(S) is the model's prediction when considering only features in set S, and f(S∪{i}) is the model's prediction when considering both set S and feature i. The average marginal contribution of feature i in all possible feature combinations is calculated. However, this formula is computationally expensive in practical applications. Therefore, the SHAP library uses a more efficient algorithm to approximate the SHAP value. The SHAP value of the ensemble decision tree model used in this patent is calculated in this way.
[0108] By plotting feature importance plots (such as SHAP summary plots), we can show the overall impact of all features on the model's prediction results. The horizontal axis represents the absolute value of the SHAP value, the vertical axis represents the feature name, and the color represents the size of the feature value. This can help us understand which features are most critical to power system load forecasting and identify the main influencing factors.
[0109] For a specific prediction sample, a forceplot is plotted to visually demonstrate the direction (positive or negative) and magnitude of each feature's contribution to the prediction result of that sample. For example, for load prediction at a specific time point, we can see whether the increase in temperature contributes to increasing or decreasing the load prediction, and the degree of contribution, thus helping us understand the basis of the model's decision in this specific situation. It is also a means of verifying the reliability of the model. The direction in which the model interprets a certain feature should be consistent with the direction of change of the feature and the load in the past statistics.
[0110] Plotting a SHAPdependence plot shows the relationship between a specific feature and the model output. It also allows observation of the interaction between this feature and other features, helping to further understand the impact of complex relationships between features on load forecasting. For example, it can analyze the nonlinear relationship between temperature and load, and whether humidity plays a regulatory role. At the same time, the deviation correction model is imported. Through statistical calculation indicators, such as daily load rate and daily peak-to-valley difference, the degree of fit between the output results and historical data is tested. Under the condition that all boundary conditions are the same for historical and forecast data, the forecast results and indicators should match our historical data and indicators.
[0111] like Figure 1 As shown in Figure 14, based on the analysis results, similar day matching and dimensional fusion are performed to obtain the fused load prediction results, including:
[0112] 141. Based on the analysis results, calculate the matching factors for the target time period and construct a matching factor sequence. The matching factors include N-dimensional vectors of meteorological and economic characteristics, cosine similarity, typical day coefficient, and time decay weight.
[0113] 142. Based on the matching factor sequence, grey relational analysis is performed using the grey relational algorithm. By multiplying the daily load with the predicted load for the day, the hourly load for the day can be obtained. Similarly, the predicted daily and monthly loads for the month and year can be obtained, resulting in the fused load forecast.
[0114] In this embodiment of the invention, during the invention process, the prediction results of other models (such as linear regression, random forest, neural network, etc.) were compared, and it was found that based on the existing structured data, the LightGBM model showed stronger prediction accuracy and better performance than other models in power system load forecasting.
[0115] The calculation of the daily, hourly, monthly, and annual load distribution is performed, along with historical monthly maximum temperature, average temperature, and average rainfall (X_1, X_2, ..., X_n) to provide data for matching similar days. In the algorithm for similar days, a grey relational analysis is conducted on the sequence composed of three major factors: meteorological and economic similarity, typical days, and time series weights. This analysis identifies the most similar historical time series to the predicted time series. Grey relational analysis is primarily used to study systems where some information is explicit and others is implicit. The calculation of the correlation between the parent sequence (reference sequence) and the child sequence (comparison sequence) is also performed. The parent sequence can be considered the main characteristic sequence of the system's behavior. In studying the load distribution, the three major factor sequences of the predicted days can be used as the parent sequence. The child sequence is the sequence of other relevant factors compared with the parent sequence, i.e., the sequence composed of the three historical factors.
[0116] Suppose the parent sequence is X0={x0(1),x0(2),…,x0(n)} and the subsequence is Xi={xi(1),xi(2),…,xi(n)}(i=1,2,…,m), where n is the length of the sequence and m is the number of subsequences;
[0117] For each data in a sequence, divide each data in the sequence by the first data in that sequence; the processed parent sequence is X0′={x0′(1)=1,x0′(2)=x0(2) / x0(1),…,x0′(n)=x0(n) / x0(1)}, and the subsequences are processed in the same way; the purpose of this is to eliminate the influence of different dimensions on the correlation calculation;
[0118] Then the correlation coefficient was calculated:
[0119] ;
[0120] Where k=1,2,…,n, i=1,2,…,m, and ρ is the resolution coefficient, usually taken as 0.5; It is the minimum value among all differences between the parent sequence and the child sequence. It is the maximum value among all differences between the parent sequence and the child sequence; the correlation coefficient reflects the degree of correlation between the parent sequence and the child sequence at each corresponding point;
[0121] The final grey relational degree, i.e., the average of the relational coefficients, is:
[0122] ;
[0123] Where, r i This represents the degree of correlation between the parent sequence and the child sequence as a whole, where n is the grey correlation degree, i.e., the number of correlation coefficients. Grey relational degree, also known as the correlation coefficient, is a measure of the degree of correlation. The value of the correlation degree is generally between [0,1]. The higher the correlation degree, the stronger the correlation between the two sequences.
[0124] Cosine similarity is calculated for factor matching of meteorological and economic characteristics; the formula for calculating cosine similarity is as follows:
[0125] ;
[0126] Among them, meteorological characteristics such as temperature, humidity, perceived temperature, and atmospheric pressure are N-dimensional vectors A and B, which together with the N-dimensional vectors of meteorological characteristics at the prediction time can be calculated to form a cosine angle cos(A, B). The same applies to economic characteristics.
[0127] For matching typical day factors, a similarity coefficient dictionary is used between typical days. The default value is 0.15 for Monday, 0.2 for Tuesday to Thursday, 0.25 for Friday, 0.7 for Saturday, and 1.0 for Sunday. The formula for calculating the typical day matching coefficient is as follows:
[0128] in, and For the corresponding dictionary match value, Typical daily matching coefficient;
[0129] For time factor matching, a linear decay process is applied based on the distance between historical days and the predicted day. For example, the weight of the past day is 0.99, the weight of the past seven days is 0.9, and the weight of the past 365 days is 0.8. The principle of "nearer data has a larger weight, farther data has a smaller weight" is adopted, meaning that historical data closer to the predicted day has a larger weight. The periodicity of load changes and the impact of holidays are also considered. The specific calculation formula is as follows:
[0130] Where, δ equals t model power, t divided by power, t divided by The result of multiplying the three parts to the power of beta1 is 0.95, beta2 is 0.96, beta3 is 0.97, M1 is 7, M2 is 365, and M3 is 340.
[0131] The three coefficients form a sequence, and then the grey relational analysis is used to calculate the most similar date. By multiplying the day's index by the day's predicted quantity, the hourly load of that day can be obtained. Similarly, the predicted daily and monthly loads for the month and year can also be obtained.
[0132] like Figure 1 As shown in Figure 15, based on the fused load forecast results, multi-time-scale hierarchical forecasting is performed to generate load curves at arbitrary scales, including:
[0133] 151. Based on the fused load forecast results, the total annual load is predicted and decomposed into 12 monthly load curves through year-month matching to form an annual forecast framework.
[0134] 152. Based on the annual forecasting framework, the average monthly load for the target month is predicted. The daily load distribution is generated by matching the monthly and daily time groups. The daily and hourly time groups are then overlaid to refine the data to the hour level to obtain the monthly forecast results.
[0135] 153. Based on the monthly forecast results, output a refined load curve with a 15-minute granularity to obtain the daily forecast results;
[0136] 154. By freely combining different time scales through the dimensional matching algorithm, load forecast curves with user-specified accuracy can be generated.
[0137] In this embodiment of the invention, an annual load distribution benchmark is established by matching the annual total load forecast with the annual monthly load distribution. This provides a macro-level basis for long-term power planning (such as power generation dispatch and energy procurement), forming a stable monthly load allocation ratio and avoiding the problem of amplified cumulative errors caused by directly forecasting monthly data.
[0138] The monthly average load is dynamically adjusted within the annual framework. By combining historical daily load patterns (monthly-daily) and time-of-day fluctuation patterns (daily-hourly), the coherent forecasting of monthly-daily-hourly load is achieved, improving the timeliness of monthly forecasts. This is especially suitable for the impact of short-term events such as holidays or extreme weather, and supports monthly power trading and unit combination optimization.
[0139] The monthly forecast results directly generate 15-minute load curves to meet real-time scheduling requirements. Through high-precision time-division dimension matching, the forecast deviation during peak and valley periods is significantly reduced (MAPE<2%), which helps to smooth peaks and valleys and improve demand response.
[0140] It allows users to select the prediction scale as needed (e.g., quarterly → weekly → hourly), and dynamically synthesize the target curve through the dimensional matching algorithm, breaking the limitation of the fixed scale of the traditional model, realizing "one-time modeling, multi-scale output", and greatly reducing the coordination cost of cross-scale prediction;
[0141] By employing a top-down hierarchical forecasting approach based on "year → month → day" and dimensional fusion, the system ensures the stability of long-term trends while retaining the flexibility to handle short-term fluctuations, making it suitable for various scenarios such as electricity market transactions and power grid safety verification.
[0142] like Figure 1 As shown in Figure 16, based on load curves of arbitrary scale, prediction deviations are verified through backtesting with actual data, data is monitored and incremental learning is triggered to form a closed-loop optimization, including:
[0143] 161. Compare the load forecast results at any scale with the actual operating data through backtesting, calculate the deviation values of key indicators, and obtain the verification results;
[0144] 162. Based on the validation results, monitor key indicators in real time. When the limits are exceeded continuously, the incremental learning process is automatically triggered to complete the dynamic update of model parameters.
[0145] In this embodiment of the invention, the prediction results at different time scales (year / month / day) are compared with the actual operating data to calculate key indicators such as MAPE and daily load factor deviation, quantify the model prediction accuracy, comprehensively evaluate the model performance through multi-dimensional indicators (such as peak-valley deviation and trend consistency), identify systematic biases (such as long-term overestimation or insufficient capture of short-term fluctuations), and provide data support for optimization.
[0146] Real-time monitoring of key indicators (such as MAPE>5% for 3 consecutive times) automatically triggers the incremental learning mechanism, updates model parameters with the latest data, avoids delays caused by manual intervention, and the model continuously adapts to changes in load patterns (such as the deformation of the demand curve caused by the access of new energy sources), maintaining prediction accuracy. For sudden anomalies (such as extreme weather events), the model can self-correct within 48 hours, reducing prediction risks. Retraining is only initiated for periods when the deviation exceeds the limit, reducing redundant computing costs by more than 80%.
[0147] Through an automated closed loop of "prediction-validation-optimization", the model has the ability to continuously evolve, which is especially suitable for scenarios where load characteristics change rapidly during the power system transition period, and can reduce the overall prediction error in the long run.
[0148] like Figure 2 As shown, a machine learning prediction system 20 for load regulation at arbitrary time scales includes:
[0149] The acquisition module 21 is used to acquire historical load, meteorological and economic data, build a feature library, and screen key features and perform preprocessing to obtain preprocessed data.
[0150] Processing module 22 is used to construct a dimensional separation model based on preprocessed data and train and optimize it to obtain a decision tree model; based on the decision tree model, calculate the SHAP value of each feature and analyze the contribution of global features to obtain analysis results; based on the analysis results, perform similar day matching and dimensional fusion to obtain fused load prediction results; based on the fused load prediction results, perform multi-time-scale hierarchical prediction to generate load curves of arbitrary scales; based on the load curves of arbitrary scales, verify the prediction deviation through backtesting with actual data, monitor the data and trigger incremental learning to form a closed-loop optimization.
[0151] Figure 3 This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention.
[0152] like Figure 3 As shown, the electronic device may include a processor 610, a communications interface 620, a memory 630, and a communication bus 640. The processor 610, communications interface 620, and memory 630 communicate with each other via the communication bus 640. The processor 610 can call logical instructions from the memory 630 to execute machine learning prediction methods for arbitrary time scales under a unified load.
[0153] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0154] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to perform machine learning prediction methods for arbitrary time scales on the overall load provided by the above methods.
[0155] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a machine learning prediction method for arbitrary time scales to perform the overall load provided by the methods described above.
[0156] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0157] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0158] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A machine learning method for predicting loads at arbitrary time scales under unified regulation, characterized in that, include: Historical load, meteorological, and economic data are acquired, a feature library is constructed, and key features are selected and preprocessed to obtain preprocessed data. Based on the preprocessed data, a dimension separation model is constructed and trained and optimized to obtain a decision tree model; Based on the decision tree model, the SHAP value of each feature is calculated, and the contribution of the global features is analyzed to obtain the analysis results. Based on the analysis results, similar day matching and dimensional fusion are performed to obtain the fused load prediction results; Based on the fused load forecast results, multi-time-scale hierarchical forecasting is performed to generate load curves at arbitrary scales. Based on load curves of arbitrary scale, prediction deviations are verified through backtesting with actual data, data is monitored and incremental learning is triggered to form a closed-loop optimization.
2. The machine learning prediction method for centrally regulated load at arbitrary time scales according to claim 1, characterized in that, Historical load, meteorological, and economic data are acquired, a feature library is constructed, and key features are selected and preprocessed to obtain preprocessed data, including: Historical load, meteorological, and economic data at a 15-minute granularity were acquired as raw data. Based on the raw data, a three-layer feature library is constructed, which includes basic features, composite features, and scene features; The three-layer feature library is subjected to a Filter-Wrapper hybrid algorithm. First, the correlation coefficient between features and load is calculated for initial screening. Then, the feature subset is optimized by recursive feature elimination combined with the SVR model to obtain the key features to be screened. Calculate the correlation coefficient between key features and load, as well as the contribution of the SHAP value, to obtain preprocessed data.
3. The machine learning prediction method for centrally regulated load at arbitrary time scales according to claim 2, characterized in that, Based on the preprocessed data, a dimension separation model is constructed and trained and optimized to obtain a decision tree model, including: Based on the preprocessed data, a load forecasting formula is constructed by combining the concept of dimensional separation, and the basic load is decoupled from meteorological and economic fluctuations to obtain a dimensional separation model. A sliding window strategy is used to incrementally learn the dimension separation model to obtain the trained dimension separation model. The key parameters of the trained dimensional separation model are optimized through grid search and cross-validation to obtain the optimized trained dimensional separation model. Based on the optimized training dimension separation model, a decision tree model is obtained by combining the Leaf-wise tree growth strategy and the GOSS sampling method.
4. The machine learning prediction method for arbitrarily time-scaled load according to claim 3, characterized in that, Based on the decision tree model, the SHAP value of each feature is calculated, and the global feature contribution is analyzed to obtain the analysis results, including: Based on the decision tree model, the SHAP value of each feature sample is calculated to obtain the contribution of each feature sample to the prediction result. Based on the contribution of each feature sample to the prediction results, the overall influence of the features is analyzed, and the key driving factors are identified, which are the analysis results. Based on key driving factors, the daily load factor and peak-valley difference indicators of the forecast results are compared with historical statistical intervals to evaluate the rationality of the forecast and obtain the evaluation results. Based on the evaluation results, when the key driving factors deviate from the historical range by more than a preset threshold, abnormal prediction results are automatically marked. For abnormal results, feature weights are adjusted first in combination with human experience. If the deviation continues, the incremental training process of the model is triggered.
5. The machine learning prediction method for centrally regulated load at arbitrary time scales according to claim 4, characterized in that, Based on the analysis results, similar day matching and dimensional fusion are performed to obtain the fused load forecast results, including: Based on the analysis results, the matching factors for the target time period are calculated and a matching factor sequence is constructed. The matching factors include N-dimensional vectors of meteorological and economic characteristics, cosine similarity, typical day coefficient, and time decay weight. Based on the matching factor sequence, grey relational analysis is performed using the grey relational algorithm. By multiplying the daily load by the predicted load for that day, the hourly load for that day can be obtained. Similarly, the predicted daily and monthly loads for the current month and year can be obtained, resulting in the fused load forecast.
6. The machine learning prediction method for centrally regulated load at arbitrary time scales according to claim 5, characterized in that, Based on the fused load forecast results, multi-time-scale hierarchical forecasting is performed to generate load curves at arbitrary scales, including: Based on the fused load forecast results, the total annual load is predicted and decomposed into 12 monthly load curves through year-month matching to form an annual forecast framework. Based on the annual forecasting framework, the average monthly load for the target month is predicted, and the daily load distribution is generated by matching the monthly and daily time groups. The daily and hourly time groups are then overlaid to refine the data to the hour level to obtain the monthly forecast results. Based on the monthly forecast results, a refined load curve with a 15-minute granularity is output to obtain the daily forecast results; By freely combining different time scales using a dimensional matching algorithm, load forecast curves with user-specified accuracy can be generated.
7. The machine learning prediction method for centrally regulated load at arbitrary time scales according to claim 6, characterized in that, Based on load curves of arbitrary scale, prediction deviations are verified through backtesting with actual data, data is monitored and incremental learning is triggered to form a closed-loop optimization, including: The load forecast results at any scale are compared with the actual operating data through backtesting, and the deviation values of key indicators are calculated to obtain the verification results. Based on the validation results, key indicators are monitored in real time. When the limits are exceeded continuously, the incremental learning process is automatically triggered to dynamically update the model parameters.
8. A machine learning prediction system for regulated loads at arbitrary time scales, characterized in that, include: The acquisition module is used to acquire historical load, meteorological and economic data, build a feature library, and filter key features and perform preprocessing to obtain preprocessed data. The processing module is used to build a dimension separation model based on the preprocessed data and train and optimize it to obtain a decision tree model; Based on the decision tree model, the SHAP value of each feature is calculated, and the contribution of the global features is analyzed to obtain the analysis results. Based on the analysis results, similar day matching and dimensional fusion are performed to obtain the fused load prediction results; Based on the fused load forecast results, multi-time-scale hierarchical forecasting is performed to generate load curves at arbitrary scales. Based on the load curves at arbitrary scales, the forecast deviation is verified by backtesting with actual data, the data is monitored and incremental learning is triggered to form a closed-loop optimization.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the machine learning prediction method for the unified load at any time scale as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the machine learning prediction method for the unified load at any time scale as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Non-market electrical load dynamic prediction method and system based on dimension separation mechanism
CN120952448A
Port facility management and maintenance large model report review intelligent agent construction method and system
CN121581681A