Machine learning prediction method and system for tracking load at any time scale
By constructing a multimodal data feature library and a dimensional separation model, combined with decision trees and incremental learning, the shortcomings of unified load forecasting in terms of time scale and accuracy are solved, and accurate forecasting and risk management at any time scale are achieved, which is suitable for power market transactions.
Patent Information
- Application Number
- CN202510862037.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Existing unified load forecasting methods have deficiencies in accuracy and time scale, especially in the scenario of power sales companies, where it is difficult to encompass changes in all time scales. Traditional methods are also unable to effectively cope with price fluctuations and risk management in the electricity market.
A multimodal data feature library is constructed, combined with the physical logic of load changes and prediction algorithms, and a dimensional separation model and a decision tree model are adopted. The feature contribution is analyzed through SHAP values, similar day matching and dimensional fusion are performed to generate load curves of any time scale, and closed-loop optimization is formed through incremental learning.
It achieves accurate prediction of coordinated load, can accurately reflect electricity consumption behavior at any time scale, discover influencing factors, reduce memory consumption, is suitable for rapid deployment and iteration, and improves the controllability of trading strategies and risk management capabilities.
Smart Images

Figure CN120764752A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of power transaction, in particular to a machine learning prediction method and system for regulating load at any time scale. BACKGROUND
[0002] At present, there are various known regulating load prediction methods, and the scenarios are also more, which are more applied to the power grid side and the power generation side, and the method for regulating load prediction based on the power selling company scenario is less; the power selling company is the power user electricity agent subject, which is naturally sensitive to the price change of the power market, and the regulating load as demand has a great influence on the supply and demand relationship, and the fluctuation can directly cause the fluctuation of the electricity price; therefore, the power selling company is very sensitive to the change of the regulating load; at the same time, the power selling company itself participates in the power market transaction, and whether in the medium and long term market or the spot market, such as annual long-term cooperation, monthly medium and long term, even multi-day and day-ahead, the macro regulating load needs to be qualitatively and quantitatively perceived, and the corresponding decision is made according to the data owned, so as to avoid potential risks.
[0003] The common prediction method is basically based on time series power data, at most some basic weather indexes such as temperature and humidity are fused, and then traditional statistical regression or machine learning, artificial neural network and the like are directly used to predict data, and the end-to-end load prediction is carried out, the prediction result and the actual deviation are large, and the traditional prediction method cannot cover all time scales.
[0004] The present application is based on the related change characteristics analyzed by the big data in the real power selling company application scenario, a more rich multi-modal data feature library is built, a complete regulating load feature engineering is constructed, and finally the regulating load at any time scale is more accurately predicted by combining the load change physical logic and the prediction algorithm. SUMMARY
[0005] The present application provides a machine learning prediction method and system for regulating load at any time scale, which is used to solve the accuracy requirement of boundary condition analysis in medium and long term and spot transaction.
[0006] In one aspect, the present application provides a machine learning prediction method for regulating load at any time scale, comprising: obtaining historical load, meteorological and economic data, constructing a feature library, screening key features and preprocessing to obtain preprocessed data; based on the preprocessed data, constructing a dimension separation model and performing training optimization to obtain a decision tree model; based on the decision tree model, calculating the SHAP value of each feature, analyzing the contribution degree of the global feature, and obtaining an analysis result; Based on the analysis results, similar day matching and dimension fusion are performed to obtain the fused load forecast results; Based on the fused load forecast results, multi-time-scale layered forecasting is performed to generate load curves at any scale; Based on load curves of any scale, the prediction deviation is verified through actual data backtesting, data is monitored and incremental learning is triggered to form a closed-loop optimization.
[0007] Furthermore, historical load, meteorological and economic data are obtained to build a feature library, and key features are screened and preprocessed to obtain preprocessed data, including: Obtain historical load, meteorological, and economic data at 15-minute granularity as raw data; Based on the original data, a three-layer feature library is constructed, including basic features, composite features, and scene features; The Filter-Wrapper hybrid algorithm is used for the three-layer feature library. The correlation coefficient between the feature and the load is calculated for preliminary screening. Then, the feature subset is optimized by recursive feature elimination combined with the SVR model to obtain the key features for screening. The correlation coefficient between key features and load and the contribution of SHAP value are calculated to obtain the preprocessed data.
[0008] Furthermore, based on the preprocessed data, a dimension separation model is constructed and trained and optimized to obtain a decision tree model, including: Based on the preprocessed data, the load forecasting formula is constructed in combination with the dimensional separation concept, and the base load is decoupled from meteorological and economic fluctuations to obtain a dimensional separation model. The dimensional separation model is incrementally learned using a sliding window strategy to obtain a trained dimensional separation model. The trained dimension separation model is optimized through grid search and cross-validation to obtain the optimized trained dimension separation model; Based on the optimized training dimension separation model, the Leaf-wise tree growing strategy and GOSS sampling method are combined to obtain the decision tree model.
[0009] Furthermore, based on the decision tree model, the SHAP value of each feature is calculated and the global feature contribution is analyzed to obtain the analysis results, including: Based on the decision tree model, the SHAP value of each feature sample is calculated to obtain the contribution of each feature sample to the prediction result; Based on the contribution of each feature sample to the prediction result, the overall influence of the feature is analyzed to identify the key driving factors, i.e. the analysis results; Based on the key driving factors, the daily load rate and the peak-valley difference index of the prediction results are compared with the historical statistical interval to evaluate the rationality of the prediction, and the evaluation results are obtained. Based on the evaluation results, when the key driving factors deviate from the historical interval by more than a preset threshold, the abnormal prediction results are automatically marked, and the feature weight is adjusted based on artificial experience for abnormal results. If it continues to deviate, the model incremental training process is triggered.
[0010] Further, based on the analysis results, similar day matching and dimension fusion are performed to obtain the fused load prediction results, including: Based on the analysis results, the matching factor of the target period is calculated, and a matching factor sequence is formed, the matching factor includes N-dimensional vector calculation cosine similarity of meteorological and economic characteristics, typical day coefficient and time decay weight; Based on the matching factor sequence, gray correlation analysis is performed using the gray correlation algorithm, and the daily dimension is multiplied by the predicted amount of the day to obtain the daily load. Similarly, the monthly and annual prediction daily, monthly load can be obtained, and the fused load prediction result is obtained.
[0011] Further, based on the fused load prediction result, multi-time scale hierarchical prediction is performed to generate an arbitrary scale load curve, including: Based on the fused load prediction result, the annual total load is predicted, which is decomposed into 12 monthly load curves through month-by-month dimension matching to form an annual prediction framework; Based on the annual prediction framework, the monthly average load of the target month is predicted, and the daily load distribution is generated by combining the month-by-day dimension matching, and the day-by-hour dimension is superimposed to the hour level to obtain the monthly prediction result; Based on the monthly prediction result, a refined load curve with a granularity of 15 minutes is output to obtain the daily prediction result; By freely combining different time scales through dimension matching algorithm, a load prediction curve with user-specified precision is generated.
[0012] Further, based on the arbitrary scale load curve, the prediction deviation is verified through backtesting of actual data, the data is monitored and the incremental learning is triggered to form a closed-loop optimization, including: The arbitrary scale load prediction result is compared with the actual operation data through backtesting to calculate the key indicator deviation value to obtain the verification result; Based on the verification result, the key indicators are monitored in real time, and the incremental learning process is automatically triggered when the continuous overrun is completed to update the model parameters dynamically.
[0013] On the other hand, a machine learning prediction system for tuning load to any time scale includes: The acquisition module is used to obtain historical load, meteorological and economic data, build a feature library, filter key features and perform preprocessing to obtain preprocessed data; The processing module is used to construct a dimension separation model based on the preprocessed data and perform training optimization to obtain a decision tree model; based on the decision tree model, the SHAP value of each feature is calculated and the global feature contribution is analyzed to obtain the analysis results; based on the analysis results, similar day matching and dimension fusion are performed to obtain the fused load forecast results; based on the fused load forecast results, multi-time scale hierarchical forecasting is performed to generate arbitrary scale load curves; based on the arbitrary scale load curves, the forecast deviation is verified through actual data backtesting, the data is monitored and incremental learning is triggered to form a closed-loop optimization.
[0014] On the other hand, the present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, it implements a machine learning prediction method for coordinated loads at any time scale as described in any of the above-mentioned methods.
[0015] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a machine learning prediction method for coordinated loads at any time scale as described in any of the above.
[0016] On the other hand, the present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above-mentioned machine learning prediction methods for coordinated loads at any time scale.
[0017] The machine learning prediction method and system for centralized load at any time scale provided by this invention can identify many relevant factors and sensitivities affecting the load while predicting it. This facilitates identifying important directional trends in the centralized load when important boundary conditions change, such as sudden climate change or the La Niña / El Niño transition, and can predict maximum and minimum values based on extreme values. Accurate boundaries / confidence intervals provide trading risk boundaries, making trading behavior and strategies more controllable. Furthermore, the core prediction principle of "dimensional separation" is relatively simple, and the machine learning algorithm used is highly efficient, consuming little memory, and suitable for rapid deployment and iteration. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1 1 is a flow chart of a method for machine learning prediction of coordinated loads at any time scale provided by an embodiment of the present invention; Figure 2 Schematic diagram of a machine learning prediction system for centralized load regulation at any time scale provided by an embodiment of the present invention; Figure 3 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0020] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0021] Figure 1 This is one of the flow charts of the machine learning prediction method for the coordinated load at any time scale provided by an embodiment of the present invention.
[0022] like Figure 1 As shown, the embodiment of the present invention provides a machine learning prediction method for centralized load at any time scale, and the method mainly includes the following steps: 11. Obtain historical load, meteorological and economic data, build a feature library, screen key features and perform preprocessing to obtain preprocessed data; 12. Based on the preprocessed data, a dimension separation model is constructed and trained and optimized to obtain a decision tree model; 13. Based on the decision tree model, calculate the SHAP value of each feature and analyze the global feature contribution to obtain the analysis results; 14. Based on the analysis results, similar day matching and dimension fusion are performed to obtain the fused load forecast results; 15. Based on the fused load forecast results, multi-time scale layered forecasting is performed to generate load curves of any scale; 16. Based on load curves of any scale, the prediction deviation is verified through actual data backtesting, the data is monitored and incremental learning is triggered to form a closed-loop optimization.
[0023] In the embodiment of the present invention, the complete physical logic behind the unified dispatching load is first sorted out. It represents the "load" situation of the dispatching caliber of Guangdong Province, reflecting the sum of the actual electricity demand of all the smallest electricity-consuming units in the Guangdong power system, including the total power of various types of electricity-consuming equipment such as industry, commerce, and residents; due to the real-time nature of unified dispatching and the accuracy of publicly disclosed information to every 15 minutes, the prediction usually only involves a key time-series load indicator of the day or month, or even only the year, such as minimum load, maximum load, and average load. Since it reflects "demand", no matter at which time scale, its load is a comprehensive reflection of the electricity consumption behavior at the current time scale, and the electricity consumption behavior is affected by many factors. Therefore, to meet forecasting needs at any time scale, we adopt the core principle of "dimension separation, where quantities can be broken down into bases and superimposed fluctuations, with the dimensions determined by similar behaviors." This involves attempting to separate the underlying load from the centralized load. We then examine the fluctuations in electricity load caused by single or multiple factors, such as meteorological and economic factors. Each influencing factor (hereafter referred to as a characteristic) can be accurately quantified by studying its correlation and sensitivity with the centralized load. Forecasting requires only meteorological data, some expected economic data, or other comprehensive indicators to calculate the future impact of these different characteristics on the centralized load. The importance of these characteristics can also be explained using the SHAP value theory from game theory, providing comprehensive interpretability and facilitating subsequent model optimization and review. This provides substantial support for decision-making regarding centralized load and subsequent trading.
[0024] like Figure 1 As shown in 11, historical load, meteorological and economic data are obtained, a feature library is constructed, and key features are screened and preprocessed to obtain preprocessed data, including: 111. Obtain historical load, meteorological and economic data with 15-minute granularity as raw data; 112. Based on the original data, a three-layer feature library is constructed, which includes basic features, composite features and scene features; 113. The Filter-Wrapper hybrid algorithm is used for the three-layer feature library. The correlation coefficient between the feature and the load is calculated for preliminary screening. Then, the feature subset is optimized by recursive feature elimination combined with the SVR model to obtain the key features for screening. 114. Calculate the correlation coefficient between key features and load and the contribution of SHAP value to obtain preprocessed data.
[0025] In an embodiment of the present invention, historical data on power system load is collected, including: load values at different time granularities (such as hours, days, etc.), data on factors affecting the load, meteorological data including temperature, humidity, wind speed, wind force level, wind direction, sunlight, cloud cover, atmospheric pressure, rainfall, etc. in various cities, date types including weekdays, weekends, holidays, etc., economic activity indicators including industrial output value, business activity index, etc., population data, load ratios in various regions, etc.; Delete samples with a large number of missing values, and use statistical analysis or domain knowledge to determine and correct or eliminate abnormal data points. Filter outliers based on the 3σ principle, where σ represents the standard deviation and μ represents the mean. If the absolute distance between the data and the mean is greater than three times the standard deviation, that is, the range between [-∞, μ-3σ] and [μ+3σ, +∞], this part of the data is an outlier. For outliers, the average power consumption of the same type of date is used to fill the gap. Extract useful features based on domain knowledge and data characteristics. These features include basic and composite features of time series, meteorology, and economics. Basic features are derived from data available during the raw data preparation phase, while composite features are derived from statistics and empirical summaries. For time series data, time features (hours, days of the week, months, etc.) and lag features (load values at several past time points) are extracted, and appropriate time window statistics and other operations are performed on the meteorological data, including but not limited to separate labeling of structured time series data, and data cleaning, stratification, and cross-cutting.
[0026] The specific calculation method of characteristic indicators is as follows: The data of all historical days are classified into typical days and typical time periods. Typical days are divided into typical Mondays, typical weekdays, typical weekends, and different holidays, Spring Festival, New Year's Day, and Qingming Festival. At the same time, these typical days are divided into different load levels (the load weight of typical days is calculated, and the load weight of weekends / Saturdays compared to weekdays will be marked accordingly after statistics). The input of categorical feature variables is used. The typical time period is divided into peak, flat and valley periods according to Guangdong's peak, flat and valley periods. At the same time, data cross-reconstruction features are performed between days. The daily load of the past week is multiplied by the daily peak-to-valley difference rate of the past week to reconstruct a new time series indicator feature. Among them, time series characteristics are divided into three categories, daily load characteristic indicators, monthly load characteristic indicators and annual load characteristic indicators; daily load characteristics include but are not limited to: daily average load, daily maximum load, daily minimum load, daily load rate, etc.; monthly load characteristics include but are not limited to: monthly maximum load, monthly average daily load, monthly maximum peak-to-valley difference, monthly average daily load rate, cumulative maximum load utilization hours; annual load characteristics include but are not limited to: annual maximum load, annual average load, seasonal imbalance coefficient, annual maximum load utilization hours; meteorological characteristics include basic temperature, humidity, wind speed, wind force level, wind direction, light, cloud cover, atmospheric pressure, rainfall; according to these basic meteorological characteristics, the time series characteristics are divided into three categories, daily load characteristic indicators, monthly load characteristic indicators and monthly load characteristic indicators; monthly load characteristics include but are not limited to: monthly maximum load, monthly average daily load, monthly maximum peak-to-valley difference, monthly average daily load rate, cumulative maximum load utilization hours; annual load characteristics include but are not limited to: annual maximum load, annual average load, seasonal imbalance coefficient, annual maximum load utilization hours; meteorological characteristics include basic temperature, humidity, wind speed, wind force level, wind direction, light, cloud cover, atmospheric pressure, rainfall, etc. In addition to the image features, the following were manually added: perceived temperature, urban heat effect coefficient, 2-day accumulated temperature effect, 3-day accumulated temperature effect, and 4-day accumulated temperature effect. Economic features include GDP, CPI, energy consumption intensity, primary, secondary, and tertiary industry income, export investment, and national economic income. Some complex composite features have undergone special processing within the system, and the time series will distinguish between typical working days, weekends, and statutory and adjusted holidays. In terms of data structure, some indicators will be directly bucketed and crossed, such as multiplying temperature and humidity, and then converted into categorical feature tracking. At the same time, temperatures will also be aggregated by region, such as the regional impact of light and the regional impact of temperature. After constructing the above features, the feature library is complete: the feature library will count the changing relationships of each feature, such as the daily, monthly, and annual trends of time-sharing temperature and unified load, and scatter plots. This can quickly determine their correlation and sensitivity, and use the piecewise linear regression method to automatically divide the linear change interval into 1, 2, 3, or even more segments, quantifying the impact of the temperature in each interval on the composite change. This is done for simple features and also for composite features. The correlation is based on the Filter-Wrapper hybrid feature selection algorithm to select the optimal feature. The filter part: calculates the correlation coefficient between the feature and the load: ; in, Representative features values, X represents the feature mean, Represents the amount of electricity Values, Y represents the average load of the unified regulation, and irrelevant features are eliminated according to the size of the correlation coefficient; Wrapper part: First, a data set with n features is obtained after filtering through Filter. Second, SVR is used as the base model of recursive feature elimination (RFE). Third, given the number of features k (k <n),最后,第一轮对所有特征进行训练,根据MSE指标给出每个特征的排名,剔除得分最小的特征,此时特征减少为n-1,继续以n-1个特征进行迭代,直到特征保留为k个,即最终进入模型训练的特征; Finally, a multidimensional table of characteristics and loads can be statistically calculated, as shown below: ; Uncertain features are primarily tracked in certain scenarios, primarily processing textual information. Therefore, a standard data storage format has been designed to convert unstructured data into structured data, thereby enabling the use of additional large-scale model reasoning capabilities and multimodal input. The specific feature storage table is as follows: ; By tracking these features, we can know which features can be input into the prediction model as features, thereby pre-screening the model input features and avoiding repeated algorithm iterations that waste too much computing resources and information loss.
[0027] like Figure 1 As shown in 12, based on the preprocessed data, a dimension separation model is constructed and trained and optimized to obtain a decision tree model, including: 121. Based on the pre-processed data, the load forecasting formula is constructed in combination with the dimensional separation idea, and the base load is decoupled from meteorological and economic fluctuations to obtain a dimensional separation model. 122. The dimension separation model is incrementally learned using a sliding window strategy to obtain a trained dimension separation model. 123. Optimize key parameters of the trained dimension separation model through grid search and cross validation to obtain the optimized trained dimension separation model; 124. Based on the optimized training dimension separation model, the Leaf-wise tree growing strategy and GOSS sampling method are combined to obtain the decision tree model.
[0028] In an embodiment of the present invention, in an actual application scenario, new data on power system load and related factors are continuously collected. When a certain amount of new data is accumulated, incremental learning is prepared, and the collected data is divided into an initial training set and a test set. The initial training set is used to build a basic version of the model, and the test set is used to preliminarily evaluate the model performance. During the training process, the idea of "enhanced learning / lifelong learning" is used to continuously slide the time window by setting the step size. Each time the training set is imported for training, the generalization of the test set is tested to improve the prediction accuracy. If the sliding step size is set to 90 days and the validation set size is set to 30 days, then the first step will be to fit the model. The decision tree is iterated on the data from January 1, 2020, to be verified on the data from July 1, 2020, to be used for data iteration. In the second step, the data from April 1, 2020, to be fitted, is verified on the data from October 1, 2020, to be used for data iteration again, until all the data sets are exhausted. This step-by-step, iterative process can support data sets of any size. The only issue is the length of training time. However, there is no need to worry about the future expansion of data. Parallel computing is also very efficient and accuracy is guaranteed. Initially set the tree depth (max_depth), learning rate (learning_rate), number of trees (num_leaves), objective function (objective), number of cycles (num_iterations) of the LightGBM model. These parameters are initially set based on previous experience and preliminary understanding of the data, and will be adjusted later. GridSearch combined with cross-validation (such as k-fold cross-validation) is used to tune the parameters of the LightGBM model on the initial training set. During the parameter tuning process, the evaluation indicator of focus is the mean absolute percentage error (MAPE). This indicator is sensitive to relative error and will not change due to the global scaling of the target variable. It is suitable for problems with large dimensional differences in the target variables. The parameter combination that minimizes this indicator is selected. The calculation formula of MAPE is: ; in, For the The true load value of the sample (the actual observed value), For the The predicted load value (model output value) of samples, nsamples is the total number of samples (the total number of data points involved in the evaluation), is a very small constant (such as ), used to avoid the denominator being zero and ensure numerical stability; Use the initial training set data to train the LightGBM model according to the determined parameter settings; during the training process, observe the changes in the model's training loss and validation loss to determine whether the model is overfitting or underfitting, and further adjust the parameters or data preprocessing methods as needed; use LightGBM's native commands for incremental learning, specifying the previously trained model file (such as the 'init_model' parameter) and the path to the new data (such as the 'data' parameter) in the command; at the same time, integrate the preprocessed new data with the previous training data, or directly use the new data as incremental data; parameter adjustment and training: use the grid search method again to fine-tune the model parameters to adapt to the characteristics and changes of the new data.
[0029] The specific method of grid search is as follows: Determine the hyperparameters that need to be optimized and their candidate value ranges. These parameters and their value ranges are set based on the nature of the coordinated load data type, the size of the dataset, and existing experimental experience. After defining the LightGBM model, initialize the grid search using the public GridSearchCV method, specifying the model, parameter grid, cross-validation folds, and scoring criteria; Call the fit method to fit the training set. The model will traverse all parameter combinations and perform cross-validation on each set of parameters. After the grid search is completed, the best parameter combination is obtained through the best_params_ attribute, and then the model is reinitialized with the optimal parameters and trained on the entire training set. During the training process, parameters such as the learning rate can be adjusted appropriately so that the model can better learn information from new data while retaining the original knowledge. Through iterative training, the model gradually updates and optimizes its prediction capabilities to adapt to the changing patterns of power system load. The updated model's performance is then continuously re-evaluated on the test set. Error metrics (MAPE and other indicators) between the predicted and actual load values are calculated and compared with the previous model's performance to verify whether incremental learning has improved the model's prediction accuracy. This allows the model to be checked for generalization on new data, ensuring that the model has not overfitted to the new data and lost its grasp of the overall data patterns. Finally, a final model is output.
[0030] like Figure 1 As shown in 13, based on the decision tree model, the SHAP value of each feature is calculated and the global feature contribution is analyzed to obtain the analysis results, including: 131. Based on the decision tree model, calculate the SHAP value of each feature sample to obtain the contribution of each feature sample to the prediction result; 132. Based on the contribution of each feature sample to the prediction result, analyze the overall influence of the feature and identify the key driving factors, i.e. the analysis results; 133. Based on the key driving factors, the daily load rate and peak-to-valley difference indicators of the forecast results are compared with the historical statistical intervals to evaluate the rationality of the forecast and obtain the evaluation results; 134. Based on the evaluation results, when the key driving factors deviate from the historical range by more than the preset threshold, the abnormal prediction results are automatically marked, and the feature weights of the abnormal results are adjusted based on manual experience. If the deviation continues, the model incremental training process is triggered.
[0031] In an embodiment of the present invention, the SHAP library is installed in a Python environment, and the SHAP interpreter is imported: TreeExplainer (an interpreter specially designed for tree-based models such as LightGBM) is imported into the code; then the SHAP value is calculated: the trained LightGBM model and interpreter are used to calculate the SHAP value for the samples in the training set or test set. The SHAP value represents the contribution of each feature to the model prediction result, that is, the impact of the value of the feature on the model output considering all other features; the calculation formula of the SHAP value is based on the concept of Shapley value, which is a concept in game theory used to distribute benefits in cooperative games; in the context of machine learning model interpretation, the SHAP value is used to distribute the contribution of the model prediction to each feature; for a model with n features, the SHAP value of the i-th feature is It can be calculated by the following formula: ; Where N is the set of all features, S is a subset of N excluding feature i, |S| is the number of features in set S, f(S) is the model's prediction when considering only the features in set S, and f(S∪{i}) is the model's prediction when considering set S and feature i. This calculates the average marginal contribution of feature i across all possible feature combinations. However, this formula is computationally expensive in practice, so the SHAP library uses a more efficient algorithm to approximate the SHAP value. This is how the SHAP value of the ensemble decision tree model used in this patent is calculated. By drawing a feature importance plot (such as SHAPsummaryplot), we can show the overall impact of all features on the model prediction results. The horizontal axis of the plot represents the absolute value of the SHAP value, the vertical axis represents the feature name, and the color represents the size of the feature value. This can help us understand which features are most critical to power system load forecasting and identify the main influencing factors. For a specific prediction sample, a forceplot is drawn to visually display the direction (positive or negative) and magnitude of each feature's contribution to the sample's prediction result. For example, for a load forecast at a specific time point, you can see whether a temperature increase contributes to an increase or decrease in the load forecast, as well as the extent of the contribution. This helps us understand the model's decision-making basis in this specific situation. It is also a means of verifying the reliability of the model. The model's interpretation of a particular feature should be consistent with past statistical characteristics and the direction of load changes. Draw a SHAPdependenceplot to show the relationship between a specific feature and the model output. At the same time, you can also observe the interaction between this feature and other features, which helps to further understand the impact of complex relationships between features on load forecasting. For example, analyze the nonlinear relationship between temperature and load, and whether humidity plays a regulatory role in it. At this step, import the deviation correction model at the same time. Through statistical calculation indicators such as daily load rate and daily peak-to-valley difference, test the degree of fit between the output results in the statistical indicators and the forecast and history. When all kinds of boundary conditions in history and forecast are the same, the forecast results and indicators should be consistent with our historical data and indicators.
[0032] like Figure 1 As shown in 14, based on the analysis results, similar day matching and dimension fusion are performed to obtain the fused load forecast results, including: 141. Based on the analysis results, calculate the matching factor for the target period and form a matching factor sequence. The matching factor includes the cosine similarity calculated from the N-dimensional vector of meteorological and economic characteristics, the typical day coefficient, and the time decay weight. 142. Based on the matching factor sequence, the grey correlation algorithm is used to perform grey correlation analysis. By multiplying the daily outline with the daily forecast, the time-sharing load of the day can be obtained. Similarly, the forecast daily and monthly loads of the month and year can be obtained to obtain the fused load forecast result.
[0033] In the embodiment of the present invention, during the invention process, the prediction results of other models (such as linear regression, random forest, neural network, etc.) were also compared, and it was found that based on the existing structured data, the LightGBM model has stronger prediction accuracy and better performance than other models in power system load forecasting; Calculate the daily, monthly, and annual distribution of the unified dispatching load. Simultaneously, calculate historical monthly maximum temperature, average temperature, average rainfall, and other indicators (X_1, X_2, …X_n) to serve as the basis for matching the distribution with similar-day data. In the distribution similarity algorithm, a gray correlation analysis is performed on the sequence composed of three factors: meteorological and economic characteristic similarity, typical days, and time series weights. This analysis identifies the historical time series most similar to the predicted time series. Gray correlation analysis is primarily used to study systems with partially clear and partially ambiguous information. The correlation between the parent sequence (reference sequence) and the child sequence (comparison sequence) is calculated. The parent sequence can be considered the main characteristic sequence of system behavior. When studying unified dispatching load, the three factor sequence of the unified dispatching forecast day can be used as the parent sequence. The child sequence is the sequence of other relevant factors compared with the parent sequence, i.e., the sequence composed of the three historical factors. Assume that the parent sequence is X0={x0(1),x0(2),…,x0(n)}, and the subsequence is Xi={xi(1),xi(2),…,xi(n)}(i=1,2,…,m), where n is the length of the sequence and m is the number of subsequences; For each data in the sequence, each data in the sequence is divided by the first data in the sequence; the processed parent sequence is X0′={x0′(1)=1,x0′(2)=x0(2) / x0(1),…,x0′(n)=x0(n) / x0(1)}. Similarly, the subsequences are also processed in the same way; the purpose of this is to eliminate the influence of different dimensions on the correlation calculation; The correlation coefficient is then calculated: ; Where k = 1, 2, ..., n, i = 1, 2, ..., m, ρ is the resolution coefficient, usually 0.5; is the minimum value of all differences between the parent sequence and the child sequence, It is the maximum value of all differences between the parent sequence and the child sequence; the correlation coefficient reflects the degree of correlation between the parent sequence and the child sequence at each corresponding point; The final grey correlation degree, that is, the average value of the correlation coefficient is: ; Among them, r i It represents the overall correlation between the parent sequence and the child sequence, n is the grey correlation degree, i.e. the number of correlation coefficients, is the grey correlation coefficient; the value range of the correlation is generally between [0,1]. The larger the correlation, the higher the degree of correlation between the two sequences. The cosine similarity is calculated for the factor matching of meteorological characteristics and economic characteristics. The calculation formula of cosine similarity is as follows: ; Among them, meteorological characteristics such as temperature, humidity, perceived temperature, and atmospheric pressure are an N-dimensional vector A, and the cosine angle between it and the N-dimensional vector B of meteorological characteristics at the forecast time can be calculated as cos(A, B). The same is true for economic characteristics. For the matching of typical day factors, there will be a similarity coefficient dictionary between typical days. The default is 0.15 for Monday, 0.2 for Tuesday to Thursday, 0.25 for Friday, 0.7 for Saturday, and 1.0 for Sunday. The formula for calculating the typical day matching coefficient is:
[0034] in, and is the corresponding dictionary matching value, is the typical day matching coefficient; For the matching of time factors, linear attenuation is performed according to the distance between the historical day and the forecast day. For example, the weight of the past day is 0.99, the weight of the past seven days is 0.9, and the weight of the past 365 days is 0.8. The principle of "larger for the near and smaller for the far" is adopted, so the closer the historical data to the forecast day, the greater the weight. At the same time, the periodicity of load changes and the impact of holidays are taken into consideration. The specific calculation formula is as follows:
[0035] Where δ is equal to T-mold Power, t divided by Power, t divided by The result of multiplying these three parts is: beta1 is 0.95, beta2 is 0.96, beta3 is 0.97, M1 is 7, M2 is 365, and M3 is 340; The three major coefficients form a sequence, and then the grey correlation degree is used to calculate the most similar date. By using the day's outline x the day's predicted quantity, the day's time-sharing load can be obtained; by analogy, the predicted daily and monthly loads for the month and year can also be obtained.
[0036] like Figure 1 As shown in 15, based on the fused load forecast results, multi-time scale hierarchical forecasting is performed to generate load curves of any scale, including: 151. Based on the integrated load forecast results, the annual total load is predicted and decomposed into 12 monthly load curves by matching the annual and monthly load curves to form an annual forecast framework; 152. Based on the annual forecast framework, the average monthly load of the target month is forecasted. The daily load distribution is generated by combining the monthly and daily distributions. The daily and hourly distributions are then refined to the hourly level to obtain the monthly forecast results. 153. Based on the monthly forecast results, output the refined load curve with 15-minute granularity to obtain the daily forecast results; 154. Different time scales are freely combined through the dimension matching algorithm to generate a load forecast curve with user-specified accuracy.
[0037] In the embodiment of the present invention, by matching the annual total load forecast with the annual monthly plan, a full-year load distribution benchmark is established, providing a macro basis for long-term power planning (such as power generation scheduling and energy procurement), forming a stable monthly load distribution ratio, and avoiding the cumulative error amplification problem caused by directly forecasting monthly data; Dynamically adjust the monthly average load within the annual framework, combining historical daily load patterns (monthly and daily overview) and time-of-day fluctuation patterns (daily and time-of-day overview) to achieve consistent load forecasts at the monthly, daily, and hourly levels. This improves the timeliness of monthly forecasts, especially adapting to the impact of short-term events such as holidays or extreme weather, and supports monthly power trading and unit portfolio optimization. Directly generate 15-minute load curves based on monthly forecast results to meet real-time scheduling needs. Through high-precision time-of-day dimension matching, significantly reduce peak and valley period forecast deviations (MAPE < 2%), facilitating peak shaving and valley filling and demand response. Allows users to select the forecast scale (e.g., quarterly, weekly, hourly) as needed, dynamically synthesizes the target curve through a dimension-quantity matching algorithm, breaking the fixed-scale limitations of traditional models and achieving "one-time modeling, multi-scale output," significantly reducing the coordination cost of cross-scale forecasting. Through top-down hierarchical forecasting and dimension-quantity fusion of "year→month→day", the stability of long-term trends is guaranteed while retaining the flexibility of short-term fluctuations. It is suitable for multiple scenarios such as power market transactions and power grid security verification.
[0038] like Figure 1 As shown in 16, based on load curves of any scale, the prediction deviation is verified by backtesting actual data, monitoring data and triggering incremental learning to form a closed-loop optimization, including: 161. Backtest and compare the load forecast results of any scale with the actual operation data, calculate the deviation value of key indicators to obtain verification results; 162. Based on the verification results, key indicators are monitored in real time. When they exceed the limit continuously, the incremental learning process is automatically triggered to complete the dynamic update of model parameters.
[0039] In the embodiment of the present invention, the forecast results at different time scales (year / month / day) are compared with the actual operating data, key indicators such as MAPE and daily load factor deviation are calculated, the model forecast accuracy is quantified, and the model performance is comprehensively evaluated through multi-dimensional indicators (such as peak-to-valley deviation and trend consistency). Systematic deviations (such as long-term overestimation or insufficient capture of short-term fluctuations) are identified to provide data support for optimization. Real-time monitoring of key indicators (e.g., MAPE > 5% for three consecutive times) automatically triggers an incremental learning mechanism, using the latest data to update model parameters, avoiding manual intervention delays. The model continuously adapts to changes in load patterns (e.g., demand curve deformation caused by the integration of new energy sources), maintaining forecast accuracy, and can self-correct the model within 48 hours for sudden anomalies (e.g., extreme weather events), reducing forecast risks. Retraining is only initiated during periods of excessive deviation, reducing redundant computing costs by over 80%. Through the automated closed loop of "prediction-verification-optimization", the model has the ability to continuously evolve, which is particularly suitable for scenarios where load characteristics change rapidly during the transformation period of the power system, and can reduce the overall prediction error in the long term.
[0040] like Figure 2 As shown, a machine learning prediction system 20 for load regulation at any time scale includes: The acquisition module 21 is used to obtain historical load, meteorological and economic data, build a feature library, and screen key features and perform preprocessing to obtain preprocessed data; The processing module 22 is used to construct a dimension separation model based on the preprocessed data and perform training optimization to obtain a decision tree model; based on the decision tree model, calculate the SHAP value of each feature and analyze the global feature contribution to obtain the analysis result; based on the analysis result, perform similar day matching and dimension fusion to obtain the fused load forecast result; based on the fused load forecast result, perform multi-time scale hierarchical forecasting to generate an arbitrary scale load curve; based on the arbitrary scale load curve, verify the forecast deviation through actual data backtesting, monitor the data and trigger incremental learning to form a closed-loop optimization.
[0041] Figure 3 It is a structural diagram of an electronic device provided by an embodiment of the present invention.
[0042] like Figure 3 As shown, the electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communications interface 620, and the memory 630 communicate with each other via the communication bus 640. The processor 610 may call the logic instructions in the memory 630 to execute a machine learning prediction method for the coordinated load at any time scale.
[0043] In addition, the logic instructions in the aforementioned memory 630 can be implemented in the form of a software functional unit and, when sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0044] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the machine learning prediction method for the coordinated load provided by the above methods at any time scale.
[0045] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a machine learning prediction method for performing coordinated load prediction on any time scale provided by the above methods.
[0046] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0047] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0048] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A machine learning prediction method for centralized load at any time scale, characterized in that: include: Obtain historical load, meteorological and economic data, build a feature library, filter key features and perform preprocessing to obtain preprocessed data; Based on the preprocessed data, a dimension separation model is constructed and trained and optimized to obtain a decision tree model; Based on the decision tree model, the SHAP value of each feature is calculated and the global feature contribution is analyzed to obtain the analysis results; Based on the analysis results, similar day matching and dimension fusion are performed to obtain the fused load forecast results; Based on the fused load forecast results, multi-time-scale layered forecasting is performed to generate load curves at any scale; Based on load curves of any scale, the prediction deviation is verified through actual data backtesting, data is monitored and incremental learning is triggered to form a closed-loop optimization.
2. The method for machine learning prediction of coordinated load at any time scale according to claim 1 is characterized in that: Obtain historical load, meteorological, and economic data, build a feature library, filter key features, and perform preprocessing to obtain preprocessed data, including: Obtain historical load, meteorological, and economic data at 15-minute granularity as raw data; Based on the original data, a three-layer feature library is constructed, including basic features, composite features, and scene features; The Filter-Wrapper hybrid algorithm is used for the three-layer feature library. The correlation coefficient between the feature and the load is calculated for preliminary screening. Then, the feature subset is optimized by recursive feature elimination combined with the SVR model to obtain the key features for screening. The correlation coefficient between key features and load and the contribution of SHAP value are calculated to obtain the preprocessed data.
3. The method for machine learning prediction of coordinated load at any time scale according to claim 2 is characterized in that: Based on the preprocessed data, a dimension separation model is constructed and trained and optimized to obtain a decision tree model, including: Based on the preprocessed data, the load forecasting formula is constructed in combination with the dimensional separation concept, and the base load is decoupled from meteorological and economic fluctuations to obtain a dimensional separation model. The dimensional separation model is incrementally learned using a sliding window strategy to obtain a trained dimensional separation model. The trained dimension separation model is optimized through grid search and cross-validation to obtain the optimized trained dimension separation model; Based on the optimized training dimension separation model, the Leaf-wise tree growing strategy and GOSS sampling method are combined to obtain the decision tree model.
4. The method for machine learning prediction of coordinated load at any time scale according to claim 3 is characterized in that: Based on the decision tree model, the SHAP value of each feature is calculated and the global feature contribution is analyzed to obtain the analysis results, including: Based on the decision tree model, the SHAP value of each feature sample is calculated to obtain the contribution of each feature sample to the prediction result; Based on the contribution of each feature sample to the prediction result, the overall influence of the feature is analyzed to identify the key driving factors, i.e. the analysis results; Based on key driving factors, the daily load rate and peak-to-valley difference indicators of the forecast results are compared with the historical statistical intervals to evaluate the rationality of the forecast and obtain the evaluation results; Based on the evaluation results, when the key driving factors deviate from the historical range by more than the preset threshold, the abnormal prediction results are automatically marked, and the feature weights of the abnormal results are adjusted based on manual experience. If the deviation continues, the model incremental training process is triggered.
5. The method for machine learning prediction of coordinated load at any time scale according to claim 4 is characterized in that: Based on the analysis results, similar day matching and dimension fusion are performed to obtain the fused load forecast results, including: Based on the analysis results, the matching factors of the target period are calculated and a matching factor sequence is formed. The matching factors include the cosine similarity, typical day coefficient and time decay weight calculated from the N-dimensional vector of meteorological and economic characteristics; Based on the matching factor sequence, the grey correlation algorithm is used to perform grey correlation analysis. By multiplying the daily forecast by the forecasted quantity, the time-of-day load can be obtained. Similarly, the forecasted daily and monthly loads for the current month and the current year can be obtained, and the fused load forecast result can be obtained.
6. The method for machine learning prediction of coordinated load at any time scale according to claim 5 is characterized in that: Based on the fused load forecast results, multi-time-scale layered forecasting is performed to generate load curves at any scale, including: Based on the integrated load forecast results, the annual total load is predicted and decomposed into 12 monthly load curves through annual and monthly matching to form an annual forecast framework; Based on the annual forecast framework, the average monthly load is forecasted for the target month. The daily load distribution is generated by combining the monthly and daily distributions. The daily and hourly distributions are then refined to the hourly level to obtain the monthly forecast results. Based on the monthly forecast results, a refined load curve with 15-minute granularity is output to obtain the daily forecast results; Through the dimension matching algorithm, different time scales are freely combined to generate a load forecast curve with user-specified accuracy.
7. The method for predicting the coordinated load with machine learning at any time scale according to claim 6 is characterized in that: Based on load curves of any scale, we verify prediction deviations through real data backtesting, monitor data, and trigger incremental learning to form a closed-loop optimization, including: Backtest and compare the load forecast results of any scale with the actual operation data, calculate the deviation value of key indicators to obtain verification results; Based on the verification results, key indicators are monitored in real time. When they exceed the limit continuously, the incremental learning process is automatically triggered to complete the dynamic update of model parameters.
8. A machine learning prediction system for load regulation at any time scale, characterized by: include: The acquisition module is used to obtain historical load, meteorological and economic data, build a feature library, filter key features and perform preprocessing to obtain preprocessed data; The processing module is used to construct a dimension separation model based on the preprocessed data and perform training optimization to obtain a decision tree model; Based on the decision tree model, the SHAP value of each feature is calculated and the global feature contribution is analyzed to obtain the analysis results; Based on the analysis results, similar day matching and dimension fusion are performed to obtain the fused load forecast results; Based on the fused load forecast results, multi-time-scale hierarchical forecasting is performed to generate load curves of any scale. Based on the load curves of any scale, the forecast deviation is verified through actual data backtesting, the data is monitored and incremental learning is triggered to form a closed-loop optimization.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the machine learning prediction method for the coordinated load at any time scale as described in any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for machine learning prediction of coordinated loads at any time scale as claimed in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Short-term load prediction method based on similar day optimization screening
CN110163429A
Short-term load prediction method based on TCN and IPSO-LSSVM combined model
CN111860979A
Time-phased load prediction method based on machine learning
CN117767281A
Power grid evaluation method and system based on power load decomposition
CN118607767A
Power grid load prediction system and method based on machine learning
CN119398274A
Cited By
Provincial dispatching non-marketization unit output prediction method and system
CN122338736A