Load prediction method and system based on ensemble learning and hyper-parameter intelligent optimization
By employing ensemble learning and intelligent hyperparameter optimization, a multi-source heterogeneous feature set is constructed and the model hyperparameters are optimized. This solves the problems of parameter sensitivity and insufficient feature fusion in existing power load forecasting methods, and achieves high-precision, real-time load forecasting.
Patent Information
- Application Number
- CN202610133867.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-15
AI Technical Summary
Existing power load forecasting methods are inadequate in terms of the systematic nature of parameter optimization, model effectiveness, and multi-source feature fusion, making it difficult to meet the requirements of smart grids for high precision, high stability, and strong adaptability.
A method based on ensemble learning and intelligent hyperparameter optimization is adopted. A multi-source heterogeneous feature set is constructed through multi-source input data. The hyperparameters of the ensemble learning base model are optimized collaboratively using intelligent optimization algorithms. Combined with dynamic weighted ensemble and lightweight residual correction, load forecasting is achieved.
It improves the stability and accuracy of forecasts, reduces computational costs, meets the real-time requirements of short-term load forecasting, enhances the forecasting adaptability to complex scenarios such as extreme weather and holidays, and achieves efficient characterization of load patterns.
Smart Images

Figure CN122051935A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power load forecasting, specifically relating to a load forecasting method and system based on ensemble learning and intelligent hyperparameter optimization. Background Technology
[0002] In the field of power load forecasting, existing technologies are mainly divided into two categories: traditional machine learning methods and deep learning methods. However, these methods have significant limitations in practical applications.
[0003] In the field of electricity load forecasting, the first generation of traditional machine learning methods primarily relies on time series analysis, regression methods, wavelet transform, least squares estimation, autoregressive moving average (ARIMA) models, support vector machines (SVM), and classification regression trees (CART). These methods require establishing statistical relationships between energy consumption and input variables using historical data, thus heavily depending on large-scale historical datasets. The second generation of forecasting methods, however, incorporates artificial intelligence techniques such as artificial neural networks (ANN), random forests, gradient boosting, fuzzy logic, genetic algorithms, and particle swarm optimization (PSO).
[0004] However, these traditional machine learning methods generally suffer from a serious deficiency in parameter configuration sensitivity. For example, in ensemble methods of extreme learning machines, the choice of network structure parameters directly affects prediction accuracy; in the application of radial basis function (RBF) neural networks, the parameter settings of the second-order training algorithm have a decisive impact on short-term load forecasting performance. Studies have shown that the performance of artificial neural network models is highly dependent on the appropriate selection of training data, the design of the learning algorithm, and the optimized configuration of the network structure; different parameter combinations can lead to orders of magnitude differences in prediction errors. Comparative studies have found that even with the same neural network architecture, significant differences in prediction performance exist when trained using the Levenberg-Marquardt (LM) method, the gradient descent (GD) method, and the gradient descent momentum (GDM) method, with the LM method showing the best performance.
[0005] Furthermore, while techniques such as the combination of support vector machines and fuzzy control, and the PSO-SVM hybrid method have achieved good results in specific scenarios, the tuning process of hyperparameters such as kernel function selection, penalty parameter setting, and convergence criteria of optimization algorithms is extremely complex and lacks a systematic optimization framework, resulting in a serious lack of generalization ability of the model in different application scenarios.
[0006] In recent years, deep learning technology has become a research hotspot due to its success in computer vision and natural language processing. In load forecasting, recurrent neural networks such as Long Short-Term Memory (LSTM) have been widely used to process residential electricity consumption data. Research shows that LSTM-based sequence-to-sequence (S2S) architectures can handle minute-level and hour-level resolution data for individual residential customers and can be used for short-term individual customer electricity consumption forecasting. Deep residual networks (ResNet) have also proven effective in short-term load forecasting. Deep neural networks, deep convolutional networks, and other architectures have also been proposed. These represent the second type of deep learning method currently existing in the field of electricity load forecasting.
[0007] However, these deep learning methods face the fundamental problem of insufficient effectiveness in practical applications: High data demand and high training cost: Deep learning models typically require massive amounts of training data to fully learn the complex nonlinear relationships of load patterns. However, in actual power distribution network applications, high-quality labeled data is often difficult to obtain, making model training difficult.
[0008] Risk of overfitting and weak generalization ability: Deep network structures are prone to overfitting when training data is limited, leading to a significant drop in the model's predictive performance on new data. Although LSTM can capture long-term dependencies in time series, its generalization ability remains limited when dealing with the high volatility and randomness of residential loads.
[0009] High computational complexity and poor real-time performance: Complex architectures such as deep residual networks require a large amount of computing resources for training and inference, which makes it difficult to meet real-time requirements in short-term load forecasting scenarios that require fast response.
[0010] Insufficient interpretability: The "black box" nature of deep learning models makes the prediction results lack interpretability, making it difficult to provide a reliable basis for power grid dispatching decisions.
[0011] While existing research considers additional information such as weather conditions and user behavior to improve prediction accuracy, it still lacks a systematic design framework for feature engineering. Statistical techniques such as quantile regression have been used to enhance prediction performance and improve probabilistic load forecasting, but these methods have failed to fully integrate the complementary advantages of multi-source heterogeneous features.
[0012] More critically, existing methods generally lack adaptive hyperparameter optimization mechanisms. Whether it's ensemble learning algorithms like random forests, gradient boosting decision trees (GBDT), and XGBoost in traditional machine learning methods, or various neural network architectures in deep learning, their hyperparameter settings typically rely on human experience or simple grid search, making it difficult to maintain optimal performance under dynamically changing load patterns. Furthermore, the lack of multi-objective optimization functions designed specifically for load forecasting tasks makes it impossible to balance the trade-off between prediction accuracy and model stability.
[0013] In summary, the two existing categories of power load forecasting technologies have significant shortcomings in terms of the systematic nature of parameter optimization, model effectiveness, and multi-source feature fusion, making it difficult to meet the actual needs of smart grids for high-precision, high-stability, and highly adaptive load forecasting. Summary of the Invention
[0014] The purpose of this invention is to overcome the significant shortcomings of existing power load forecasting methods in terms of the systematic nature of parameter optimization, model effectiveness, and multi-source feature fusion. This invention proposes a load forecasting method and system based on ensemble learning and intelligent hyperparameter optimization.
[0015] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a load forecasting method based on ensemble learning and intelligent hyperparameter optimization, comprising the following steps: Acquire multi-source input data; multi-source input data should include at least historical load data, meteorological data, and calendar information; A multi-source heterogeneous feature set is constructed based on multi-source input data; the multi-source heterogeneous feature set includes at least time periodic features, meteorological correlation features, holiday identification features, and historical load statistics features; Intelligent optimization algorithms are used to perform hyperparameter co-optimization on multiple ensemble learning base models to obtain the optimal hyperparameter configuration for each ensemble learning base model. Based on the optimal hyperparameter configuration and multi-source heterogeneous feature set, each ensemble learning base model is trained to obtain each trained ensemble learning base model and its fitting parameters, and the performance evaluation results of each trained ensemble learning base model are obtained simultaneously. Based on the training and performance evaluation results of each post-training ensemble learning base model, a dynamic weighted ensemble strategy is adopted to fuse the prediction results of each post-training ensemble learning base model to obtain the preliminary ensemble prediction value. The residuals are calculated based on the preliminary integrated forecast values and the actual load values. A residual feature set is constructed based on the residuals, and a residual correction model is trained to obtain the residual correction model and its fitting parameters. The deviation values are predicted by the residual correction model, and the deviation values are superimposed with the preliminary integrated forecast values to complete the systematic deviation correction and obtain the final load forecast values.
[0016] Furthermore, when constructing a multi-source heterogeneous feature set based on multi-source input data, a standardized dataset is used for construction; The standardized dataset is obtained by preprocessing multi-source data. Data preprocessing includes outlier removal, missing value imputation, and data standardization. Outlier removal uses the interquartile range method, missing value imputation uses linear interpolation, and data standardization uses Z-score standardization or min-max standardization. After obtaining the final load forecast, a rolling window update mechanism is used to dynamically update the standardized dataset, multi-source heterogeneous feature set, fitting parameters of the ensemble learning base model, and fitting parameters of the residual correction model based on the final load forecast and newly acquired measured data, thereby completing the adaptive iterative update of the forecast model.
[0017] Furthermore, the time periodicity feature is achieved through sine-cosine coding, which maps time information such as hours, dates, and days of the week to a continuous space; Meteorological correlation characteristics were obtained through Pearson correlation coefficient analysis. The screening criteria were that the absolute value of the correlation between meteorological elements and load was greater than a preset threshold, and the meteorological correlation characteristics also included derived cooling degree day and heating degree day characteristics.
[0018] Furthermore, the intelligent optimization algorithm is a particle swarm optimization algorithm. The particle swarm optimization algorithm adopts an adaptive inertia weight mechanism, inertia weight is dynamically adjusted with the number of iterations, and the particle swarm optimization algorithm introduces an early stopping mechanism. When the improvement of the global optimal solution in consecutive preset iterations is less than a preset threshold, the hyperparameter optimization process is terminated in advance.
[0019] Furthermore, the multiple ensemble learning base models include random forest, gradient boosting decision tree, and extreme gradient boosting; the hyperparameter search space of each ensemble learning base model includes at least two of the following: number of decision trees, tree depth, learning rate, sample sampling ratio, and feature sampling ratio.
[0020] Furthermore, the dynamic weighted integration strategy is specifically as follows: Performance scores are calculated based on the mean absolute percentage error (MAPE) of each post-trained ensemble learning base model on the validation set. The performance scores are negatively correlated with MAPE. The performance scores are normalized using the Softmax function to obtain the dynamic weights of each ensemble learning base model. The prediction results of each model are summed with their corresponding dynamic weights to obtain the preliminary integrated prediction value.
[0021] Furthermore, the residual feature set includes the lagged term features of the residuals, the rolling statistical features of the residuals, and the time pattern features of the residuals; The residual correction model is a lightweight regression model, specifically a ridge regression model or a shallow gradient boosting model. The rolling window update mechanism uses a fixed time window with a window size of 30-60 days. During the update, historical data is replaced according to the first-in-first-out principle. When the final load forecast value MAPE exceeds the preset threshold, an emergency retraining process is triggered.
[0022] Secondly, the present invention provides a load forecasting system based on ensemble learning and intelligent hyperparameter optimization, comprising: The dataset acquisition module is used to acquire multi-source input data; the multi-source input data includes at least historical load data, meteorological data, and calendar information. The feature set construction module is used to construct a multi-source heterogeneous feature set based on multi-source input data; the multi-source heterogeneous feature set includes at least time periodic features, meteorological correlation features, holiday identification features, and historical load statistics features; The training model module is used to perform hyperparameter co-optimization on multiple ensemble learning base models using intelligent optimization algorithms to obtain the optimal hyperparameter configuration for each ensemble learning base model; based on the optimal hyperparameter configuration and multi-source heterogeneous feature set, each ensemble learning base model is trained to obtain each trained ensemble learning base model and its fitting parameters, and the performance evaluation results of each trained ensemble learning base model are obtained simultaneously. The preliminary prediction module is used to fuse the prediction results of each post-trained ensemble learning base model and the performance evaluation results using a dynamic weighted ensemble strategy to obtain the preliminary ensemble prediction value. The final prediction module is used to calculate the residuals based on the preliminary integrated prediction values and the actual load values, construct a residual feature set based on the residuals and train the residual correction model to obtain the residual correction model and its fitting parameters; predict the deviation value through the residual correction model, and complete the systematic deviation correction by superimposing the deviation value with the preliminary integrated prediction value to obtain the final load prediction value.
[0023] Thirdly, the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a load forecasting method based on ensemble learning and hyperparameter intelligent optimization.
[0024] Fourthly, the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements a load forecasting method based on ensemble learning and intelligent hyperparameter optimization.
[0025] Compared with the prior art, the present invention has the following beneficial technical effects: This invention proposes a load forecasting method based on ensemble learning and intelligent hyperparameter optimization, addressing the parameter sensitivity problem of traditional machine learning and improving forecast stability. It employs an intelligent optimization algorithm to collaboratively optimize the hyperparameters of multiple ensemble learning base models, replacing manual experience or simple grid search for parameter setting. This ensures that each base model is always in an optimal parameter configuration state, significantly reducing the impact of parameter fluctuations on forecast performance and avoiding orders-of-magnitude differences in forecast errors. It balances forecast accuracy and timeliness, adapting to practical deployment needs. By adopting ensemble learning to replace complex deep learning architectures, and combining a two-stage design of dynamic weighted ensemble and lightweight residual correction, it significantly reduces the computational cost of model training and inference while maintaining forecast accuracy. This avoids the dependence of deep learning on massive labeled data and the risk of overfitting, meeting the real-time requirements of short-term load forecasting. It achieves systematic fusion of multi-source heterogeneous features, enhancing the ability to characterize load patterns. A multi-dimensional feature set covering time periodicity, meteorological correlation, holiday identification, and historical load statistics is constructed, fully leveraging the complementary advantages of multi-source data to improve forecast adaptability in complex scenarios such as extreme weather and holidays, addressing the lack of systematic feature fusion in existing methods. Attached Figure Description
[0026] The accompanying drawings described herein are for illustrative purposes only and are not intended to limit the scope of the invention in any way. Furthermore, the shapes and proportions of the components in the drawings are merely illustrative to aid in understanding the invention and do not specifically limit the shapes and proportions of the components of the invention.
[0027] In the attached diagram: Figure 1 This is a flowchart of the load forecasting method based on ensemble learning and intelligent hyperparameter optimization according to the present invention.
[0028] Figure 2 This is a simplified structural diagram of the load forecasting system based on ensemble learning and intelligent hyperparameter optimization of the present invention.
[0029] Figure 3 This is a schematic diagram of an electronic device based on the load prediction method of the present invention, which is based on ensemble learning and intelligent hyperparameter optimization.
[0030] Figure 4 This is a flowchart of the load forecasting method based on ensemble learning and intelligent hyperparameter optimization in an embodiment of the present invention. Detailed Implementation
[0031] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0032] Example 1 See Figure 1 A load forecasting method based on ensemble learning and intelligent hyperparameter optimization includes the following steps: The process involves acquiring and preprocessing multi-source input data to obtain a standardized dataset. The multi-source input data includes at least historical load data, meteorological data, and calendar information. A multi-source heterogeneous feature set is constructed based on the standardized dataset. This feature set includes at least time-periodic features, meteorological correlation features, holiday identification features, and historical load statistical features. An intelligent optimization algorithm is used to collaboratively optimize the hyperparameters of multiple ensemble learning base models, obtaining the optimal hyperparameter configuration for each model. Based on the optimal hyperparameter configuration and the multi-source heterogeneous feature set, each ensemble learning base model is trained, yielding each trained ensemble learning base model and its fitting parameters. Simultaneously, the performance evaluation results of each trained ensemble learning base model are acquired. Based on the trained ensemble learning base models and their performance evaluation results, a dynamic weighted ensemble strategy is used to fuse the prediction results of each model, obtaining a preliminary ensemble prediction value. The residuals are calculated based on the preliminary ensemble prediction value and the actual load value. A residual feature set is constructed based on the residuals, and a residual correction model is trained, obtaining the residual correction model and its fitting parameters. The deviation value is predicted using the residual correction model, and the deviation value is superimposed with the preliminary ensemble prediction value to complete systematic deviation correction, resulting in the final load prediction value.
[0033] At the data level, this embodiment integrates heterogeneous data from multiple sources, including historical loads, meteorological data, and calendar data, and constructs diverse feature sets encompassing time periods, meteorological correlations, and holiday markers to fully explore potential data relationships and enhance the richness and representativeness of the model input. At the model level, intelligent optimization algorithms are used to collaboratively tune the hyperparameters of multiple ensemble learning base models, ensuring that each model performs optimally. A dynamic weighted ensemble strategy is then used to fuse the prediction results, effectively reducing the bias of individual models and improving prediction robustness. At the accuracy level, a residual correction mechanism is introduced, utilizing residual features to further correct initial prediction biases, significantly improving the final prediction accuracy.
[0034] When constructing a multi-source heterogeneous feature set based on multi-source input data, a standardized dataset is used. The standardized dataset is obtained by preprocessing the multi-source data, including outlier removal, missing value imputation, and data standardization. Outlier removal uses the interquartile range method, missing value imputation uses linear interpolation, and data standardization uses Z-score standardization or min-max standardization. After obtaining the final load forecast, a rolling window update mechanism is used to dynamically update the standardized dataset, multi-source heterogeneous feature set, fitting parameters of the ensemble learning base model, and fitting parameters of the residual correction model based on the final load forecast and newly acquired measured data, thereby completing the adaptive iterative update of the forecast model.
[0035] This embodiment features dynamic adaptive capabilities, ensuring long-term forecast reliability. Through a rolling window update mechanism, the dataset, feature set, and model fitting parameters are dynamically updated based on new measured data. Combined with performance monitoring triggering emergency retraining, the model continuously adapts to dynamic changes in load patterns, significantly reducing deployment and maintenance costs and providing high-precision, high-stability load forecasting support for smart grid dispatching decisions. At the adaptive level, the rolling window update mechanism dynamically adjusts the dataset, feature set, and model parameters, enabling the model to quickly adapt to changes in load patterns and achieve long-term stable forecasts, providing reliable decision support for power dispatching and energy planning.
[0036] The load forecasting method in this embodiment includes four main stages: data preprocessing, multi-model construction, dynamic integration, and forecast output. Each module forms a complete forecasting architecture.
[0037] Process begins: The system initiates a short-term load forecasting task to obtain the forecast demand at the current time T, with the goal of forecasting the load value for the next 1-7 days.
[0038] Step 1: Update Historical Data: The system updates the historical load database, extracting historical load data from TN to T as the modeling data source, where N represents the historical data window length (N=60). Simultaneously, it acquires meteorological data (temperature, humidity, wind speed, precipitation, air pressure, etc.) and calendar information (date, time, day of the week, holiday identifiers, etc.) for the corresponding time period. The data is aligned by timestamps to form a complete multi-source dataset. Data quality is checked, outliers are removed using the IQR method, missing values are filled using linear interpolation, and meteorological characteristics are standardized.
[0039] Step 2: Multi-source Feature Construction: Constructing a multi-dimensional feature representation system. First, periodic encoding of time features is performed, mapping hourly indices to a unit circle using a sine-cosine transform to generate two time-period features. Next, weather feature processing is performed, calculating the Pearson correlation coefficient between each meteorological element and historical loads, and selecting significant features with an absolute correlation value greater than 0.3. Then, holiday features are constructed, setting three layers of identification: a binary identifier for whether it is a holiday, a holiday type code, and a gradual weighting effect for the three days before and after the holiday. Finally, rolling window statistical features are constructed based on preceding loads, generating approximately 10-15 feature variables in total.
[0040] Step 3: Dataset Partitioning: Divide the feature data into training and validation sets based on time. The training set consists of historical data from time TN to T-1, used for model parameter learning; the validation set consists of data from T-1 to the most recent day (96 time points), used for hyperparameter optimization and model performance evaluation. The prediction set consists of data features at time T, used to generate future load forecasts.
[0041] Step 4: Model Training – PSO Hyperparameter Optimization: Enter the model training phase and perform PSO hyperparameter optimization for the three base learners: RF, GBDT, and XGBoost. For each model, the hyperparameter search space is defined as follows: RF: n_estimators (5-300), max_depth (5-30), min_samples_split (2-20), min_samples_leaf (1-10); GBDT: n_estimators (5-300), learning_rate (0.01-0.3), max_depth (3-15), subsample (0.6-1.0); XGBoost: n_estimators (5-300), learning_rate (0.01-0.3), max_depth (3-15), subsample (0.6-1.0), colsample_bytree (0.6-1.0). Initialize the particle swarm, setting the particle count to 50 and the maximum number of iterations to 20. Each particle represents a set of hyperparameter configurations, and the particle position and velocity are randomly initialized. In each iteration, the hyperparameter configuration of each particle is configured to train the model, and MAPE and stability indices are calculated on the validation set based on the multi-objective function. Evaluate particle fitness. Update individual optimal positions and global optimal positions. Employ adaptive inertia weights. Dynamically adjust the search strategy and update the formula using speed and position. and Iterative optimization is employed. An early stopping mechanism is introduced, terminating the process when the improvement in the global optimum is less than 0.001 over 10 consecutive generations. After PSO optimization, the optimal hyperparameter configurations for RF, GBDT, and XGBoost are obtained respectively.
[0042] Step 5: Ensemble Learning – Training Prediction Models: Using the optimal hyperparameters obtained through PSO optimization, train three base learners (RF, GBDT, and XGBoost) on the full training set. After training, obtain the prediction results (y_pred_RF, y_pred_GBDT, y_pred_XGB) of the three models on the validation set, and calculate the MAPE value for each model. Calculate the performance score based on the MAPE. Dynamic weights are calculated using Softmax normalization, specifically by first calculating... To avoid numerical overflow, normalization is performed again. The weight coefficients of the three models are obtained. The predictions from the three models are then weighted and integrated using dynamic weights to obtain the integrated prediction value for the first stage. .
[0043] Step 6: Generate Predictions: Apply the trained ensemble model to the prediction set (feature data at time T) to generate preliminary future load predictions.
[0044] Step 7: Prediction Correction – Bias Model Correction: To further improve prediction accuracy, a second-stage residual correction model is constructed. First, the ensemble prediction residuals on the validation set are calculated. Then, residual features are constructed, including the lag terms of the residuals. The rolling statistical features of the residuals are the mean and standard deviation of the residuals at the previous 12 time points, as well as the temporal pattern features of the residuals, such as the average residual pattern statistically analyzed hourly. The original features and residual features are merged to form an enhanced feature set. A lightweight residual model is trained (in this embodiment, Ridge regression (alpha=1.0) is used as the residual model; shallow XGBoost can also be used), with the goal of learning systematic bias patterns. The residual model is used to correct the predicted value at time T, generating a residual correction value. The final prediction result is This completes the second phase of systematic deviation correction.
[0045] Step 8: Determine whether to stop: The system checks whether the termination conditions of the prediction task have been met. If the current time has reached the end time of the prediction cycle or a stop command has been received, the process ends, and the final prediction results, including the prediction timestamp, prediction load value, model version, and other information, are output and saved to the database; if the termination conditions have not been met, the rolling update process begins.
[0046] Step 9: Train the bias model using historical forecast errors: During continuous forecasting, as real load data is continuously acquired, the system accumulates historical forecast error data. Periodically (e.g., daily or weekly), the bias model is retrained using the accumulated historical forecast errors to update the statistical and temporal patterns of the residual characteristics, enabling the bias correction model to better adapt to changes in load patterns.
[0047] Step 10: Obtain Rolling Data: The system acquires the latest measured load values and corresponding meteorological data, adding them to the historical database. A fixed-window rolling strategy is adopted, removing the oldest data points while maintaining a constant training window size (e.g., 30 days of data). Simultaneously, it checks whether the hyperparameter optimization cycle (e.g., every 7 days) has been reached; if so, it marks the need to re-execute PSO optimization.
[0048] Step 11: Return to Step 1: The system returns to the "Update Historical Data" step and begins a new prediction cycle. In this new cycle, if the hyperparameter optimization period has not been reached, the cached optimal hyperparameters are used directly for fast prediction; if the optimization period has been reached, the complete PSO hyperparameter optimization process is re-executed, and feature importance is re-evaluated, eliminating redundant features with importance below the threshold to achieve adaptive model updates. The system continuously monitors prediction performance metrics, and triggers an emergency retraining mechanism when MAPE exceeds a preset threshold to ensure long-term stability of prediction quality.
[0049] Through the cyclical execution of the above process, the system achieves continuous rolling forecasting of short-term loads. By employing multi-source feature fusion, adaptive hyperparameter optimization, two-stage cascaded forecasting, and a rolling update mechanism, it ensures high accuracy, high stability, and strong adaptability in forecasting. The entire process fully leverages the complementary advantages of historical load data, meteorological data, and calendar information. Through intelligent feature engineering and model optimization techniques, it effectively addresses the problems of traditional methods, such as high parameter sensitivity, insufficient effectiveness of deep learning, lack of systematic feature fusion, and insufficient model adaptability, providing an efficient and reliable technical solution for load forecasting.
[0050] The load forecasting method based on ensemble learning and intelligent hyperparameter optimization employs a two-stage cascaded architecture based on multi-source feature fusion and adaptive hyperparameter optimization to achieve load forecasting. (See [link to relevant documentation]). Figure 4 The overall technical architecture includes four core modules: 1. Multi-source heterogeneous feature intelligent fusion system This module constructs a multi-dimensional feature representation system. For time features, a periodic encoding technique is used to map hourly time information to a continuous space through a sine-cosine transform. For example, the hourly feature is encoded as follows:
[0051] For weather characteristics, the correlation between meteorological elements and load is analyzed using Pearson correlation coefficients. Strongly correlated characteristics are dynamically screened, and derived characteristics such as cooling degree day (CDD) and heating degree day (HDD) are constructed. For holiday characteristics, a three-layer identification system is established, including whether it is a holiday, the type of holiday, and the pre- and post-holiday effects. For historical load characteristics, statistical characteristics based on a preceding rolling window are constructed.
[0052] 2. Adaptive Hyperparameter Co-optimization Based on PSO This module designs an optimization function that integrates prediction accuracy and stability for three base learners: RF, GBDT, and XGBoost.
[0053] Where MAPE is the mean absolute percentage error, and Stability is the prediction stability index (the coefficient of variation of the absolute percentage error). The PSO algorithm is used for hyperparameter optimization, and an adaptive inertia weighting mechanism is introduced.
[0054] The weights are dynamically adjusted with the number of iterations, maintaining a large global search capability in the early stages and enhancing the local fine-grained search capability in the later stages. An early stopping mechanism is also introduced, terminating the optimization process prematurely when the global optimum has not significantly improved over several consecutive iterations, thereby improving optimization efficiency.
[0055] Three: Two-stage cascaded prediction architecture The first stage employs a dynamic weighting mechanism to integrate the prediction results from RF, GBDT, and XGBoost. The MAPE performance score of each model is calculated based on its performance on the validation set. Dynamic weights are obtained through Softmax normalization:
[0056] The ensemble prediction results are as follows:
[0057] The second stage involves constructing residual features, including lagged terms of historical residuals, rolling statistical features, and time pattern features of the residuals. A lightweight residual model (ridge regression model) is then trained to correct for systematic biases. The final prediction result is as follows: .
[0058] 4. Adaptive Mechanism under Rolling Updates A fixed-window rolling update strategy is adopted, maintaining the training set within a fixed time window (e.g., 30 days), and updating it according to the FIFO principle each time new data is acquired. Feature importance evaluation and hyperparameter optimization are performed periodically (e.g., every 7 days), and the cached optimal parameters are used for fast prediction during non-optimization periods. Simultaneously, performance monitoring is implemented, triggering an emergency retraining mechanism when the prediction error exceeds a threshold, ensuring the model adapts quickly to changes in load patterns.
[0059] This embodiment employs a load forecasting method based on ensemble learning and intelligent hyperparameter optimization. Adaptive optimization and a two-stage architecture achieve a balance between accuracy and efficiency. This embodiment uses PSO-based adaptive hyperparameter optimization to design a multi-objective function that integrates MAPE and stability. A two-stage cascaded architecture is adopted: the first stage uses Softmax to dynamically weight and integrate RF, GBDT, and XGBoost; the second stage uses a lightweight residual model to eliminate systematic biases. This avoids the high computational cost of deep learning, enabling fast training, real-time inference, and large-scale deployment. Multi-source heterogeneous feature fusion comprehensively characterizes load patterns, constructing a complete feature system: periodic encoding of temporal features to capture multi-scale patterns; dynamic selection of meteorological features; multi-level holiday identification; and rolling window statistical features based on preceding loads. Characterizing loads from three dimensions—time, meteorology, and historical loads—it improves the forecasting capability for complex scenarios (extreme weather, holidays, sudden changes). A rolling update mechanism ensures long-term stability, using a fixed 60-day rolling update window. After acquiring new data, the training set is automatically updated, feature importance is re-evaluated, hyperparameters are optimized, and redundant features are dynamically removed. Real-time monitoring is introduced, triggering emergency retraining when MAPE exceeds a threshold. This enables the model to continuously learn the latest patterns, maintain good generalization ability in different scenarios, and significantly reduce deployment and maintenance costs.
[0060] The core technical solution of this embodiment contains several replaceable technical paths to adapt to different scenarios or achieve performance improvements. At the feature engineering level, the periodic encoding of time features can be replaced with Fourier transform or wavelet transform to capture periodicity more precisely; correlation analysis of meteorological features can be replaced with nonlinear measurement methods such as mutual information and maximum information coefficient, or autoencoder feature learning. At the hyperparameter optimization level, the PSO algorithm can be replaced with Bayesian optimization, genetic algorithms, TPE, etc.; MAPE in the objective function can be replaced with indicators such as RMSE, MAE, and quantile loss. At the model ensemble level, the base model can be replaced with other algorithms such as LightGBM, CatBoost, and SVR; the dynamic weighting mechanism can be replaced with Stacking, Blending, or attention mechanism fusion; the residual correction model can be replaced with linear regression, ridge regression, or recurrent neural networks such as LSTM and GRU. At the rolling update level, the fixed window can be replaced with progressively expanding windows, exponentially weighted moving average windows, or online learning and incremental learning algorithms; feature importance evaluation can be replaced with SHAP values, permutation importance, etc. These replacement methods, while maintaining the consistency of the overall framework, provide flexible options for optimization in different application scenarios.
[0061] This embodiment proposes an adaptive hyperparameter co-optimization mechanism based on particle swarm optimization (PSO). It designs a multi-objective function that integrates mean absolute percentage error (MAPE) and prediction stability. By introducing adaptive inertia weights and an early stopping mechanism, it achieves efficient parameter optimization and ensures the stability of the model under different load scenarios.
[0062] This embodiment uses a traditional machine learning ensemble method to replace deep learning. Through a two-stage cascaded prediction architecture, the prediction results of RF, GBDT, and XGBoost are integrated in the first stage using a dynamic weighting mechanism. In the second stage, a lightweight residual model is built to perform systematic bias correction. This not only ensures prediction accuracy but also significantly reduces computational complexity and improves the effectiveness of the model.
[0063] This embodiment constructs an intelligent fusion system of multi-source heterogeneous features, including periodic encoding of time features (sine-cosine transform), correlation analysis and dynamic selection of weather features, multi-level identification of holidays, and rolling window statistical features based on previous loads, forming a multi-dimensional feature representation system to comprehensively depict the load change pattern.
[0064] This embodiment designs an adaptive mechanism under rolling updates, adopting a fixed window rolling update strategy. Each time new data is acquired, the importance of features is re-evaluated and hyperparameters are optimized to enable the model to quickly adapt to changes in load patterns and maintain the stability of long-term prediction performance.
[0065] Example 2 See Figure 2 A load forecasting system based on ensemble learning and intelligent hyperparameter optimization includes: The dataset acquisition module is used to acquire multi-source input data; the multi-source input data includes at least historical load data, meteorological data, and calendar information. The feature set construction module is used to construct a multi-source heterogeneous feature set based on multi-source input data; the multi-source heterogeneous feature set includes at least time periodic features, meteorological correlation features, holiday identification features, and historical load statistics features; The training model module is used to perform hyperparameter co-optimization on multiple ensemble learning base models using intelligent optimization algorithms to obtain the optimal hyperparameter configuration for each ensemble learning base model; based on the optimal hyperparameter configuration and multi-source heterogeneous feature set, each ensemble learning base model is trained to obtain each trained ensemble learning base model and its fitting parameters, and the performance evaluation results of each trained ensemble learning base model are obtained simultaneously. The preliminary prediction module is used to fuse the prediction results of each post-trained ensemble learning base model and the performance evaluation results using a dynamic weighted ensemble strategy to obtain the preliminary ensemble prediction value. The final prediction module is used to calculate the residuals based on the preliminary integrated prediction values and the actual load values, construct a residual feature set based on the residuals and train the residual correction model to obtain the residual correction model and its fitting parameters; predict the deviation value through the residual correction model, and complete the systematic deviation correction by superimposing the deviation value with the preliminary integrated prediction value to obtain the final load prediction value.
[0066] Example 3 See Figure 3An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements a load forecasting method based on ensemble learning and intelligent hyperparameter optimization: acquiring multi-source input data and preprocessing it to obtain a standardized dataset; the multi-source input data includes at least historical load data, meteorological data, and calendar information; constructing a multi-source heterogeneous feature set based on the standardized dataset; the multi-source heterogeneous feature set includes at least time-periodic features, meteorological correlation features, holiday identification features, and historical load statistical features; using an intelligent optimization algorithm to perform hyperparameter co-optimization on multiple ensemble learning base models to obtain the optimal hyperparameter configuration for each ensemble learning base model; and training each ensemble learning base model based on the optimal hyperparameter configuration and the multi-source heterogeneous feature set to obtain each trained ensemble learning base model and its approximate model. The system integrates parameters and simultaneously acquires the performance evaluation results of each trained ensemble learning base model. Based on the training ensemble learning base models and their performance evaluation results, a dynamic weighted ensemble strategy is used to fuse the prediction results of each training ensemble learning base model to obtain preliminary ensemble predictions. The residuals are calculated based on the preliminary ensemble predictions and the actual load values. A residual feature set is constructed based on the residuals, and a residual correction model is trained to obtain the residual correction model and its fitting parameters. The deviation values are predicted using the residual correction model, and the deviation values are superimposed with the preliminary ensemble predictions to complete systematic deviation correction, resulting in the final load prediction value. Based on the final load prediction value and newly acquired measured data, a rolling window update mechanism is used to dynamically update the standardized dataset, the multi-source heterogeneous feature set, the fitting parameters of the ensemble learning base model, and the fitting parameters of the residual correction model, completing the adaptive iterative update of the prediction model.
[0067] Example 4 A computer-readable storage medium stores a computer program that, when executed by a processor, implements a load forecasting method based on ensemble learning and intelligent hyperparameter optimization: acquiring and preprocessing multi-source input data to obtain a standardized dataset; the multi-source input data includes at least historical load data, meteorological data, and calendar information; constructing a multi-source heterogeneous feature set based on the standardized dataset; the multi-source heterogeneous feature set includes at least time-periodic features, meteorological correlation features, holiday identification features, and historical load statistical features; employing an intelligent optimization algorithm to perform hyperparameter co-optimization on multiple ensemble learning base models to obtain the optimal hyperparameter configuration for each ensemble learning base model; training each ensemble learning base model based on the optimal hyperparameter configuration and the multi-source heterogeneous feature set to obtain each trained ensemble learning base model and its fitting parameters; and... The system synchronously acquires the performance evaluation results of each post-trained ensemble learning base model. Based on the performance evaluation results, a dynamic weighted ensemble strategy is used to fuse the prediction results of each model to obtain a preliminary ensemble prediction. The residuals are calculated based on the preliminary ensemble prediction and the actual load value. A residual feature set is constructed based on the residuals, and a residual correction model is trained to obtain the residual correction model and its fitting parameters. The deviation value is predicted using the residual correction model, and the deviation value is superimposed with the preliminary ensemble prediction to complete the systematic deviation correction, resulting in the final load prediction value. Based on the final load prediction value and newly acquired measured data, a rolling window update mechanism is used to dynamically update the standardized dataset, the multi-source heterogeneous feature set, the fitting parameters of the ensemble learning base model, and the fitting parameters of the residual correction model, thus completing the adaptive iterative update of the prediction model.
[0068] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, optical storage, etc.) containing computer-usable program code.
[0069] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1A device that provides the functions specified in one or more boxes.
[0070] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0071] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0072] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the present invention.
Claims
1. A load forecasting method based on ensemble learning and intelligent hyperparameter optimization, characterized in that, Includes the following steps: Acquire multi-source input data; the multi-source input data includes at least historical load data, meteorological data, and calendar information; A multi-source heterogeneous feature set is constructed based on the multi-source input data; the multi-source heterogeneous feature set includes at least time periodic features, meteorological correlation features, holiday identification features, and historical load statistics features; An intelligent optimization algorithm is used to perform hyperparameter co-optimization on multiple ensemble learning base models to obtain the optimal hyperparameter configuration for each ensemble learning base model. Based on the optimal hyperparameter configuration and the multi-source heterogeneous feature set, each ensemble learning base model is trained to obtain each trained ensemble learning base model and its fitting parameters, and the performance evaluation results of each trained ensemble learning base model are obtained simultaneously. Based on the training and performance evaluation results of each post-training ensemble learning base model, a dynamic weighted ensemble strategy is used to fuse the prediction results of each training and ensemble learning base model to obtain a preliminary ensemble prediction value. The residuals are calculated based on the preliminary integrated forecast values and the actual load values. A residual feature set is constructed based on the residuals, and a residual correction model is trained to obtain the residual correction model and its fitting parameters. The deviation values are predicted using the residual correction model, and the deviation values are superimposed with the preliminary integrated forecast values to complete the systematic deviation correction and obtain the final load forecast values.
2. The load forecasting method based on ensemble learning and intelligent hyperparameter optimization according to claim 1, characterized in that, When constructing a multi-source heterogeneous feature set based on the multi-source input data, a standardized dataset is used for construction. The standardized dataset is obtained by preprocessing the multi-source data. The data preprocessing includes outlier removal, missing value imputation, and data standardization. Outlier removal uses the interquartile range method, missing value imputation uses linear interpolation, and data standardization uses Z-score standardization or min-max standardization. After obtaining the final load forecast, based on the final load forecast and the newly acquired measured data, a rolling window update mechanism is used to dynamically update the standardized dataset, the multi-source heterogeneous feature set, the fitting parameters of the ensemble learning base model, and the fitting parameters of the residual correction model, thereby completing the adaptive iterative update of the prediction model.
3. The load forecasting method based on ensemble learning and intelligent hyperparameter optimization according to claim 1, characterized in that, The time periodicity feature is achieved through sine-cosine coding, which maps time information such as hours, dates, and days of the week to a continuous space. The meteorological correlation features were obtained through Pearson correlation coefficient analysis. The screening criteria were that the absolute value of the correlation between meteorological elements and load was greater than a preset threshold, and the meteorological correlation features also included derived cooling degree day and heating degree day features.
4. The load forecasting method based on ensemble learning and intelligent hyperparameter optimization according to claim 1, characterized in that, The intelligent optimization algorithm is a particle swarm optimization algorithm. The particle swarm optimization algorithm adopts an adaptive inertia weight mechanism, inertia weight is dynamically adjusted with the number of iterations, and the particle swarm optimization algorithm introduces an early stopping mechanism. When the improvement of the global optimal solution in consecutive preset iterations is less than a preset threshold, the hyperparameter optimization process is terminated in advance.
5. The load forecasting method based on ensemble learning and intelligent hyperparameter optimization according to claim 1, characterized in that, The multiple ensemble learning base models include random forest, gradient boosting decision tree, and extreme gradient boosting; the hyperparameter search space of each ensemble learning base model includes at least two of the following: number of decision trees, tree depth, learning rate, sample sampling ratio, and feature sampling ratio.
6. The load forecasting method based on ensemble learning and intelligent hyperparameter optimization according to claim 1, characterized in that, The dynamic weighted integration strategy is specifically as follows: Performance scores are calculated based on the mean absolute percentage error (MAPE) of each post-trained ensemble learning base model on the validation set. The performance scores are negatively correlated with MAPE. The performance scores are normalized using the Softmax function to obtain the dynamic weights of each ensemble learning base model. The prediction results of each model are summed with their corresponding dynamic weights to obtain the preliminary integrated prediction value.
7. The load forecasting method based on ensemble learning and intelligent hyperparameter optimization according to claim 1, characterized in that, The residual feature set includes the residual lag term features, the residual rolling statistical features, and the residual time pattern features; The residual correction model is a lightweight regression model, specifically a ridge regression model or a shallow gradient boosting model. The rolling window update mechanism uses a fixed time window with a window size of 30-60 days. During the update, historical data is replaced according to the first-in-first-out principle. When the final load forecast value MAPE exceeds the preset threshold, an emergency retraining process is triggered.
8. A load forecasting system based on ensemble learning and intelligent hyperparameter optimization, characterized in that, include: The dataset acquisition module is used to acquire input data from multiple sources; The multi-source input data includes at least historical load data, meteorological data, and calendar information; The feature set construction module is used to construct a multi-source heterogeneous feature set based on the multi-source input data; the multi-source heterogeneous feature set includes at least time periodic features, meteorological correlation features, holiday identification features, and historical load statistics features; The model training module is used to perform hyperparameter co-optimization on multiple ensemble learning base models using an intelligent optimization algorithm to obtain the optimal hyperparameter configuration of each ensemble learning base model; based on the optimal hyperparameter configuration and the multi-source heterogeneous feature set, each ensemble learning base model is trained to obtain each trained ensemble learning base model and its fitting parameters, and the performance evaluation results of each trained ensemble learning base model are obtained simultaneously. The preliminary prediction module is used to fuse the prediction results of each post-trained ensemble learning base model and the performance evaluation results using a dynamic weighted ensemble strategy to obtain a preliminary ensemble prediction value. The final prediction module is used to calculate the residual based on the preliminary integrated prediction value and the actual load value, construct a residual feature set based on the residual and train a residual correction model to obtain the residual correction model and its fitting parameters; predict the deviation value through the residual correction model, and superimpose the deviation value with the preliminary integrated prediction value to complete the systematic deviation correction and obtain the final load prediction value.
9. An electronic device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the load forecasting method based on ensemble learning and hyperparameter intelligent optimization as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the load forecasting method based on ensemble learning and hyperparameter intelligent optimization as described in any one of claims 1-7.