Heterogeneous integrated short-term multi-element load prediction method and system
The heterogeneous Stacking integrated prediction model is constructed through multiple seasonal trend decomposition and non-dominant sorting genetic optimization algorithms, which solves the problems of complex coupling relationships and seasonal differentiation characteristics of multiple load prediction in integrated energy systems, and improves prediction accuracy and interpretability.
Patent Information
- Application Number
- CN202510534508.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art has problems such as insufficient processing of complex coupling relationships, insufficient mining capabilities for seasonal differentiation characteristics, learning device combination optimization problems and insufficient interpretability in the prediction of multiple loads in integrated energy systems.
Multiple seasonal trend decomposition methods are used to decompose multiple loads into periodic sequences and trend sequences of multiple time scales to construct differentiated input features. The optimal learner combination was screened through a non-dominant sorting genetic optimization algorithm with elite strategies, a heterogeneous Stacking integrated prediction model was constructed, and attribution analysis was performed through the Shapley value additive interpretation method.
It improves the accuracy and stability of multi-load prediction, enhances the interpretability of the model, and facilitates decision-making and optimization in practical applications.
Smart Images

Figure CN120047017A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of integrated energy system load forecasting, and more specifically, to a heterogeneous integration short-term multi-load forecasting method and system. Background Art
[0002] With the increasing global energy demand, the greenhouse gas emissions have risen significantly, leading to severe challenges to green and sustainable development due to global environmental deterioration and climate change. Building a new power system with new energy as the main body is one of the key measures to reduce carbon emissions and improve energy utilization efficiency. The integrated energy system (IES) promotes the complementary coordination and cascade utilization of energy by coupling multiple heterogeneous energy systems, improving energy utilization efficiency. However, the randomness and uncertainty brought by energy coupling and the introduction of new energy in the IES make it a huge challenge to achieve accurate and stable multi-load forecasting in the IES.
[0003] The forecasting of multi-loads (such as electric load, heat load, and cooling load) in the IES is a key link to achieve energy optimization scheduling and efficient resource utilization. However, different from traditional single energy systems, the multi-loads in the IES have significant diversity, complexity, and coupling. The multi-loads are not only affected by their own energy consumption patterns but also have non-linear coupling relationships with other load types in the system. In addition, the influence of external factors such as seasons and holidays on the load exacerbates the complexity of its time series characteristics, including significant strong volatility and non-stationarity, which makes it a huge challenge to accurately forecast the multi-loads in the IES.
[0004] Traditional load forecasting methods usually target a single energy type and are mainly divided into forecasting methods based on mathematical statistics (such as regression analysis, time series method) and forecasting methods based on artificial intelligence (such as random forest, support vector machine, deep learning). Although these methods show a certain degree of accuracy in single load forecasting, they are unable to handle the complex coupling relationships of multi-loads in the IES. In addition, recent research has begun to introduce ensemble learning strategies to improve forecasting accuracy by combining multiple learners. However, the existing ensemble learning models generally have the following problems: (1) It is difficult to fully exploit the complex coupling characteristics among multi-loads in the IES: Most studies directly adopt single load forecasting methods and fail to deeply explore the strong coupling and interaction laws among multi-loads such as electricity, cooling, and heat in the IES, ignoring the synergy generated by energy complementarity and equipment coupling among multi-loads in the IES, resulting in the prediction model being difficult to accurately capture the potential energy consumption patterns of users in practical applications and the prediction accuracy being reduced; (2)Insufficiencies of decomposition techniques in high-frequency component prediction: Although signal decomposition techniques (such as VMD) can reduce data complexity, the prediction accuracy of high-frequency components is still low. Especially when the load sequence contains significant volatility and non-stationary characteristics, existing decomposition strategies are prone to introducing additional errors, thus affecting the overall prediction effect. (3)Lack of balance between the diversity and accuracy of model selection: In ensemble learning, usually only the model accuracy is concerned, while the diversity of base learners is ignored; or after introducing diversity, the model becomes too complex and the computational cost is too high. (4)Insufficiencies in model transparency and interpretability: As the complexity of prediction models continues to increase, existing methods are difficult to provide clear explanations, which limits the credibility and usability of the models in actual scenarios.
[0005] Therefore, how to improve the accuracy and stability of IES multi-load forecasting and provide reliable technical support for the optimal operation of integrated energy systems is a technical problem that those skilled in the art urgently need to solve. Summary of the Invention
[0006] In view of this, the present invention provides a heterogeneous integrated short-term multi-load forecasting method and system, which solves the problems existing in the background technology.
[0007] To achieve the above object, the present invention provides the following technical solutions: A heterogeneous integrated short-term multi-load forecasting method includes the following steps: S1: Based on multiple seasonal trend decomposition, decompose the multi-load historical sequence into periodic sequences and trend sequences on multiple time scales, and construct differentiated input features. S2: With diversity and accuracy as optimization objectives, screen the optimal learner combination suitable for different seasons and loads through the non-dominated sorting genetic optimization algorithm with an elite strategy, and construct the most excellent heterogeneous Stacking integrated forecasting model. S3: Use the Shapley value additivity interpretation method for attribution analysis to measure the contribution degree of each input feature to the most excellent heterogeneous Stacking integrated forecasting model from both global and individual dimensions.
[0008] Optionally, it further includes: Before constructing the differentiated input features, perform multi-load characteristic analysis, which specifically includes the following steps: Analyze the multi-load using the autocorrelation coefficient calculation formula of Equation (1) to test whether the multi-load historical sequence has short-term and long-term repeating patterns. (1); In the formula: represents the time delay, represents the time delay of The autocorrelation coefficient of the load denotes t the load value at time denotes the historical load mean value denotes the load value at time denotes the historical load variance; the value range of the autocorrelation coefficient is [-1, 1], the larger it is, the stronger the correlation; E denotes taking the mean value; Use the Spearman rank correlation analysis calculation formula in Equation (2) to analyze the historical data of multi - load in different seasons, and obtain the coupling characteristics of multi - load in different seasons and the correlation between multi - load and meteorological factors; (2); In the formula: r denotes the rank correlation coefficient between different loads, denotes the rank difference between two data variables, n denotes the total number of observed samples; r The value range of is [-1, 1], and the larger it is, the stronger the correlation.
[0009] Optionally, in S1, constructing differential input features specifically includes the following steps: Using the multi - seasonal trend decomposition algorithm in Equation (3), additively decompose the historical multi - load sequence into a trend component, a residual component, and multiple seasonal components; (3); In the formula: denotes the historical multi - load sequence, denotes the trend component of the load sequence, denotes the residual component of the load sequence, denotes the p th p seasonal component corresponding to the th time - scale period existing in the sequence;
[0010] Optionally, in S1, the multi - seasonal trend decomposition is divided into an inner loop and an outer loop, with the inner loop nested in the outer loop. The specific decomposition process is as follows: Set the original sequence as and the number of each time - scale period existing in the sequence as p , and obtain the period array q sorted from smallest to largest time - scale; at the same time, all seasonal components Set the initial value of d to 0, and the non-seasonal component x has an initial value of l ; In the -th outer loop and the -th inner loop, first, update the non-seasonal component , where is the seasonal component of the k -1-th time the outer loop decomposed to obtain the i -th time scale period. When k = 1, is the initial value 0 of the seasonal component; second, use the STL algorithm to decompose the non-seasonal component to obtain the seasonal component of the -th time scale decomposed in the -th outer loop, and update the non-seasonal component ; when the traversal of all q period elements in the period array p is completed, the inner loop terminates and a new round of outer loop starts again until the l -th outer loop, the outer loop ends, and p multiple seasonal components are decomposed as the periodic characteristics of each time scale of the multivariate load; represents the non-seasonal component after the first update, and represents the non-seasonal component after the second update; According to the p multiple seasonal components and the original sequence, calculate the trend component and the residual component, and use the trend component as the long-term change trend characteristic of the multivariate load.
[0011] Optionally, in S2, the individual learner in the heterogeneous Stacking ensemble prediction model is the base learner, and the learner that combines the results of the preliminary prediction individual learners and performs secondary prediction to output the final prediction result is the meta-learner; the ensemble learning method based on Stacking is specifically: Randomly divide the data set into k subsets of equal size and non-overlapping, denoted as , , respectively define and as the K -th fold test set and training set in k -fold cross-validation; the first-layer prediction model contains K base learners, and use for the training set kA base model is obtained through algorithm training ; among them, is the feature vector of the n th sample, is the predicted value of the n th sample, m is the number of features included, and each feature vector can be expressed as ; For each sample K in the k th fold test set of k-fold cross-validation , the prediction result of the base learner is denoted as ; after the cross-validation process ends, a new data set is formed based on the output results of K base learners, denoted as ; Based on the newly formed data set , train the second-layer meta-learner model to obtain the optimal prediction result.
[0012] Optionally, in S2, the heterogeneous Stacking ensemble prediction model selects Catboost, DT, KNN, Lasso, RF, Ridge, SVM, XGBoost, LSTM, and GRU to construct a diverse heterogeneous base model library.
[0013] Optionally, in S2, to construct the most excellent heterogeneous Stacking ensemble prediction model, the specific steps are as follows: Use the selective ensemble method of evolutionary multi-objective optimization to screen the base learners, and the optimization problem is shown in the following formula: (22); In the formula: and are the objective functions for measuring the accuracy and diversity of individual learners respectively; The definition formula of the prediction accuracy objective function is: (23); In the formula: n represents the number of samples in the training set, represents the prediction output of the base learners selected after integration for the i th training sample in the training set, represents the i th actual observation value in the training set, represents the average value of all actual observation values in the training set; is The specific definition formula, that is, the specific expression of the accuracy objective function of the individual learner in the evolutionary multi-objective optimization algorithm; The Pearson correlation coefficient is used to measure the diversity, and the calculation formula of the Pearson correlation coefficient between any two base learners is as follows: (24); In the formula: and respectively represent the prediction errors of any two base learners, represents the covariance between any two errors, represents the variance of the calculated error; According to formula (25), calculate the correlation coefficient between any two selected base learners, and take the average value of the obtained correlation coefficients as the diversity index of the ensemble model; (25); In the formula: is The specific definition formula, that is, the specific expression of the diversity objective function of the individual learner in the evolutionary multi-objective optimization algorithm; represents the number of selected base learners; Convert the maximization multi-objective problem in formula (22) into the minimization optimization problem in formula (26): (26); Use the non-dominated sorting genetic algorithm to solve the Pareto front of the multi-objective problem, and select base learners to construct the best heterogeneous Stacking ensemble prediction model.
[0014] Optionally, the specific process of the non-dominated sorting genetic algorithm to achieve selective integration is as follows: Perform binary encoding on the heterogeneous base learners, randomly initialize the chromosomes of each individual, and generate an initial population N with a size of , and use it as the parent population; Perform binary tournament selection, crossover, and mutation operations on the parent population to generate an offspring population with the same size as , and fuse it with the parent population to obtain a population N with a size of 2 , that is ; For the population Decode the individuals in it to determine the selected base learners, evaluate the integrated prediction effect of the selected base learners based on the training set, and calculate two optimization objectives to obtain the fitness, and then perform fast non-dominated sorting and calculate the crowding degree; According to the non-dominated relationship and crowding degree of the population individuals, select a new parent population with a size of N ; ; Repeat the above operations until the number of iterations reaches the maximum number of evolutionary generations, stop the optimization, and obtain the Pareto optimal solution set; Decode the binary chromosome string to obtain the heterogeneous model library for selective integration.
[0015] The present invention also provides a heterogeneous integrated short-term multi-load forecasting system, which applies the heterogeneous integrated short-term multi-load forecasting method described in any one of the above, including a feature construction module, a multi-objective evolutionary optimization module, and an interpretability analysis module connected in sequence; The feature construction module is used to decompose the multi-load historical sequence into periodic sequences and trend sequences on multiple time scales based on the multi-seasonal trend decomposition, and construct differentiated input features; The multi-objective evolutionary optimization module is used to take diversity and accuracy as optimization objectives, and screen the optimal learner combinations suitable for different seasons and loads through the non-dominated sorting genetic optimization algorithm with an elite strategy to construct the most excellent heterogeneous Stacking integrated prediction model; The interpretability analysis module is used to perform attribution analysis through the Shapley value additivity interpretation method, and measure the contribution degree of each input feature to the most excellent heterogeneous Stacking integrated prediction model from both global and individual dimensions.
[0016] It can be seen from the above technical solutions that compared with the prior art, the present invention discloses a heterogeneous integrated short-term multi-load forecasting method and system, which has the following beneficial effects: (1) The present invention adopts the MSTL (multi-seasonal trend decomposition) method to decompose the multi-load into periodic sequences and trend sequences on multiple time scales, which can more accurately capture the load change characteristics under different seasons and cycles, and improve the prediction accuracy; (2) The present invention proposes a multi-objective evolutionary algorithm based on diversity and accuracy indicators. Through the non-dominated sorting genetic optimization algorithm with an elite strategy, the optimal learner combinations suitable for different seasons and loads are screened, and a heterogeneous Stacking integrated prediction model is constructed, which improves the prediction accuracy and stability; (3) The present invention adopts the SHAP (Shapley value additivity interpretation) method for attribution analysis, measures the contribution degree of each input feature to the prediction model from both global and individual dimensions, improves the interpretability of the prediction model, and facilitates decision-making and optimization in practical applications. Description of the Drawings
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.
[0018] Figure 1 It is a flowchart of the heterogeneous integration short-term multi-load forecasting method provided by the present invention; Figure 2a It is an autocorrelation coefficient diagram of the annual electricity load of IES with a 96h time delay provided by the present invention; Figure 2b It is an autocorrelation coefficient diagram of the annual cooling load of IES with a 96h time delay provided by the present invention; Figure 2c It is an autocorrelation coefficient diagram of the annual heating load of IES with a 96h time delay provided by the present invention; Figure 3 It is a schematic diagram of the Stacking ensemble learning framework provided by the present invention; Figure 4 It is a coding method diagram of the heterogeneous base model provided by the present invention; Figure 5a It is the MAPE index of the multi-load annual test set provided by the present invention; Figure 5b It is the R 2 index provided by the present invention; Figure 6 It is a comparison diagram of the MAPE of the electricity load prediction of Model 0~Model 4 provided by the present invention; Figure 7 It is a structural diagram of the heterogeneous integration short-term multi-load forecasting system provided by the present invention. Detailed implementation manners
[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0020] The current short-term multi-load forecasting methods have the following technical problems when dealing with the load of integrated energy systems: First, the complex coupling relationships of multi-loads are inadequately handled. Existing methods do not sufficiently consider the complex coupling characteristics among various loads such as electricity, cooling, and heating in integrated energy systems, resulting in a reduction in forecasting accuracy. Second, the ability to mine seasonal differentiation features is insufficient. The multi-loads in different seasons have significant periodic and trend changes, but existing methods fail to fully utilize these differentiation features, thus limiting the generalization performance of the model. Third, there is the problem of optimizing the combination of learners. Traditional forecasting methods usually adopt a single model or a fixed combination model, lacking a dynamic optimization mechanism based on diversity and accuracy, and it is difficult to select the optimal learner combination for different seasons and load types, leading to inaccurate forecasting results. Finally, the interpretability is insufficient. Existing load forecasting methods are mostly "black-box" models, lacking the analysis of the global and individual contributions of input features, which limits users' trust and understanding of the forecasting results.
[0021] Embodiment 1: To overcome the above technical problems, this embodiment provides a heterogeneous integrated short-term multi-load forecasting method, as Figure 1 shown, which specifically includes the following steps: S1: Based on the multiple seasonal trend decomposition, decompose the multi-load historical sequence into periodic sequences and trend sequences on multiple time scales, and construct differentiated input features; S2: With diversity and accuracy as the optimization objectives, screen the optimal learner combination applicable to different seasons and loads through the non-dominated sorting genetic optimization algorithm with an elite strategy, and construct the most excellent heterogeneous Stacking integrated forecasting model; S3: Adopt the Shapley value additive interpretation method for attribution analysis, and measure the contribution degrees of each input feature to the most excellent heterogeneous Stacking integrated forecasting model from both global and individual dimensions.
[0022] Aiming at the characteristics of IES multi-load forecasting, based on the Figure 1 shown process, this embodiment is based on the multiple seasonal trend decomposition (Multiple Seasonal Trend decomposition using Loess, MSTL), decomposes the load into periodic sequences and trend sequences on multiple time scales, thereby constructing differentiated input features, and fully mines the long-term change laws and differentiated periodicities of multi-loads in different seasons. Further, a heterogeneous Stacking integrated model based on multi-objective evolution is adopted, with diversity and accuracy as the optimization objectives, to find the optimal learner combination and improve the forecasting accuracy and stability. In addition, the SHAP interpretation method is introduced to conduct attribution analysis on the forecasting model at both the global and individual levels, greatly improving the transparency and interpretability of the model.
[0023] To fully explore the differential potential change laws and complex coupling characteristics of diversified loads in the integrated energy system in each season, the following will elaborate on Figure 1 the process shown below, providing a novel, practical, and efficient solution for the prediction of diversified loads in IES.
[0024] I. Construction of input features for diversified loads 1. Analysis of diversified load characteristics 1) Autocorrelation analysis The autocorrelation coefficients function (ACF) calculation formula in Equation (1) is used to analyze the diversified load, to test whether the historical sequence of the diversified load has short-term and long-term repetitive patterns, and to determine the importance of historical load for prediction; (1); In the formula: represents the time delay, represents the load autocorrelation coefficient with a time delay of , represents t the load value at time represents the mean value of historical loads, represents the load value at time represents the variance of historical loads; the value range of the autocorrelation coefficient is [-1, 1], the larger, the stronger the correlation; E represents taking the mean value.
[0025] Figure 2a , Figure 2b , Figure 2c Figure 46 shows the autocorrelation coefficient diagram of the annual diversified load of IES with a 96h time delay, where the shaded part is the 95% confidence interval. It can be seen that when the time delay is in the range of [0, 12], the autocorrelation coefficient of the diversified load decreases with the increase of the time delay, and when the time delay is small, the autocorrelation coefficient of the diversified load is large; when the time delay is in the range, the autocorrelation coefficient of the diversified load first decreases and then increases with the increase of the time delay, and the peak value of the autocorrelation coefficient diagram corresponding to the time delay gradually decreases with the increase of the number of days of the time. This confirms that the diversified load of IES at a certain moment depends not only on the load at the adjacent moment but also on the load at the same moment of the adjacent day. Therefore, when predicting the diversified load, the short-term repetitive pattern and the daily cycle periodicity of the diversified load should be fully considered.
[0026] 2) Analysis of the coupling characteristics of multiple loads in each season In the IES, each energy subsystem realizes the coordinated complementarity between multiple energies through energy conversion devices. The change in the energy consumption behavior pattern of IES users in different seasons will affect the energy utilization method between different energy forms in the IES, making the strength of the coupling relationship between the multiple loads in the IES change with the seasons. The differences in meteorological conditions in different seasons will affect the energy consumption habits of IES users, and the differences in user energy consumption habits will further affect the energy conversion and utilization strategies between various energy forms within the IES, making the load show seasonal, periodic and other phenomena. Therefore, analyzing the coupling characteristics of multiple loads in the IES in different seasons and the correlation between multiple loads and meteorological factors is of great significance for feature selection in multiple load forecasting.
[0027] Use the Spearman rank correlation analysis calculation formula with low computational complexity and wide application range in Equation (2) to analyze the historical data of multiple loads in different seasons, and obtain the coupling characteristics of multiple loads in different seasons and the correlation between multiple loads and meteorological factors; (2); In the formula: r represents the rank correlation coefficient between different loads, represents the rank difference between two data variables, n represents the total number of observed samples; r The value range of is [-1, 1], and the larger
[0028] 2. Construct differential cycle-trend multiple load features based on MSTL To further explore the differential potential change laws between multiple loads in different seasons, the MSTL decomposition method is used to decompose the historical sequence of multiple loads, and the obtained load components are used to construct the periodic features and trend features of each load on multiple time scales.
[0029] 1) Multiple seasonal trend decomposition MSTL can be understood as additively decomposing a time series into a trend component, a residual component and multiple seasonal components through multiple iterations of the STL (Seasonal-Trend decomposition using Loess) algorithm, and its expression is shown in Equation (3): (3); In the formula: represents the historical sequence of multiple loads, represents the trend component of the load sequence, represents the residual component of the load sequence, represents the pcorresponding to a time scale period p seasonal components; The decomposition of multiple seasonal trends is divided into an inner loop and an outer loop. The inner loop is nested within the outer loop. The specific decomposition process is as follows: ① Parameter initialization: Set the original sequence as , the number of time scale periods existing in the sequence is p , and the period array q is obtained by sorting in ascending order of time scale; At the same time, the initial value of all seasonal components is set to 0, the initial value of the non-seasonal component d is x , and the number of outer loop iterations is l ; ② Iterative decomposition: Take the -th outer loop and the -th inner loop as an example. First, update the non-seasonal component , where is the seasonal component of the k -1-th outer loop decomposition for the i -th time scale period. When k = 1, is the initial value 0 of the seasonal component; Secondly, use the STL algorithm to decompose the non-seasonal component to obtain the seasonal component of the -th time scale decomposed in the -th outer loop, and update the non-seasonal component ; When the traversal of all q period elements in the period array p is completed, the inner loop terminates and a new round of outer loop starts until the l -th outer loop, the outer loop ends, and p multiple seasonal components are obtained; represents the non-seasonal component after the first update, represents the non-seasonal component after the second update; ③ Calculate the trend component and the residual component based on the p multiple seasonal components and the original sequence.
[0030] Among the different load components generated by MSTL decomposition, the multiple seasonal component represents the periodic repetition patterns of the multivariate load at each time scale, and the trend component represents the long-term change law of the multivariate load, and both show differences among different loads in different seasons. The residual component reflects the random changes and irregular fluctuations of the multivariate load sequence caused by data anomalies and emergencies, and cannot be used to analyze the potential change law of the multivariate load. Therefore, in this embodiment, the multiple seasonal component obtained by MSTL decomposition is used as the periodic feature of the multivariate load at each time scale, and the trend component is used as the long-term change trend feature of the multivariate load.
[0031] 2) Construct the differential input features of the multivariate load for each season To fully explore the differential periodicity and differential long-term change trend characteristics between different seasons and different loads, this embodiment uses Spearman rank correlation analysis to screen the strongly correlated sets of the multivariate load with the trend component and the multiple seasonal component in different seasons, and constructs the differential input features for each season.
[0032] II. Multivariate load forecasting using the Stacking ensemble method based on multi-objective optimization 1. Ensemble learning method based on Stacking In the heterogeneous Stacking ensemble prediction model, the individual learners for the preliminary prediction are the base learners, and the learner that combines the results of the preliminary prediction individual learners and performs the secondary prediction to output the final prediction result is the meta-learner; the ensemble learning method based on Stacking is specifically as follows: ① Randomly divide the dataset into k subsets of equal size and non-overlapping, denoted as , , and define and as the K th test set and training set in the k -fold cross-validation; the first-layer prediction model contains K base learners, and the training set is used to train the base model k using algorithms; among them, is the feature vector of the n th sample, is the predicted value of the n th sample, m is the number of features included, and each feature vector can be expressed as ; ② For each sample K in the k th test set in the , the prediction results of the base learners are denoted as ; after the cross - validation process ends, a new data set is formed based on the output results of K base learners, denoted as ; ③ Based on the newly formed data set , train the second - layer meta - learner model to obtain the optimal prediction result. Using K k - fold cross - validation can reduce the risk of overfitting while ensuring that the meta - learner has sufficient training samples. In this embodiment, the 5 - fold cross - validation method is adopted, and the Stacking ensemble learning framework is as shown in Figure 3 .
[0033] 2. Construct a heterogeneous base model library One of the keys to constructing a high - performance integrated multi - load prediction model is to construct diverse base learners. The heterogeneous Stacking ensemble prediction model in this embodiment selects 10 base learners, namely Catboost, DT, KNN, Lasso, RF, Ridge, SVM, XGBoost, LSTM, and GRU, to fully utilize the performance advantages of different algorithms and construct a diverse heterogeneous base model library.
[0034] DT, RF, Catboost, and XGBoost are tree - structured models. They use a recursive method to divide the data set into different regions by splitting features, and finally make predictions based on the target variables in each region.
[0035] Among them, DT selects the optimal feature for splitting through the Gini coefficient (CART), and the goal is to minimize the impurity after each split. The calculation formula of CART is shown in Equation (4): (4); In the formula: represents the probability that the sample belongs to class i .
[0036] RF obtains more stable and accurate results by constructing multiple decision trees and voting on the prediction results of each tree. The calculation process is shown in Equation (5): (5); In the formula: T represents the number of DTs, represents the t th prediction value of the DT.
[0037] Catboost and XGBoost are variants of Gradient Boosting Decision Tree (GBDT), which improve the model performance by gradually fitting the residuals. XGBoost introduces a regularization term to prevent overfitting, while Catboost adopts special processing for categorical features, making it more efficient in dealing with categorical features.
[0038] Among them, the objective function optimized by XGBoost is shown in Equation (6): (6); In the formula, L is the loss function; is the regularization term, which is used to describe the complexity of the penalty model.
[0039] Lasso is a variant of linear regression. Through L1 regularization, the coefficients of some features are set to zero for feature selection. Its loss function is shown in Equation (7): (7); In the formula, is the penalty coefficient, p represents the number of features, represents the model parameter (coefficient) vector in the j th element.
[0040] Ridge is a variant of linear regression. It uses L2 regularization to penalize large coefficients and prevent overfitting. Its main advantage is that it will not make the coefficients completely zero. Its loss function is shown in Equation (8): (8); KNN belongs to the distance metric model and is an instance-based learning method. When making a prediction, it selects the K nearest samples and determines the predicted value according to the majority voting result of their labels. The specific process is shown in Equations (9)-(10): (9); (10); In the formula, y is the category of the predicted data point x ; is the distance from x to the nearest neighbor data point i ; I is the indicator function, which is equal to 1 when and 0 otherwise; calculate the average value of the labels of K data points, and the label corresponding to this average value is the prediction result. Among them, zRepresents the target value of the data point to be predicted x , represents the i th nearest neighbor data point for x the target value.
[0041] SVM distinguishes samples of different classes by finding an optimal hyperplane in a high-dimensional space. The goal of SVM is to maximize the margin between classes, finding a decision boundary that maximizes the distance between classes, as shown in Equation (11): (11); In the formula, is the normal vector of the hyperplane; b is the bias term; is the label; is the feature vector.
[0042] LSTM and GRU are variants of a special Recurrent Neural Network (RNN), designed to address the problem of vanishing or exploding gradients that occur in traditional RNNs during long sequence learning. By introducing a gating mechanism to control the flow of information, long-term dependencies can be remembered.
[0043] The LSTM cell calculations are shown in Equations (12)-(17): (12); (13); (14); (15); (16); (17); In the formula, is the Sigmoid activation function; is the input to this layer at time t ; and are weight matrices; is the bias. represents the input gate output, represents the weight matrix of the input gate for the current input hidden state, represents the weight matrix of the input gate for the previous hidden state, represents the previous hidden state, represents the input gate bias term; represents the forget gate output, represents the candidate memory cell, Denote the weight matrix of candidate memory units, Denote the weight matrix of candidate memory units, Denote the bias term of candidate memory units, Denote the state of the memory unit at the current time, Denote the state of the memory unit at the previous time, Denote the output gate, Denote the weight matrix of the output gate, Denote the weight matrix of the output gate, Denote the bias term of the output gate, Denote the hidden state (i.e., output) at the current time.
[0044] The GRU cell is calculated as shown in equations (18)-(21): (18); (19); (20); (21); In the formula, and are the weight matrices of the reset gate; is the bias. Denote the value of the reset gate at the current time, Denote the input at the current time, Denote the hidden state at the previous time, Denote the value of the update gate at the current time, Denote related to the input Related weight matrix, Denote the weight matrix related to the hidden state at the previous time, Denote the candidate hidden state at the current time, Denote the weight matrix related to the current input, Denote the weight matrix related to the hidden state and the reset gate output at the previous time, Denote the hidden state at the current time.
[0045] 3. Selective Ensemble Based on Evolutionary Multi-Objective Optimization In ensemble learning, accuracy and diversity are two conflicting objectives. If all the constructed heterogeneous models are directly used for ensemble, it will not only result in a high complexity and high computational cost of the ensemble model, but also lead to problems such as reduced ensemble prediction performance and low prediction accuracy. Therefore, in order to generate "good and different" base learners in ensemble learning, selective ensemble learning is adopted to remove redundant base learners and select diverse and high-performance base learners from the heterogeneous base model library, thus saving the computational consumption of the ensemble learning model.
[0046] In this embodiment, an evolutionary multi-objective optimization-based selective integration method is used to screen the base learners, and the optimization problem is shown as follows: (22); In the formula: and are the objective functions for measuring the accuracy and diversity of individual learners respectively; the key to solving the multi-objective optimization problem is to clarify the decision variables, constraint conditions, and objective functions to be optimized.
[0047] All the base learners in the constructed heterogeneous library are binary encoded. Each bit of the encoding indicates whether the corresponding base learner is selected. 1 means the learner is selected, and 0 means it is not selected. This set of binary variables is used as the decision variables. As shown in Figure 4 , it is the encoding diagram of the heterogeneous base model. In this diagram, 1 to 10 represent the learners Catboost, DT, KNN, Lasso, RF, Ridge, SVM, XGBoost, LSTM, and GRU in sequence. And to ensure that the scale of the integrated model is within an appropriate range, the number of base learners selected is used as the constraint condition of the optimization problem.
[0048] For the definition of the objective function in the optimization problem, in this embodiment, the coefficient of determination (R-Square, R 2 ) is selected as the prediction accuracy index. The larger it is, the higher the prediction accuracy of the base learner. The definition formula of the prediction accuracy objective function is: (23); In the formula: n represents the number of samples in the training set. represents the predicted output of the selected base learners integrated for the i th training sample in the training set. represents the i th actual observed value in the training set. represents the average value of all actual observed values in the training set; is the specific definition formula of , that is, the specific expression of the individual learner accuracy objective function in the evolutionary multi-objective optimization algorithm; For the diversity index, in this embodiment, the Pearson correlation coefficient is used to measure the diversity. The greater the difference between two base learners, the smaller the error correlation coefficient of the predicted outputs. The Pearson correlation coefficient calculation formula for any two base learners is as follows: (24); In the formula: , respectively represent the prediction errors of any two base learners, represents the covariance between any two errors, represents the variance of the calculated error; Calculate the correlation coefficient between any two selected base learners according to Equation (25), and take the average of the obtained correlation coefficients as the diversity index of the ensemble model; (25); In the formula: is the specific definition formula of, that is, the specific expression of the diversity objective function of individual learners in the evolutionary multi-objective optimization algorithm; represents the number of selected base learners; In summary, the maximization multi-objective problem in Equation (22) is converted into the minimization optimization problem in Equation (26): (26); Adopt the Non-dominated sorting genetic algorithmⅡ (NSGA-Ⅱ) to solve the Pareto front of the multi-objective problem, and select base learners to construct the best quality Stacking ensemble prediction model.
[0049] Specifically, the specific process of realizing selective integration based on the non-dominated sorting genetic algorithm is as follows: ① Perform binary encoding on heterogeneous base learners, randomly initialize the chromosomes of each individual, and generate an initial population N with a scale of , and use it as the parent population; ② Perform binary tournament selection, crossover and mutation operations on the parent population to generate an offspring population with the same size as , and fuse it with the parent population to obtain a population N with a scale of 2 , that is, ; ③ Decode the individuals in the population to determine the selected base learners, evaluate the ensemble prediction effect of the selected base learners based on the training set, and calculate two optimization objectives to obtain the fitness, and perform fast non-dominated sorting and calculate the crowding degree accordingly; Select a new parent population N with a scale of according to the non-dominated relationship and crowding degree of the population individuals; ④ Repeat the above operations until the number of iterations reaches the maximum number of generations of evolution, and stop the optimization.
[0050] Optimize the above operations by setting appropriate population size and number of iterations to obtain the Pareto optimal solution set, where any Pareto solution corresponds to a binary variable combination for the selection of base learners; subsequently, decode the binary chromosome string to obtain a heterogeneous model library for selective integration.
[0051] Through the above selective optimization and integration operations, finally select base learners for the final integration. The results of the selective integration learning algorithm for multi-objective optimization are shown in Table 1. It can be seen from the data analysis that when the combination of base learners is 1, 5, 6, 7, 9 and the meta-learner is 1, at this time, the prediction effect of the spring electrical load is the best and the model diversity is the optimal, and the prediction accuracy and diversity values are (0.9734, 0.3523) respectively.
[0052] Table 1 Prediction results of electrical load for different combinations of learners
[0053] Obtain the predicted values in the test set samples according to the selected heterogeneous model library . Selecting a reasonable integration strategy is an important means to improve the prediction effect. Generally speaking, weighted integration is a relatively effective fusion strategy. In this embodiment, the final prediction result is obtained by using the average weighted method, and the calculation process is shown in Equation (27): . (27).
[0054] Embodiment 2: This embodiment uses the IES electrical, cooling, and heating load data from 00:00 on March 1, 2020 to 23:00 on February 28, 2021 at the Tempe campus of Arizona State University in the United States. The data sampling interval is 1h. The meteorological data comes from the National Solar Radiation Database of the United States. The federal legal holidays in the United States are used as the selection rule for holidays in the dataset, and at the same time, three calendar rules of time, day, and month are considered as input features.
[0055] To fully consider the seasonal load changes, the multivariate load data is divided according to the four seasons, and the data for each quarter is divided into a training set and a test set according to a ratio of 7:3, and then the training set is further divided by 5-fold cross-validation. This embodiment uses the data for rolling prediction with a prediction step of 1h.
[0056] This embodiment uses The criterion filters out abnormal data in the historical load, treats the abnormal values as missing values, and uses cubic spline interpolation to fill in the missing values. To eliminate the influence of different dimensions of input features on the prediction results, Min-Max normalization is adopted to make the input data within the range of [0,1]. The calculation process is shown in Equation (28): (28); In the formula, x is a certain sequence in the input data; and are the maximum and minimum values of this sequence, respectively.
[0057] (1) Comparison with the prediction results of a single prediction model To verify the effectiveness of the heterogeneous Stacking ensemble prediction model based on multi-objective evolution proposed in this patent, the used Stacking ensemble model is compared with each learner in it, namely Catboost, DT, KNN, Lasso, RF, Ridge, SVM, XGBoost, LSTM, and GRU. Figure 5a 、 Figure 5b show the evaluation indexes of the annual test set of the multi-load.
[0058] The Stacking ensemble framework combines diverse base learners, fully leveraging the advantages of each algorithm to observe data from different data spaces and structures, and makes up for the limitations of different base learners in dealing with specific data. Then, the meta-learner performs average weighting on the prediction results of the base learners to further improve the prediction accuracy. Taking the cooling load as an example, the of this embodiment has a 1.12% improvement compared to Ridge with the highest prediction accuracy among single models, and the MAPE and RMSE are reduced by 72.60% and 62.18% respectively. It can be seen from Figure 5a that compared with the single-learner prediction method, the MAPE indexes of the proposed method for predicting electric, cooling, and heating loads all decrease, and the decrease in MAPE for predicting electric and cooling loads is the most obvious. In Figure 5b , the indexes of the proposed method for multi-load prediction are all improved compared with those of the single-learner prediction, and the indexes for predicting the three types of loads are all close to 1. indexes for predicting the three types of loads are all close to 1.
[0059] (2) Comparison with the results of different feature set construction methods To verify the effectiveness of the proposed method for constructing the multi-source load input feature set, comparative experiments Model 0 to Model 4 are designed. Among them, the input feature types adopted by Model 0 and Model 1 are the commonly used input features in existing research. Model 0 includes the historical load features, time features, and holiday features of each load. Model 1 adds the strongly correlated meteorological features of each load in different seasons compared to Model 0. Model 2 adds the periodic features of each load at different time scales in different seasons compared to Model 1. Model 3 adds the trend features of each load in different seasons compared to Model 1. Model 4 adds the periodic features and trend features of each load at different time scales in different seasons compared to Model 1. Figure 6 It is a comparison chart of the MAPE indicators of the electric load prediction models for Model 1 to Model 4.
[0060] From Figure 6 it can be seen that compared with Model 1, the decline range of the MAPE of the electric load in different seasons for Model 2 is 13.160% - 40.511%. Compared with Model 1, the decline range of the MAPE of the electric load in different seasons for Model 3 is 29.843% - 38.668%. Compared with Model 1, the decline range of the MAPE of the electric load in different seasons for Model 4 is 48.276% - 58.637%. Compared with Model 0, the decline range of the MAPE of the electric load in different seasons for Model 4 is 49.200% - 59.841%. In summary, the proposed strategy for constructing the differential cycle-trend multi-source load features based on MSTL in different seasons fully exploits the differential cycle-trend characteristics among different seasons and different loads of the multi-source load, effectively improving the prediction accuracy of the multi-source load.
[0061] (3)Comparative analysis with different base learners To further verify the effectiveness of the multi-objective selection of the optimal learner combination, Table 1 presents the electric load prediction results for different quarters and different learner combination methods. Among them, Stacking Model 1 uses the random selection of learner combination method, and Stacking Model 2 uses the multi-objective optimization to select the optimal learner combination method.
[0062] As can be seen from Table 1, different combinations of learners have a great impact on the prediction results of the Stacking ensemble model. In autumn, the combination of DT-KNN-RF-Ridge-LSTM selected in this paper has a 32.45% improvement compared with the random combination method. In summer, although the random combination of Catboost-DT-Lasso has a certain advantage in prediction time compared with the combination selected in this paper due to the reduction in the number of learners, its prediction accuracy also drops significantly. It is reduced by 15.31% compared with the method in this paper. Thus, it proves the rationality of using multi-objective optimization to select learners with higher prediction accuracy and better diversity for combination. When predicting the cooling and heating loads in the remaining quarters, the use of multi-objective optimization to select the optimal learner combination is more superior in prediction performance than the random selection method. It is proved again that the diversity and difference of the learner combination will make the integrated result more robust and accurate, thus greatly improving the prediction effect.
[0063] Embodiment 3: corresponding to Figure 1 the method described above, an embodiment of the present invention provides a heterogeneous integrated short-term multi-load prediction system for Figure 1 For the specific implementation of the method in, a heterogeneous integrated short-term multi-load prediction system provided by an embodiment of the present invention can be applied to a computer terminal or various mobile devices, such as Figure 7 as shown, specifically including a feature construction module, a multi-objective evolutionary optimization module, and an interpretability analysis module connected in sequence; The feature construction module is used to decompose the multi-load historical sequence into periodic sequences and trend sequences of multiple time scales based on the multi-seasonal trend decomposition, and construct differentiated input features; The multi-objective evolutionary optimization module is used to take diversity and accuracy as optimization objectives, and screen the optimal learner combination suitable for different seasons and loads through the non-dominated sorting genetic optimization algorithm with elitist strategy, and construct the most excellent heterogeneous Stacking integrated prediction model; The interpretability analysis module is used to conduct attribution analysis through the Shapley value additivity interpretation method, and measure the contribution degree of each input feature to the most excellent heterogeneous Stacking integrated prediction model from both global and individual dimensions.
[0064] To sum up, in order to improve the prediction accuracy of IES multi-load, deeply explore the complex coupling and potential change laws among IES multi-loads, the present invention proposes a heterogeneous integrated short-term multi-load prediction method and system considering differentiated cycle-trend characteristics, and the remarkable effects can be summarized as follows: (1) Compared with the single-load prediction method, the multi-load prediction method proposed by the present invention effectively overcomes the problem that the generalization ability of a single model has limitations, and the MAPE indicators of the electric, cooling, and heating loads in the annual test set have all decreased, verifying the rationality and effectiveness of the method in this paper.
[0065] (2) In the comparison of different feature set construction methods, the present invention proposes to use the multi-load decomposition method based on MSTL to construct differentiated input features in different seasons, which can effectively explore the differentiated potential change laws and complex coupling characteristics among multi-loads in different seasons, and its prediction error is lower and the accuracy is higher than that of the feature construction method.
[0066] (3) Compared with other learner combination methods, the selective Stacking integration method based on evolutionary multi-objective optimization proposed in the present invention can effectively calculate the optimal learner combination, improving the efficiency and prediction performance of the overall integrated model.
[0067] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and reference can be made to the description of the method part for relevant parts.
[0068] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A heterogeneous integrated short-term multivariate load forecasting method, characterized in that: The following steps are involved: S1: Based on multi-season trend decomposition, the multivariate load historical series is decomposed into multi-time scale period series and trend series to construct differentiated input features; S2: Taking diversity and accuracy as optimization goals, the optimal learner combination suitable for different seasons and loads is selected through the non-dominated sorting genetic optimization algorithm with elite strategy to build the best quality Stacking integrated prediction model; S3: The Shapley value additive interpretation method is used for attribution analysis to measure the contribution of each input feature to the most optimal Stacking ensemble prediction model from both global and individual dimensions.
2. A heterogeneous integrated short-term multivariate load forecasting method according to claim 1, characterized in that: Also includes: Before constructing the differentiated input features, a multivariate load characteristic analysis is performed, which includes the following steps: The autocorrelation coefficient calculation formula of formula (1) is used to analyze the multivariate loads to test whether the multivariate load historical series has short-term and long-term repetitive patterns; (1); Where: Indicates time delay, Indicates the time delay is The load autocorrelation coefficient, express t The load value at the moment, represents the historical load mean, express The load value at the moment, Represents the historical load variance; the autocorrelation coefficient ranges from [-1,1]; E It means to find the mean value; The Spearman rank correlation analysis formula (2) is used to analyze the historical data of multivariate loads in different seasons to obtain the coupling characteristics of multivariate loads in different seasons and the correlation between multivariate loads and meteorological factors. (2); Where: r represents the rank correlation coefficient between different loads, represents the rank difference between two data variables, n represents the total number of observed samples; r The value range is [-1,1].
3. The heterogeneous integrated short-term multivariate load forecasting method according to claim 1 is characterized in that: In S1, constructing differentiated input features includes the following steps: Using the multi-seasonal trend decomposition algorithm of formula (3), the multivariate load history series is additively decomposed into trend component, residual component and multi-seasonal component; (3); Where: represents the multivariate load history series, represents the trend component of the load series, represents the residual component of the load series, The load sequence p The time scale period corresponds to p Seasonal component; Spearman rank correlation analysis was used to screen the strongly correlated sets of multivariate loads and trend components and multiple seasonal components in different seasons, and to construct differentiated input characteristics for each season.
4. The heterogeneous integrated short-term multivariate load forecasting method according to claim 1 is characterized in that: In S1, the multiple seasonal trend decomposition is divided into an inner loop and an outer loop, and the inner loop is nested in the outer loop. The specific decomposition process is: Set the original sequence to , the number of periods of each time scale of the sequence is p , sort the periodic array according to the time scale from small to large q ; At the same time, all seasonal components The initial value of is set to 0, and the non-seasonal component d The initial value is x , the number of outer loops is l ; In progress k The outer cycle i In the second inner loop, first, update the non-seasonal component ,in, For the k -1 times the outer loop decomposes the i The seasonal component of the time scale cycle is k =1, The initial value of the seasonal component is 0; secondly, the STL algorithm is used to decompose the non-seasonal component , get the k The first i Seasonal component of time scale , update the non-seasonal component ; When the cycle array is completed q All p After traversing the elements of the cycle, the inner loop execution terminates and a new outer loop starts again until the first l When the outer cycle is completed, the decomposition is p Multiple seasonal components are used as the periodic characteristics of each time scale of the multivariate load; among them, , , represents the non-seasonal component after the first update, represents the non-seasonal component after the second update; according to p The trend component and residual component are calculated from the multiple seasonal components and the original sequence, and the trend component is used as the long-term trend characteristic of the multivariate load.
5. The heterogeneous integrated short-term multivariate load forecasting method according to claim 1 is characterized in that: In S2, the individual learner that makes the initial prediction in the heterogeneous Stacking ensemble prediction model is the base learner, and the learner that combines the initial prediction results of the individual learner and performs secondary prediction to output the final prediction result is the meta learner; the specific ensemble learning method based on Stacking is: Randomly divide the data set Divide into k There are equal-sized and non-overlapping subsets, denoted by , , respectively define and for K In the cross validation k The first layer prediction model contains K base learner, for the training set use k The algorithm is trained to obtain the base model ;in, For the n The feature vector of the samples, For the n The predicted value of samples, m is the number of features contained, and each feature vector can be expressed as ; for K The first fold in cross validation k Fold test set Each sample in , base learner The prediction result is recorded as ; After the cross-validation process is completed, a new data set is formed based on the output results of K base learners, denoted as ; Based on the newly constructed dataset , train the second-layer meta-learner model to obtain the optimal prediction result.
6. A heterogeneous integrated short-term multivariate load forecasting method according to claim 1, characterized in that: In S2, the heterogeneous Stacking ensemble prediction models select Catboost, DT, KNN, Lasso, RF, Ridge, SVM, XGBoost, LSTM and GRU to build a diverse heterogeneous base model library.
7. A heterogeneous integrated short-term multivariate load forecasting method according to claim 5, characterized in that: In S2, the best quality Stacking ensemble prediction model is constructed, which includes the following steps: The selective ensemble method of evolutionary multi-objective optimization is used to screen the base learners. The optimization problem is as follows: (22); Where: and are the objective functions for measuring the accuracy and diversity of individual learners, respectively; Prediction accuracy objective function The definition of is: (23); Where: n represents the number of samples in the training set, Indicates the selected After the base learners are integrated, the i The predicted output of training samples is Indicates the training set i The actual observed values, represents the average value of all actual observations in the training set; for The specific definition of is the specific expression of the accuracy objective function of the individual learner in the evolutionary multi-objective optimization algorithm; The Pearson correlation coefficient is used to measure diversity. The Pearson correlation coefficient calculation formula for any two base learners is as follows: (24); Where: , denote the prediction errors of any two base learners, represents the covariance between any two errors, represents the variance of the calculation error; According to formula (25), the correlation coefficient between any two selected base learners is calculated, and the average of the obtained correlation coefficients is taken as the diversity index of the integrated model; (25); Where: for The specific definition of is the specific expression of the diversity objective function of the individual learner in the evolutionary multi-objective optimization algorithm; Indicates the number of selected base learners; The maximization multi-objective problem of formula (22) is transformed into the minimization optimization problem of formula (26): (26); The non-dominated sorting genetic optimization algorithm is used to solve the Pareto frontier of multi-objective problems and select The base learners are used to build the best quality Stacking ensemble prediction model.
8. The heterogeneous integrated short-term multivariate load forecasting method according to claim 5 is characterized in that: The specific process of non-dominated sorting genetic optimization algorithm to achieve selective integration is: Binary encode the heterogeneous base learner and randomly initialize the chromosome of each individual to generate a scale of N The initial population , and serve as the parent population; For parent population Perform binary tournament selection, crossover, and mutation operations to generate a Own population of the same size , and with the parent population The fusion scale is 2 N Population ,Right now ; For populations Decode the individuals in the population, determine the selected base learner, evaluate the ensemble prediction effect of the selected base learner based on the training set, and calculate the fitness of the two optimization objectives, so as to perform fast non-dominated sorting and calculate the crowding degree; according to the non-dominated relationship and crowding degree of the population individuals, select the scale of N The new parent population ; Repeat the above operation until the number of iterations reaches the maximum evolutionary generation, stop the optimization, and obtain the Pareto optimal solution set; decode the binary chromosome string to obtain a selectively integrated heterogeneous model library.
9. A heterogeneous integrated short-term multi-element load forecasting system, using a heterogeneous integrated short-term multi-element load forecasting method as claimed in any one of claims 1 to 8, characterized in that: It includes a feature construction module, a multi-objective evolutionary optimization module and an interpretability analysis module which are connected in sequence; The feature construction module is used to decompose the multivariate load history series into multi-time scale period series and trend series based on multi-season trend decomposition, and construct differentiated input features; The multi-objective evolutionary optimization module is used to select the optimal learner combination suitable for different seasons and loads through the non-dominated sorting genetic optimization algorithm with elite strategy, with diversity and accuracy as optimization goals, to build the best quality Stacking integrated prediction model; The interpretability analysis module is used to perform attribution analysis through the Shapley value additive interpretation method, and to measure the contribution of each input feature to the most excellent quality Stacking ensemble prediction model from both global and individual dimensions.
Citation Information
Patent Citations
Comprehensive energy system load prediction method considering multivariate load coupling characteristics
CN111950793A
Medium and long term power load prediction method based on MSTL and LSTM models
CN114595861A
Short-term generating capacity prediction method based on ensemble learning
CN116307119A
Disease auxiliary prediction system based on S-NStackingV balance optimization integrated framework
CN117198508A
Multivariable time series prediction model construction method based on man-machine collaborative optimization
CN118313487A
Cited By
High-speed rail seat comfort automatic adaptation control method based on passenger behavior habit machine learning
CN120986478A
Layered analysis and interpretable modeling method, system and equipment for multi-source driving factors of power system and medium
CN121980275A
Hierarchical analysis and interpretable modeling method, system and device for multi-source driving factors of power system and medium
CN121980275B