Multi-element load joint prediction method fusing decomposition and reconstruction technology and multi-task learning
By combining decomposition and reconstruction technology and multi-task learning method, multi-energy loads in integrated energy systems are predicted, and the problem of insufficient load prediction accuracy and stability in the prior art is solved, thereby achieving higher precision load prediction.
Patent Information
- Application Number
- CN202510432615.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-04-08
AI Technical Summary
Existing load prediction methods are difficult to effectively capture dynamic coupling information in multi-energy loads in integrated energy systems, resulting in insufficient prediction accuracy and stability.
The combined multi-load prediction method of fusion decomposition and reconstruction technology and multi-task learning is adopted to decompose and reconstruct the load data through variational modal decomposition and mixed characteristic evaluation method, and combined with the CNN-ECA-MMoE multi-task learning model, the model's dynamic coupled feature capture capability of multi-load is improved.
It improves the accuracy and stability of load prediction, reduces interference caused by task differences in multi-task learning, and enhances the model's ability to recognize important patterns.
Smart Images

Figure CN119965861A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of comprehensive energy load forecasting, and specifically relates to a multi-load joint forecasting method integrating decomposition and reconstruction technology with multi-task learning. Background Art
[0002] Each subsystem of the traditional energy system is relatively independent and closed, lacking a coordination mechanism. This decentralized structure limits the efficient linkage between different energy forms, resulting in low energy utilization efficiency and difficulty in flexibly responding to changes in modern energy demand. With the deep integration of multiple energy systems such as electricity, gas, and heat, an integrated energy system (IES) with the power system as the core has been formed. This integrated energy system combines multiple energy sources through a systematic and integrated approach, optimizes energy production, transmission, storage, and use, and promotes the development of renewable energy, laying a physical foundation for the construction of the energy Internet. Accurate load forecasting is crucial to optimizing energy configuration and ensuring the smooth operation of the system. Load forecasting not only helps the energy sector to formulate efficient supply and demand plans, but also ensures the rational allocation and use of multiple energy forms. It is an important basis for maintaining the stability of the energy system and improving overall energy efficiency.
[0003] With the rapid development of computer hardware and artificial intelligence technology, machine learning, especially deep learning, has shown significant advantages in load forecasting. Deep learning models can automatically extract features from large amounts of data, capture complex nonlinear relationships, and adapt to dynamic load changes. For example, the long short-term memory network LSTM can effectively capture time dependencies and long-term trends; the convolutional neural network CNN is good at extracting local features and spatiotemporal patterns; the Transformer is good at global dependency modeling, can efficiently capture long-distance relationships, and process sequence data in parallel. These technologies can extract features from a large amount of historical data, capture complex nonlinear relationships, and adapt to dynamically changing load patterns without the need to establish complex mathematical models. However, in integrated energy systems with high coupling and complex dynamic characteristics, these methods often ignore the dynamic information hidden in multi-energy loads, cannot fully capture the potential complex patterns in load data, and have insufficient ability to identify the fine structure retained after filtering out external disturbances inside the load and the relationship between load and noise. They show obvious limitations in accurately identifying and separating the inherent characteristics of the load itself, and it is difficult to achieve ideal prediction accuracy and stability. Therefore, in recent years, many scholars have devoted themselves to studying how to effectively capture the dynamic coupling information hidden in multi-energy loads, and explored the use of feature processing technology to separate effective information to improve the prediction accuracy of the model.
[0004] In the research of feature processing technology in recent years, load decomposition has become one of the widely used technologies. By decomposing complex load time series into multiple modal components, it can effectively remove noise in the data and separate uncertainty factors, thereby improving the stability and processability of load data. However, the above decomposition-prediction method still faces challenges, because the large number of modal components after decomposition will significantly increase the computational complexity of the model, resulting in decreased computational efficiency and error accumulation. In order to solve this problem, some researchers have tried to incorporate the decomposition-reconstruction method into the prediction process. The decomposition-reconstruction-prediction strategy effectively reduces the computational complexity of the model while retaining the decomposition benefits, and enhances the prediction model's ability to identify important patterns. However, the current reconstruction methods often rely on a single feature as a reference, and it is difficult to fully consider the multi-dimensional information of the modal components from multiple perspectives, resulting in a lack of sufficient comprehensiveness in the reconstructed data, making it difficult to fully realize the potential of load decomposition technology.
[0005] In order to improve the model's ability to capture the dynamic coupling characteristics of multi-energy loads, some researchers have tried to apply multi-task learning (MTL) to multi-energy load forecasting. Multi-task learning is a typical inductive transfer mechanism. By training multiple related tasks at the same time, it can make full use of the shared information between tasks, thereby improving the model's ability to understand complex systems. At present, most studies use the hard sharing mechanism in multi-task learning. This method fuses the independent information of subtasks through the shared layer and abstracts the global features for representation, thereby improving the model's generalization ability for each task. However, the hard sharing mechanism allows all subtasks to share the underlying feature extraction layer, which causes the subtasks to interfere with each other due to differences, making it difficult to extract specific information exclusive to each subtask, and lacks the flexibility to adjust the requirements of each task.
[0006] In summary, based on the inherent coupling characteristics of multivariate load sequences, using efficient feature processing technology to improve data representation capabilities and combining advanced deep learning networks and multi-task learning strategies are the key to achieving comprehensive energy load joint prediction. However, the existing decomposition and reconstruction technology as a feature processing method does not integrate modal characteristics from multiple angles, and still has deficiencies in the diversity of feature expression and task correlation mining. At the same time, the existing multi-task learning method based on parameter hard sharing mechanism is easily disturbed by the weak correlation of sub-tasks, which further limits the learning effect of the task. The patent application with publication number CN113822481A discloses a comprehensive energy load forecasting method based on multi-task learning strategy and deep learning. It only adopts the structure of the existing multi-task learning model MMoE, and does not improve the model. The model's feature extraction and recognition capabilities are limited; at the same time, it does not process the input data, which is easily affected by data noise, resulting in insufficient prediction accuracy.
[0007] Therefore, how to avoid the interference of weak related information on subtasks and achieve better prediction results is a technical problem that technicians in this field urgently need to solve. Summary of the invention
[0008] Taking into account the deficiencies of existing research, this paper proposes a multi-load joint forecasting method that integrates decomposition and reconstruction technology with multi-task learning. This method combines decomposition and reconstruction technology with multi-task learning strategy, while eliminating redundant information and reducing noise interference on multi-task learning, it can provide clearer and more accurate feature representation for different subtasks, thereby achieving more accurate multi-load forecasting.
[0009] To achieve the above purpose, the technical solution adopted by the present invention is as follows.
[0010] The multi-load joint forecasting method integrating decomposition and reconstruction technology with multi-task learning includes the following steps: Step S1, obtaining the historical load sequence of electric load, cooling load and heating load in the integrated energy system, forming a historical multi-load sequence and forming a historical multi-load sequence library; at the same time, obtaining the data sequence of the characteristics of the influencing factors affecting the electric load, cooling load and heating load, forming an influencing factor characteristic library; Step S2, preliminarily decomposing the historical load sequences of electric load, cooling load and heating load in the historical multivariate load sequence by using the variational mode decomposition method, and obtaining K mode components for each load; Step S3, reconstructing the K modal components of each load obtained in step S2 from multiple angles by using a hybrid characteristic evaluation method, and converting the K modal components corresponding to each load into a set of three types of aggregate sequences; then updating the three types of aggregate sequences corresponding to each load into a historical multivariate load sequence library to obtain an updated historical multivariate load sequence library; Step S4, normalizing the influencing factor feature library obtained in step S1 and the updated historical multivariate load sequence library obtained in step S3 with a time step length M as the sequence length, and then merging them to obtain multiple groups of three-dimensional feature tensors and form a three-dimensional feature tensor library, and divide them into a training set and a test set in a ratio of 8:2; Step S5, constructing a CNN-ECA-MMoE multi-task learning model to output the results of feature sharing and feature extraction; The CNN-ECA-MMoE multi-task learning model includes multiple expert subnets and multiple gating networks. The expert subnets include a two-dimensional CNN layer, a pooling layer, and an efficient channel attention mechanism ECA connected in sequence; the expert subnets are used to learn specific feature representations of input data; the gating network is used to dynamically adjust the output weights of each expert subnet according to the input data, and combine the output features of each expert subnet; Step S6, construct three independent long short-term memory networks LSTM, embed them into the back end of the CNN-ECA-MMoE multi-task learning model as the task-specific sub-network of the model, and connect two fully connected layers after each long short-term memory network LSTM to obtain the CNN-ECA-MMoE-LSTM multi-task learning model, and train the model using the training set divided in step S4; The LSTM network is used to capture the temporal dependency of the results of feature sharing and feature extraction; the fully connected layer is used to perform feature transformation on the output of the LSTM network, output the prediction result of the multivariate load, and perform denormalization processing to obtain the prediction value of the same scale as the input data; Step S7, according to the method in step S1 to step S4, obtain the three-dimensional feature tensor corresponding to the time to be predicted; input the three-dimensional feature tensor corresponding to the time to be predicted into the CNN-ECA-MMoE-LSTM multi-task learning model trained in step S6 to extract features and capture the timing dependency, and obtain the prediction results of electric load, cooling load and heating load.
[0011] Furthermore, the historical multi-load sequence in step S1 includes a historical electric load sequence, a historical cooling load sequence, and a historical heating load sequence; the influencing factor characteristics in step S1 are meteorological factor characteristics associated with the electric load, cooling load, and heating load, and the meteorological factor characteristics include air temperature, cloud type, dew point, relative humidity, solar zenith angle, surface albedo, pressure, precipitation, wind direction, and wind speed; the meteorological factor characteristics are analyzed by a Pearson correlation analysis method, and the Pearson correlation coefficients between each meteorological factor characteristic and the electric load, cooling load, and heating load are calculated respectively, and the absolute values of the Pearson correlation coefficients between each meteorological factor characteristic and the three loads are taken and then averaged to obtain the correlation coefficients between each meteorological factor characteristic and the three loads, and the correlation coefficients are sorted in descending order, and the meteorological factor characteristics corresponding to the first three groups of correlation coefficients are selected to construct an influencing factor feature library.
[0012] Furthermore, the meteorological factor characteristics corresponding to the first three groups of correlation coefficients are temperature, precipitation, and dew point.
[0013] Furthermore, the specific steps of step S3 are: evaluating the K modal components of each load obtained in step S2 from three perspectives: complex characteristics, coupling characteristics, and frequency characteristics, respectively, calculating the complex characteristics score, coupling characteristics score, and frequency characteristics score of each modal component, and integrating the three scores to obtain a hybrid characteristics evaluation score; then classifying each modal component according to the hybrid characteristics evaluation score, and aggregating the modal components belonging to the same component category corresponding to each load, thereby converting the K modal components corresponding to each load into a group of three types of aggregation sequences, and obtaining three types of aggregation sequences corresponding to each load; the three types of aggregation sequences corresponding to each load contain three sequence components, namely, a trend component, a fluctuation component, and a random component; and then updating the three types of aggregation sequences corresponding to each load to the historical multivariate load sequence library obtained in step S1, replacing the historical multivariate load sequence before reconstruction, and obtaining an updated historical multivariate load sequence library.
[0014] Furthermore, in step S3, the specific method for classifying each modal component according to the mixed characteristic evaluation score is: classifying the modal components with a mixed characteristic evaluation score greater than or equal to 0 and less than 0.4 as trend components, classifying the modal components with a mixed characteristic evaluation score greater than or equal to 0.4 and less than 0.7 as fluctuation components, and classifying the modal components with a mixed characteristic evaluation score greater than or equal to 0.7 and less than 1 as random components.
[0015] Furthermore, the number of the expert subnetworks is 2 to 6, and the number of the gated networks is 3.
[0016] Furthermore, the efficient channel attention mechanism ECA captures the channel global information of the input feature map through the global maximum pooling layer GAP, and then adaptively learns the local interaction relationship between channels through a one-dimensional convolutional layer, generates attention weights using the Sigmoid function, and dynamically adjusts the response strength of each channel of the input feature map using the attention weights, thereby obtaining an output feature map containing key features.
[0017] Advantages and beneficial effects of the present invention:
[0018] (1) This paper comprehensively evaluates the complex characteristics, coupling characteristics and frequency characteristics of each modal component, introduces the hybrid characteristic evaluation method (HFEM), comprehensively considers the information of the modal components from multiple angles, reconstructs each modal component to obtain trend component, fluctuation component and random component, and converts the modal components into three types of aggregated sequences. While eliminating redundant information and reducing noise, the representation of effective sequence information is further clarified, thereby reducing the interference caused by task differences in multi-task learning.
[0019] (2) The multi-load joint prediction method proposed in the present invention, which integrates decomposition and reconstruction technology with multi-task learning, takes into account the coupling characteristics of multiple loads and can use the correlation between multiple loads to improve the training effect of the model and improve the load prediction accuracy.
[0020] (3) The present invention breaks through the inherent limitations of multi-task learning using a hard sharing mechanism, improves the multi-task learning model MMoE, and integrates the two-dimensional CNN layer and the efficient channel attention mechanism ECA to form a CNN-ECA module to replace the expert subnet composed of the fully connected layer in the multi-task learning model MMoE, forming an improved MMoE multi-task learning model. The high-dimensional feature extraction capability of the two-dimensional CNN layer is combined with the flexible task allocation mechanism of the multi-task learning model MMoE, while introducing the efficient channel attention mechanism ECA, which enhances the inherent ability of the expert subnet to summarize differential features. Therefore, the model proposed in the present invention can take into account the differences in the degree of correlation between multiple loads, ensure that each task can obtain the most effective information, and achieve better prediction results. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 A flowchart of a multi-load joint forecasting method integrating decomposition and reconstruction technology with multi-task learning provided in an embodiment of the present invention;
[0022] Figure 2 A graph showing the correlation analysis results between the multi-element load and the meteorological factor characteristics provided by an embodiment of the present invention;
[0023] Figure 3 Three types of aggregation sequences after electrical load reconstruction provided by the embodiment of the present invention;
[0024] Figure 4 Three types of aggregation sequences after cold load reconstruction provided by the embodiment of the present invention;
[0025] Figure 5 Three types of aggregation sequences after heat load reconstruction provided by the embodiment of the present invention;
[0026] Figure 6 A schematic diagram of the structure of a multi-task learning model MMoE provided in an embodiment of the present invention;
[0027] Figure 7 A schematic diagram of the structure of an efficient channel attention mechanism ECA provided by an embodiment of the present invention;
[0028] Figure 8 A schematic diagram of the structure of the CNN-ECA-MMoE-LSTM multi-task learning model provided in an embodiment of the present invention;
[0029] Fig. 9The prediction results of the electric load under different reconstruction strategies provided by the embodiment of the present invention;
[0030] Fig.10 The prediction results of cooling load under different reconstruction strategies provided by the embodiment of the present invention;
[0031] Fig.11 The prediction results of heat load under different reconstruction strategies provided by the embodiment of the present invention;
[0032] Fig.12 The prediction result of the electric load in the ablation comparison experiment provided by the embodiment of the present invention;
[0033] Fig.13 The prediction result of the cold load in the ablation comparison experiment provided by the embodiment of the present invention;
[0034] Fig.14 The prediction result of heat load in ablation comparison experiment provided by the embodiment of the present invention;
[0035] Fig.15 The prediction results of multiple loads under different sharing mechanisms provided by the embodiment of the present invention. DETAILED DESCRIPTION
[0036] The embodiments of the present invention are further described in detail below with reference to the accompanying drawings.
[0037] like Figure 1 As shown, a multi-load joint prediction method integrating decomposition and reconstruction technology with multi-task learning of the present invention comprises the following steps:
[0038] Step S1, obtaining the historical load sequence of the electric load, cooling load and heating load in the integrated energy system, forming a historical multi-load sequence and forming a historical multi-load sequence library; at the same time, obtaining the data sequence of the characteristics of the influencing factors that affect the electric load, cooling load and heating load in the integrated energy system, forming an influencing factor characteristic library;
[0039] The historical multi-load sequence in the historical multi-load sequence library in step S1 includes the historical electric load sequence, historical cooling load sequence, and historical heating load sequence of M time steps before each prediction moment; the influencing factor characteristics are meteorological factor characteristics that are highly correlated with the electric load, cooling load, and heating load, and the meteorological factor characteristics include air temperature, cloud type, dew point, relative humidity, solar zenith angle, surface albedo, pressure, precipitation, wind direction, and wind speed; the meteorological factor characteristics are analyzed by the Pearson correlation analysis method, and the Pearson correlation coefficients between the air temperature, cloud type, dew point, relative humidity, solar zenith angle, surface albedo, pressure, precipitation, wind direction, wind speed and the electric load, cooling load, and heating load are calculated respectively, the absolute value of the Pearson correlation coefficient between each meteorological factor characteristic and the three loads is taken and then the average is calculated to obtain the correlation coefficient between each meteorological factor characteristic and the three loads, the correlation coefficients are sorted in descending order (i.e., sorted from large to small), and the meteorological factor characteristics corresponding to the first three groups of correlation coefficients are selected to construct the influencing factor characteristic library. In this embodiment, the meteorological factor characteristics corresponding to the first three groups of correlation coefficients are temperature, precipitation, and dew point.
[0040] The energy consumption behavior of users of the integrated energy system is related to many factors, including meteorological factors, economic factors, social factors, etc. Considering the difficulty of data acquisition, this embodiment only selects meteorological factors for variable correlation analysis. The Pearson correlation coefficient is used to analyze the intrinsic correlation of the three types of load sequences of electricity, cooling, and heat in the integrated energy system, as well as the correlation between the three types of load sequences of electricity, cooling, and heat in the integrated energy system and the characteristics of meteorological factors. The specific calculation formula of the Pearson correlation coefficient is: , Among them, X n Represents the data of characteristic variable X at the nth time point. Characteristic variable X is electric load, cooling load or heating load. Y n Represents the data of the characteristic variable Y at the nth time point, where the characteristic variable Y is the electric load, cooling load, heating load, air temperature, cloud type, dew point, relative humidity, solar zenith angle, surface albedo, pressure, precipitation, wind direction or wind speed; Represents the mean of the data of the feature variable X at all time points, represents the mean of the data of all time points of the characteristic variable Y, r represents the Pearson correlation coefficient, and the value range of r is [-1,1]. When the value of r is -1, it means that the characteristic variable X and the characteristic variable Y are completely negatively correlated, when the value of r is +1, it means that the characteristic variable X and the characteristic variable Y are completely positively correlated, and when the value of r is 0, it means that there is no correlation between the characteristic variable X and the characteristic variable Y; N represents the total number of time points in the historical load series, and n represents the nth time point in the historical load series.
[0041] This embodiment uses the data from January 1, 2022 to December 31, 2022 in the Integrated Energy System (IES) data set of the Tempe campus of Arizona State University for simulation. The data is provided at intervals of 1 hour and includes data on electric load, cooling load, heating load, and meteorological factor characteristic data. The meteorological factor characteristic data is obtained from the weather station at Phoenix International Airport, which is closest to the Tempe campus. The collected meteorological factor characteristics include temperature, cloud type, dew point, relative humidity, solar zenith angle, surface albedo, pressure, precipitation, wind direction, and wind speed. The Pearson correlation coefficient between the data of electric load, cooling load, heating load and the meteorological factor characteristic data is calculated, and the correlation analysis results are obtained as shown in the figure below. Figure 2 shown.
[0042] Figure 2 The correlation analysis results shown in the figure show that any two of the three loads, namely, electric load, cooling load and heating load, show significant correlation; the temperature shows strong correlation with the three loads; the precipitation is strongly correlated with the electric load and cooling load, and moderately correlated with the heating load; the dew point and pressure show moderate correlation with the three loads. This embodiment considers the correlation strength between each meteorological factor feature and the three loads, calculates the average value of the absolute value of the Pearson correlation coefficient between each meteorological factor feature and the three loads, obtains the correlation coefficient between each meteorological factor feature and the three load sequences, and sorts the correlation coefficients between all meteorological factor features and the three loads in descending order. The correlation coefficients from high to low are temperature, precipitation, dew point, pressure, relative humidity, surface albedo, solar zenith angle, wind direction, wind speed, and cloud type. In prediction, considering more feature variables does not mean higher accuracy. Inputting feature variables with low correlation will make the training effect of the model worse while bringing computational complexity. Based on this consideration, this embodiment only selects meteorological factor features that are strongly correlated with the multivariate load sequence: temperature, precipitation, and dew point.
[0043] The influencing factor feature library in step S1 includes the meteorological factor features corresponding to the M time steps before the time to be predicted, and the meteorological factor features do not need to be decomposed and reconstructed. Since the same time of adjacent days is highly correlated, considering the daily periodicity, the value of the interval time step M of the time to be predicted is selected as 24 (in hours).
[0044] Step S2, preliminary decomposition based on variational mode decomposition (VMD) method, the historical load sequences of electric load, cooling load and heating load (i.e. historical electric load sequence, historical cooling load sequence and historical heating load sequence) in the historical multivariate load sequence library obtained in step S1 are respectively preliminarily decomposed by variational mode decomposition (VMD) method, and K modal components (IMFs) are obtained for each load.
[0045] Variational mode decomposition (VMD) is an adaptive sequence decomposition method that solves the problem of modal aliasing caused by noise interference in empirical mode decomposition (EMD) through Hilbert transform and frequency bandwidth minimization. The variational mode decomposition method adaptively decomposes the sequence into multiple discrete subsequences, which reduces the complexity of the sequence to a certain extent. The solution process of variational mode decomposition is as follows: , in, represents the Lagrangian function corresponding to the problem to be solved, represents the kth modal component, represents the center frequency of the kth modal component, v is the penalty factor, K represents the total number of set modal components, represents the partial derivative with respect to time, represents the Dirac function, j represents the imaginary unit, t represents the tth time point in the historical load sequence, represents the complex exponential function, f(t) is the historical load sequence, represents the value of the Lagrange multiplier at the tth time point;
[0046] The K modal components of each load are aggregated to form the modal aggregation sequence of the corresponding load, and the residual component (Res) can be obtained by subtracting the modal aggregation sequence from the historical load sequence of the corresponding load. The calculation formula of the residual component is: , Among them, R(t) is the residual component obtained by variational mode decomposition, K is the total number of set modal components, f(t) is the historical load sequence, represents the kth modal component.
[0047] The residual component (Res) is the part of the historical load sequence that cannot be further decomposed. It contains a lot of noise and is not included in the data processing of subsequent prediction tasks, but can be used to evaluate the stability of the division of the number of modal components. The smaller the residual component, the more effective information is retained after removing the noise during the initial decomposition. The selection of the total number of modal components K directly affects the quality of the decomposition results and the performance of subsequent tasks. In this embodiment, the mean square error is used as an evaluation indicator to determine the value of the total number of modal components K, so as to obtain K modal components with stable and periodic trends. Set the total number of modal components K to [K min ,K max ] is a positive integer between [K min ,K max Among all the values between ], the mean square error between the historical load sequence and the modal aggregation sequence is set to the value corresponding to the minimum mean square error as the total number of modal components K. In this embodiment, K min =1,K max= 10. The formula for calculating the mean square error is: , Among them, R MSE represents the mean square error between the historical load sequence and the modal aggregation sequence, R(t) is the residual component obtained by variational modal decomposition, N represents the total number of time points in the historical load sequence, and t represents the tth time point in the historical load sequence.
[0048] The test experiments of the present invention show that when the variational modal decomposition method is used to decompose the historical electric load sequence, the historical cooling load sequence and the historical thermal load sequence, when the number of modal components K is 8, the mean square error between the modal aggregation sequence and the historical multivariate load sequence is minimized.
[0049] Step S3, based on the reconstruction strategy of the hybrid characteristic evaluation method (HFEM), the K modal components of each load obtained in step S2 are reconstructed from multiple angles by the hybrid characteristic evaluation method (HFEM), and the K modal components corresponding to each load are respectively converted into a group of three types of aggregate sequences; then the three types of aggregate sequences corresponding to each load are updated to the historical multivariate load sequence library obtained in step S1, to obtain an updated historical multivariate load sequence library;
[0050] The specific steps of step S3 are: evaluating the K modal components of each load obtained in step S2 from three perspectives of complex characteristics, coupling characteristics, and frequency characteristics, respectively, calculating the complex characteristics score, coupling characteristics score, and frequency characteristics score of each modal component, and integrating the three scores to obtain a hybrid characteristics evaluation score (i.e., HFEM score); then using the hybrid characteristics evaluation score as a key indicator for classification and reconstruction, classifying each modal component into a trend component, a fluctuation component, and a random component, and aggregating the modal components belonging to the same component category corresponding to each load, thereby obtaining a reconstruction sequence considering multi-modal characteristics, that is, obtaining three types of aggregation sequences corresponding to each load, and the three types of aggregation sequences corresponding to each load all contain trend components, fluctuation components, and random components; then updating the three types of aggregation sequences corresponding to each load to the historical multi-load sequence library obtained in step S1, replacing the historical electric load sequence, historical cooling load sequence, and historical heating load sequence in the historical multi-load sequence before reconstruction, and obtaining an updated historical multi-load sequence library.
[0051] This embodiment adopts a reconstruction strategy based on the hybrid characteristic evaluation method (HFEM), which considers the characteristics of the load sequence from multiple perspectives, and reconstructs the modal components obtained by the preliminary decomposition in step S2 from three perspectives: complex characteristics, coupling characteristics, and frequency characteristics, in order to improve the reconstruction quality of the load sequence, so that the reconstructed load sequence can more accurately reflect the physical characteristics and internal modes of the original load sequence. The specific steps are:
[0052] Step S31, calculating the complex characteristic score of the modal component:
[0053] The historical load sequence of each load in the historical multi-load sequence library of the integrated energy system has the characteristics of non-stationarity, nonlinearity and high complexity. The complexity can help distinguish useful load sequences from random components, so it is extremely important for the reconstruction of load sequences. High complexity means that the load sequence has higher uncertainty and disorder, and has rich detailed fluctuations. Low complexity means that the load sequence is more regular and stable, and may reflect the key patterns or main trends in the load sequence. For the measurement of complex characteristics, this embodiment uses sample entropy for calculation, and the specific method is as follows:
[0054] Step S311, define the modal component as , x1 represents the first data point of the modal component, x N represents the Nth data point of the modal component, where N is the total number of time points in the historical load series; by embedding dimension m and time delay , rearrange the modal components according to time delay and construct the phase space matrix Y G , the phase space matrix Y G The representation is as follows: , Where m is the embedding dimension, is the time delay, G is the number of rows in the matrix, ; The first data point corresponding to the modal component, The corresponding modal component data points, The corresponding modal component Data points, similarly, Corresponding to the Gth data point of the modal component, The corresponding modal component data points, The corresponding modal component The data point is the Nth data point of the corresponding modal component. This step ensures that the embedding vector extracted from the modal component can appropriately reflect the dynamic behavior of the modal component.
[0055] Step S312, in the phase space matrix Y G In the above example, each row of data elements is regarded as an embedded vector. After arranging the data elements in the embedded vector in ascending order, the arrangement pattern corresponding to the embedded vector is obtained. .
[0056] Step S313, calculate the phase space matrix Y G All possible permutations of The frequencies that appear in the modal components are denoted by , the total number of permutation patterns is recorded as , and then based on the frequency distribution, the Shannon entropy formula is used to calculate the permutation entropy of the modal component. The calculation formula is: , Where PE represents the permutation entropy of the modal component, Indicates the arrangement mode The frequencies that appear in the modal components, The natural logarithm of the frequency is used to quantify uncertainty or information. This step quantifies the complexity of the sequence and can reflect the degree of disorder of the local structure in the sequence.
[0057] Step S314, normalizing the permutation entropy of the modal component, the normalization formula is expressed as: , In the formula, represents the normalized permutation entropy, representing the complex feature score; represents the normalization factor; where the normalized permutation entropy The range of is in [0,1]. Within the range, the larger the value of the normalized permutation entropy is, the more complex the corresponding modal component is, and the smaller the value of the normalized permutation entropy is, the more regular the corresponding modal component is.
[0058] Step S32, calculating the coupling characteristic score of the modal component:
[0059] This embodiment considers the linear and nonlinear correlation between the modal components of each type of load and the historical load series of the other two types of loads, and uses this as the basis for coupling characteristic analysis. On the basis of calculating the Pearson correlation coefficient, the gray correlation analysis method (GRA) is used to calculate the load nonlinear correlation coefficient between the modal components of each type of load and the historical load series of the other two types of loads, and obtain the coupling characteristic scores of the modal components of each type of load; taking the electric load as an example, the calculation steps of the coupling characteristic scores of the modal components of the electric load are as follows:
[0060] Step S321, calculating the i-th modal component of the electric load Pearson correlation coefficient with historical cooling load series , and calculate the i-th modal component of the electric load Pearson correlation coefficient with historical heat load series .
[0061] Step S322, using the grey relational analysis method (GRA) to calculate the i-th modal component of the electric load Grey correlation coefficient between the historical cooling load sequence , and the i-th modal component of the electric load Grey correlation coefficient between the historical heat load sequence , Grey correlation coefficient is used to measure the development trend and similarity between multiple systems or multiple factors, so it is more suitable for analyzing non-stationary sequences. The calculation formula of the grey correlation coefficient is as follows: , in, represents the i-th modal component of the electric load The grey correlation coefficient between the historical cooling load sequence, represents the i-th modal component of the electric load The grey correlation coefficient between the historical heat load sequence and the historical load sequence; N represents the total number of time points in the historical load sequence, t represents the tth time point in the historical load sequence, represents the i-th modal component of the electric load The value at the tth time point in represents the value at the tth time point in the historical cooling load series, represents the value at the tth time point in the historical heat load sequence, It is expressed as the resolution coefficient. In this embodiment, the resolution coefficient The value of is 0.5; Indicates the minimum absolute value. Indicates the maximum absolute value.
[0062] Step S323, for the i-th modal component of the electric load obtained in step S321 and step S322 Pearson correlation coefficient with historical cooling load series , the i-th modal component of the electric load Pearson correlation coefficient with historical heat load series , the i-th modal component of the electric load Grey correlation coefficient between the historical cooling load sequence , the i-th modal component of the electric load Grey correlation coefficient between the historical heat load sequence The average value of these four indicators is used to calculate the i-th modal component of the electrical load The coupling characteristic score The specific calculation formula is as follows: S e , i = [ 1 − 1 4 ( ρ e , c , i + ρ e , h , i + γ e , c o o l , i + γ e , h e a t , i ) ] .
[0063] The coupling characteristic score of the modal component of the cooling load is calculated in the same way as the coupling characteristic score of the modal component of the heating load.
[0064] Step S33, calculating the frequency characteristic score of the modal component:
[0065] The frequency characteristics of the load sequence directly reflect the changes of the load sequence on different time scales. Only analyzing the complex characteristics and coupling characteristics may not capture the time resolution of the load sequence. The high-frequency components in the load sequence may represent sudden events, noise or short-term fluctuations, while the low-frequency components in the load sequence represent the long-term trend of the load sequence. In order to consider the frequency size of the modal component as a whole, this embodiment uses the center frequency as the key indicator and uses the Hilbert transform widely used in analyzing non-stationary load sequences for calculation. The specific steps are as follows:
[0066] Step S331, perform Hilbert transform on each modal component, obtain the analytical time series corresponding to the modal component, and convert the analytical time series into polar coordinate form. The calculation formula of Hilbert transform is: , in, represents the analytical signal of the ith modal component, represents the value of the i-th modal component at the t-th time point, represents the i-th modal component at time point The value at , j represents the imaginary unit, represents the time delay in the integration, Represents the time difference between two consecutive time points. is the instantaneous amplitude of the modal component, represents the complex exponential function, is the instantaneous phase of the modal component.
[0067] Step S332, based on the instantaneous phase of the modal component Calculate instantaneous frequency , the specific calculation formula is: , in, Instantaneous phase The differential of , dt represents the differential with respect to time point t.
[0068] Step S333: calculate the instantaneous frequency of the entire time period N of the historical load sequence. Average to get the center frequency of the i-th modal component , the specific calculation formula is: ; Among them, the center frequency represents the frequency characteristic score, N represents the total number of time points in the historical load sequence, Indicates the instantaneous frequency Perform an integral operation, dt represents the differential with respect to the time point t.
[0069] Step S34, calculating the mixed characteristic evaluation score:
[0070] The mixed characteristic evaluation score (i.e., HFEM score) of each modal component is calculated according to the complex characteristic score, coupling characteristic score, and frequency characteristic score of each modal component calculated in step S31 to step S33. The calculation formula of the mixed characteristic evaluation score is as follows: , In the formula, is the mixed characteristic evaluation score of the i-th modal component, Score the complexity feature, Score the coupling characteristics, Score the frequency characteristics; is the weight coefficient of complex characteristics, is the coupling characteristic weight coefficient, is the frequency characteristic weight coefficient; in this embodiment, .
[0071] Step S35, reconstructing the modal components, the specific steps are:
[0072] Classify each modal component according to the hybrid characteristic evaluation score of each modal component obtained in step S34, classify the modal components with hybrid characteristic evaluation scores in the interval [0,0.4) (i.e., greater than or equal to 0 and less than 0.4) as trend components, classify the modal components with hybrid characteristic evaluation scores in the interval [0.4,0.7) (i.e., greater than or equal to 0.4 and less than 0.7) as fluctuation components, and classify the modal components with hybrid characteristic evaluation scores in the interval [0.7,1) (i.e., greater than or equal to 0.7 and less than 1) as random components, aggregate the modal components belonging to the same component category corresponding to each load (i.e., add the modal components), thereby converting the K modal components corresponding to each load into a group of three types of aggregated sequences, and the three types of aggregated sequences corresponding to each load contain three sequence components, namely, trend component, fluctuation component, and random component.
[0073] In this embodiment, the data from January 1, 2022 to December 31, 2022 in the Integrated Energy System (IES) dataset of the Tempe campus of Arizona State University are used for testing experiments. Taking the electric load as an example, the hybrid characteristic evaluation scores of modal component 1 (IMF1) to modal component 8 (IMF8) of the electric load are calculated. The results are shown in Table 1 below.
[0074] Table 1 Calculation results of hybrid characteristic evaluation scores of modal components of electrical loads:
[0075] According to steps S31 to S35, the reconstruction of the historical multi-load sequence is completed, and three types of aggregate sequences for each load are obtained respectively. The historical load sequence of each load in the historical multi-load sequence library in step S1 is updated to three types of aggregate sequences corresponding to each type of load, and an updated historical multi-load sequence library is obtained.
[0076] The three types of aggregate sequences after reconstruction of the historical electric load sequence obtained by the test experiment in this embodiment are as follows: Figure 3 As shown in the figure, the three types of aggregate sequences after reconstruction of the historical cooling load sequence are as follows: Figure 4 As shown in the figure, the three types of aggregate sequences after reconstruction of the historical heat load sequence are as follows: Figure 5 shown.
[0077] Step S4, data normalization and reshaping, normalize the influencing factor feature library obtained in step S1 and the updated historical multivariate load sequence library obtained in step S3 with the time step M as the sequence length, and then merge them to obtain multiple groups of three-dimensional feature tensors and form a three-dimensional feature tensor library; then divide the data in the three-dimensional feature tensor library into a training set and a test set for model training and testing.
[0078] The specific process of step S4 is as follows: first, the sequence components in the three types of aggregated sequences in the updated historical multivariate load sequence library are normalized according to the component categories, and the data with a sequence length of M (corresponding to the time step) are sequentially taken out from each normalized sequence component to construct a historical multivariate load characteristic matrix, and multiple historical multivariate load characteristic matrices consisting of multiple sequences with a length of M are obtained. The dimension of the historical multivariate load characteristic matrix is [ , , ], where H r It represents the number of components after reconstruction, and its value is 3; M represents the time step, and its value is 24, in hours; represents the number of types of multivariate loads, and its value is 3. Then, the influencing factor features in the influencing factor feature library are normalized in the same way, and multiple influencing factor feature matrices consisting of multiple sequences of length M are obtained. The dimension of the influencing factor feature matrix is [ , , ], among which, Indicates the number of variables of influencing factors, the value is 3; It represents the dimension added to match the data format of the historical multivariate load characteristic matrix, and its value is 1. Finally, each historical multivariate load characteristic matrix and its corresponding influencing factor characteristic matrix are merged in the third dimension to obtain a dimension of [ , , ] will form a three-dimensional feature tensor library based on all the three-dimensional feature tensors constructed according to the data in the updated historical multivariate load sequence library; the data in the three-dimensional feature tensor library will be divided into training set and test set in a ratio of 8:2 for model training and testing.
[0079] Step S5, constructing a CNN-ECA-MMoE multi-task learning model to output the results of feature sharing and feature extraction.
[0080] The multi-task learning model MMoE consists of multiple expert subnets, multiple gate networks (Gates) composed of trainable parameters, and multiple task-specific subnets. The specific structure is as follows: Figure 6 As shown in the figure, the role of the expert subnet is to learn the specific feature representation of the input data and provide shared feature extraction capabilities for different tasks; the role of the gating network is to dynamically adjust the output weights of each expert subnet according to the input data to adapt to the needs of different tasks; the role of the task-specific subnet is to optimize for specific tasks and convert the output of the expert subnet into task-specific prediction results.
[0081] The CNN-ECA-MMoE multi-task learning model constructed by the present invention replaces the only fully connected layer in the expert subnet in the multi-task learning model MMoE with a two-dimensional CNN layer (i.e., a two-dimensional convolutional layer) to extract the spatial local features of the input data; and a pooling layer and an efficient channel attention mechanism ECA are sequentially connected after the two-dimensional CNN layer to form a CNN-ECA module as an expert subnet of the multi-task learning model, which is used to enhance the feature attention ability of the multi-task learning model MMoE; at the same time, combined with the gating network in the multi-task learning model MMoE, an improved MMoE multi-task learning model based on CNN and ECA is formed, named CNN-ECA-MMoE multi-task learning model. That is, the CNN-ECA-MMoE multi-task learning model constructed by the present invention includes multiple expert subnets and multiple gating networks, and the expert subnet includes a two-dimensional CNN layer, a pooling layer, and an efficient channel attention mechanism ECA connected in sequence.
[0082] The introduction of the efficient channel attention mechanism ECA aims to capture the dependencies between adjacent channels while strengthening the feature selection ability between channels, making the expert subnet in the multi-task learning model MMoE more targeted and accurate in processing features. In addition, the integration of the efficient channel attention mechanism ECA does not change the dimension of the shared feature data output by the multi-task learning model, so it has high adaptability. The integration of the efficient channel attention mechanism ECA can strengthen the multi-task learning model and enhance the ability to summarize differential features.
[0083] The channel attention mechanism has been proven to have great potential in improving the performance of deep convolutional neural networks. Compared with the widely used squeeze and excitation network (SENet), the efficient channel attention mechanism ECA can bring significant performance improvements to the model while adding a small number of parameters. Due to its lightweight design, it is feasible to use it in the parallel multi-task learning model MMoE with a certain computational complexity. Compared with the global channel interaction of the squeeze and excitation network (SENet), the efficient channel attention mechanism ECA adopts a local cross-channel interaction strategy based on one-dimensional convolution, which can capture the dependency between adjacent channels and adaptively select the appropriate convolution kernel size to ensure strong adaptability in network layers of different depths, thereby selectively emphasizing effective features and suppressing useless features. Based on this feature, the efficient channel attention mechanism ECA has certain similarities with the expert subnet in the multi-task learning model MMoE in realizing selective processing of information. Therefore, this embodiment embeds the efficient channel attention mechanism ECA into the expert subnet in the multi-task learning model MMoE to further enhance its inherent feature induction ability. Figure 7 As shown, the efficient channel attention mechanism ECA captures the channel global information of the input feature map through the global maximum pooling layer GAP, and then adaptively learns the local interaction relationship between channels through the one-dimensional convolution layer, generates attention weights using the Sigmoid function, and dynamically adjusts the response strength of each channel of the input feature map using the attention weights, thereby significantly improving the model's ability to express key features with lightweight calculations, and obtaining an output feature map containing key features. The specific calculation expression of the efficient channel attention mechanism ECA is as follows: , In the formula, Represents the channel descriptor output after being processed by the global maximum pooling layer. represents the input feature map, W represents the width of the input feature map, H represents the height of the input feature map, and c represents the cth channel; Represents the input feature map The value at position (o,q) and the cth channel; represents the weight vector of the cth channel, represents the Sigmoid function, Represents a one-dimensional convolution operation, which is used to calculate channel weights. Indicates the convolution kernel size; Represents the output feature map after weighting by the efficient channel attention mechanism ECA, Represents the output feature map The value at position (o,q) and the cth channel; represents the function for calculating the convolution kernel size, C represents the number of channels of the input feature map, represents the logarithmic function with base 2, Indicates that the result is an odd number. is the convolution kernel size scaling factor. In this embodiment, the convolution kernel size scaling factor Set to 2, b is the bias term coefficient, and in this embodiment, the bias term coefficient b is set to 1.
[0084] Step S6, constructing a CNN-ECA-MMoE-LSTM multi-task learning model; the specific steps are: constructing three independent long short-term memory networks LSTM, embedding them into the back end of the CNN-ECA-MMoE multi-task learning model as the task-specific sub-network of the model, and connecting two fully connected layers after each long short-term memory network LSTM to obtain a CNN-ECA-MMoE-LSTM multi-task learning model, and training the model using the training set divided in step S4.
[0085] In order to make the long short-term memory network LSTM fully utilize the spatial features extracted by the two-dimensional CNN layer and the efficient channel attention mechanism ECA in the CNN-ECA-MMoE multi-task learning model, the present invention does not flatten the output of the CNN-ECA-MMoE multi-task learning model or add a fully connected layer, but retains the original structure of the time step to the maximum extent, and multiplies the dimension representing the number of feature maps by the height of the feature map as the feature dimension of the time step. Therefore, the dimension of the data input to the long short-term memory network LSTM is ,in, It indicates the number of feature maps output by the CNN-ECA-MMoE multi-task learning model after feature sharing and feature extraction. Represents the height of the feature map input to the long short-term memory network LSTM, Represents the width of the feature map input to the long short-term memory network LSTM; this detailed processing can effectively maintain the integrity of the spatial information extracted by the two-dimensional CNN layer and avoid the loss of time step information caused by flattening.
[0086] like Figure 8As shown, the CNN-ECA-MMoE-LSTM multi-task learning model includes an input layer, a CNN-ECA-MMoE multi-task learning model, a task-specific subnetwork and two fully connected layers; wherein, the number of expert subnetworks in the CNN-ECA-MMoE multi-task learning model is 2 to 6; and the number of two-dimensional CNN layers in each expert subnetwork is 1 to 3. The specific workflow of the CNN-ECA-MMoE-LSTM multi-task learning model is as follows: a three-dimensional feature tensor (dimension: height × width × number of channels = 3 × 24 × 4, height corresponds to the number of components after reconstruction, width corresponds to the time step M, and number of channels corresponds to the number of feature maps) is input into the CNN-ECA-MMoE-LSTM multi-task learning model as input data, and multiple expert subnets in the CNN-ECA-MMoE multi-task learning model simultaneously obtain the input data and perform feature extraction and induction. In each expert subnet, the two-dimensional CNN layer extracts the spatial local features of the input data and condenses the common features of multiple channels to obtain a set of tensors with a dimension of 4 × 25 × 32, and then the pooling layer is used to extract the significant features of the tensor and reduce the With low computational complexity, a tensor with a dimension of 3×25×32 is obtained. Finally, the efficient channel attention mechanism ECA further strengthens the feature attention ability of the expert subnetwork to form a unique expression of the expert subnetwork. At the same time, the gating network dynamically adjusts the output weights of each expert subnetwork and flexibly combines the output features of each expert subnetwork, thereby obtaining three groups of feature sharing and feature extraction results (that is, a feature tensor with a dimension of 25×96); subsequently, three long short-term memory networks LSTM capture the temporal dependencies of each group of feature sharing and feature extraction results respectively, and finally, two fully connected layers perform feature transformation on the output of the long short-term memory network LSTM, output the prediction results of the multivariate load, and perform denormalization to obtain the prediction value with the same data scale as the input data.
[0087] The selection of hyperparameters of the CNN-ECA-MMoE-LSTM multi-task learning model has a crucial impact on performance. This embodiment uses the control variable method to find the optimal hyperparameters. First, the number of iterations (Epoch), batch size, and learning rate are set to fixed values, and the rest of the model is kept at a fixed number of layers. The number of layers of the two-dimensional CNN layer is first set to one layer, and the convolution kernel size in the two-dimensional CNN layer is continuously adjusted. After determining the optimal convolution kernel size, it is fixed to adjust the convolution stride. After determining the optimal convolution stride, the number of layers of the two-dimensional CNN layer is adjusted. After the test experiment is verified, when the number of layers of the two-dimensional CNN layer is 3, the convolution stride is 1, the convolution kernel size in the first two-dimensional CNN layer is 2×2, the convolution kernel size in the second two-dimensional CNN layer is 3×3, and the convolution kernel size in the third two-dimensional CNN layer is 3×3, the prediction result of the test is the most accurate. Next, when the optimal parameters of the two-dimensional CNN layer are fixed, the number of neurons in the long short-term memory network LSTM and the fully connected layer are gradually adjusted, and the remaining hyperparameters are determined in the same way. In addition, due to the use of the framework structure of the multi-task learning model MMoE, it is necessary to determine the number of expert subnets. Experimental verification shows that when the number of expert subnets is 4, the prediction accuracy of cold load and heat load reaches the highest, while the prediction of electric load reaches the best effect when the number of expert subnets is 5, but it is only 2% higher than when the number of expert subnets is 4, and the computational complexity is increased by 10%. In order to achieve a balance between computational efficiency and prediction accuracy as much as possible, this embodiment uses 4 expert subnets for prediction.
[0088] In this embodiment, the CNN-ECA-MMoE-LSTM multi-task learning model constructed in step S6 sets the number of tasks to 3, namely, 1 electric load prediction task, 1 cold load prediction task, and 1 hot load prediction task. The number of gated networks corresponds to the number of tasks, and the number of gated networks in this embodiment is 3. The number of long short-term memory networks LSTM is 3, corresponding to the prediction tasks of electric load, cold load, and hot load, respectively, and the number of layers of the long short-term memory network LSTM selected by the task-specific subnetwork of each task is 1 layer. The number of fully connected layers is 2 layers, the first fully connected layer is used to perform feature transformation on the output of the last time step of the long short-term memory network LSTM, and the second fully connected layer is used to output the prediction result of the multivariate load according to the result of the feature transformation. The weight ratio of the multi-task loss function when training the model is set to 1:1:1, that is, the proportion of information transmitted by the three tasks is the same; the activation function used when training the model is the Relu function, the optimizer is the Adam optimizer, and the loss function is the mean absolute error (MAE). The construction and training of the CNN-ECA-MMoE-LSTM multi-task learning model all use the general deep learning toolkit (Pytorch) in the Python programming language. The hyperparameters of the CNN-ECA-MMoE-LSTM multi-task learning model used in this embodiment are shown in Table 2.
[0089] Table 2 Hyperparameters of the CNN-ECA-MMoE-LSTM multi-task learning model:
[0090] When training the model, the three-dimensional feature tensor in the three-dimensional feature tensor library training set obtained in step S4 is taken out according to a batch size of 24 and input into the CNN-ECA-MMoE-LSTM multi-task learning model constructed in step S6 for iterative training.
[0091] Step S7, according to the method in step S1 to step S4, obtain the three-dimensional feature tensor corresponding to the time to be predicted; input the three-dimensional feature tensor corresponding to the time to be predicted into the CNN-ECA-MMoE-LSTM multi-task learning model trained in step S6 to extract features and capture the timing dependency, and obtain the prediction results of electric load, cooling load and heating load.
[0092] The working process of the present invention is as follows: a multi-load joint prediction method of the present invention which integrates decomposition and reconstruction technology with multi-task learning, first obtains the historical electric load sequence, historical cooling load sequence, and historical heat load sequence of the comprehensive energy system, and simultaneously obtains the data sequence of the characteristics of the influencing factors that affect the electric load, cooling load, and heat load of the comprehensive energy system, to form a historical multi-load sequence library and an influencing factor characteristic library; then, the historical electric load sequence, the historical cooling load sequence, and the historical heat load sequence are preliminarily decomposed by the variational mode decomposition method, and reconstructed based on the hybrid characteristic evaluation method (HFEM), to obtain three types of aggregated sequences for each load and update the historical multi-load sequence library; then, the influencing factors are updated. The feature library and the updated historical multi-load sequence library are normalized to obtain the historical multi-load feature matrix and the influencing factor feature matrix. The historical multi-load feature matrix and the influencing factor feature matrix are merged in the third dimension to obtain a three-dimensional feature tensor. Then, a CNN-ECA-MMoE multi-task learning model is constructed to output the results of feature sharing and feature extraction. Then, three long short-term memory networks (LSTMs) are constructed and embedded in the back end of the CNN-ECA-MMoE multi-task learning model as task-specific subnetworks, and two fully connected layers are connected after each long short-term memory network LSTM to obtain a CNN-ECA-MMoE-LSTM multi-task learning model. Finally, the CNN-ECA-MMoE-LSTM multi-task learning model finally constructed is trained, and the three-dimensional feature tensor corresponding to the time to be predicted is input into the CNN-ECA-MMoE-LSTM multi-task learning model to obtain the prediction results of electric load, cooling load, and heating load respectively.
[0093] This embodiment uses the data from January 1, 2022 to December 31, 2022 in the Integrated Energy System (IES) data set of the Tempe campus of Arizona State University for simulation. The data is provided at intervals of 1 hour, including data on electric load, cooling load, heating load, and characteristic data of meteorological factors. The characteristic data of meteorological factors are obtained from the weather station at Phoenix International Airport, which is closest to the Tempe campus. A three-dimensional feature tensor library is constructed based on the acquired data and divided into a training set and a test set in a ratio of 8:2, and input into the CNN-ECA-MMoE-LSTM multi-task learning model constructed by the present invention for training and prediction. The load forecast value of the corresponding date output by the CNN-ECA-MMoE-LSTM multi-task learning model is compared with the actual load power value of the day, and the root mean square error (RMSE) and the mean absolute percentage error (MAPE) are calculated. Two evaluation indicators, the root mean square error (RMSE) and the mean absolute percentage error (MAPE), are calculated. The following are detailed experimental results.
[0094] (1) Effectiveness analysis of the reconstruction strategy based on the hybrid characteristic evaluation method (HFEM):
[0095] In order to verify the contribution of the reconstruction strategy based on the hybrid characteristic evaluation method (HFEM) proposed in the present invention in the reconstruction of load data, the reconstruction strategy based on the hybrid characteristic evaluation method of the present invention is compared with the reconstruction based on complex characteristics, the reconstruction based on coupling characteristics, and the reconstruction based on frequency characteristics that only consider a single characteristic, and the influence of data reconstruction on the accuracy of energy load prediction under different reconstruction strategies is tested. Fig. 9 , Fig.10 , Fig.11 The prediction results of electric load, cooling load and heating load under different reconstruction strategies are shown respectively; the specific results of the two evaluation indicators of root mean square error (RMSE) and mean absolute percentage error (MAPE) under different reconstruction strategies are shown in Table 3. It can be seen that the reconstruction strategy based on the hybrid characteristic evaluation method proposed in the present invention has the smallest prediction error, followed by the reconstruction method based on complex characteristics and the reconstruction method based on frequency characteristics. The overall prediction result of the model using the coupling characteristic reconstruction method is the worst, but its accuracy in electric load prediction is second only to the reconstruction strategy based on the hybrid characteristic evaluation method proposed in the present invention. This is because the reconstruction calculation based on the coupling characteristic takes into account the similarity of all subsequences with the original sequence, so that all modal components except modal component 1 in Table 1 are summarized as the fluctuation component of the intermediate frequency, thereby simplifying the representation of the trend component, so it is only suitable for electric loads with strong regularity.
[0096] Table 3 Comparison results of prediction errors under different reconstruction strategies:
[0097] The experimental results show that in the reconstruction stage after the initial decomposition, the reconstruction strategy based on the hybrid characteristic evaluation method can more comprehensively utilize the different properties inherent in each modal component. Specifically, the coupled characteristic evaluation is more suitable for sequences with strong regularity, while the advantages of complex characteristic and frequency characteristic evaluation are to distinguish and capture medium-frequency and high-frequency features, and are good at processing sequences with high complexity. The present invention combines the advantages of the three characteristic evaluation methods and can deeply evaluate the properties of each modal component from multiple dimensions. It not only breaks through the limitations of a single indicator, but also realizes a more refined classification in the reconstruction stage, thereby improving the accuracy of the overall analysis.
[0098] (2) Model ablation experiment analysis:
[0099] In order to verify the contribution of the two-dimensional CNN layer, the efficient channel attention mechanism ECA, and the multi-task learning model MMoE to the model proposed in the present invention, this embodiment selects the LSTM model, the CNN-LSTM model, the CNN-MMoE-LSTM model, and the CNN-ECA-MMoE-LSTM multi-task learning model for prediction under the same decomposition and reconstruction strategy, and adjusts the hyperparameters of the above four models according to the control variable method to compare the prediction results of the models under the optimal parameters. Table 4 summarizes the evaluation results of the above four models. It can be seen from Table 4 that the improvement of each model from top to bottom is significant. Among all models, the prediction accuracy of each load of the LSTM model is the lowest, indicating that it has certain limitations when processing different load types alone. While the two-dimensional CNN layer captures the long-term load pattern in the time dimension, it also captures the spatial correlation information of different aggregation modes in space, extracts the changing law and internal connection of the long-term load trend and local fluctuation, and further strengthens the characterization of the multi-load coupling relationship through the fusion of different channels. Therefore, compared with the LSTM model, the CNN-LSTM model has a significant improvement in the single-task case, especially in the prediction of electric load, the MAPE evaluation index has increased by 54.1%. By integrating the multi-task learning mechanism, the prediction accuracy of the CNN-MMoE-LSTM model is further improved on the basis of the CNN-LSTM model, especially in the prediction of cold load and heat load, the prediction accuracy is improved by 49.0% and 68.5% respectively compared with the CNN-LSTM model. In the integrated energy system (IES), the user's cold and hot load demand is supplied by energy conversion equipment, and the coupling between energy conversion equipment is very strong. The traditional single-task prediction cannot extract the coupling features, while multi-task learning can accurately capture the seasonal and temporal characteristics of cold load and heat load by sharing features and capturing the correlation between loads, thereby significantly improving its prediction accuracy. In the optimization and scheduling of integrated energy systems, accurate prediction of peak and valley values is particularly important. Accurate prediction of peak and valley values can optimize the power scheduling during high load periods and the scheduling strategy of energy storage during low load periods, thereby avoiding economic losses caused by wind and solar power abandonment. However, during peak and valley moments, the forecast target changes dramatically and with poor regularity, which makes forecasting difficult. Fig.12 , Fig.13 , Fig.14It shows that after integrating the multi-task learning model MMoE, compared with the CNN-LSTM model, the predicted values of the CNN-MMoE-LSTM model in each time period are closer to the true value of the load data, especially the peak and valley values of the cold load and heat load prediction are closer to the true value. This is because the data of different tasks complement each other under multi-task learning, and the model can understand the reasons for load changes from multiple angles, thereby more accurately predicting the peak and valley values of the load. The addition of the multi-task learning model MMoE based on the multi-task architecture has the greatest improvement in the accuracy of model prediction, and its averaged MAPE evaluation index has increased by 53.3%, which is much greater than the 36.4% improvement brought by adding only the two-dimensional CNN layer. Finally, this embodiment further integrates the efficient channel attention mechanism ECA on the basis of the CNN-MMoE-LSTM model, so that the expert subnet can focus on key features more accurately. The design concept of the expert subnet is to perform deep modeling for specific tasks or specific features, and the efficient channel attention mechanism ECA strengthens the ability of the expert subnet in extracting key features by focusing on local information and important channel features. The experimental results show that after the introduction of the efficient channel attention mechanism ECA, the prediction accuracy of electric load, cooling load and heating load increased by 26.7%, 27.1% and 10.1% respectively, and the prediction accuracy of the model for load peak and valley values was also further improved.
[0100] Table 4 Comparison results of model ablation experiments:
[0101] (3) Comparison of multi-task learning sharing mechanisms:
[0102] In order to further verify the prediction ability of the multi-task learning model framework adopted in the present invention, this embodiment compares the CNN-ECA-MMoE-LSTM multi-task learning model of the present invention with the traditional multi-task learning hard sharing mechanism, the single-task learning model (the single-task learning model uses the CNN-ECA-LSTM model) and a typical soft sharing mechanism (cross-stitch network) to compare the prediction accuracy between the CNN-ECA-MMoE-LSTM multi-task learning model of the present invention and the conventional multi-task learning model and the single-task learning model; at the same time, in order to ensure the rigor of the experiment, this embodiment uses the CNN-ECA module applied in the expert subnet of the CNN-ECA-MMoE-LSTM multi-task learning model of the present invention as a feature sharing learner for all multi-task models, and still retains the same setting as the traditional multi-task learning model MMoE in the selection of the task-specific subnetwork. The data from November 23, 2022 to December 31, 2022 in the Integrated Energy System (IES) dataset of the Tempe campus of Arizona State University in the United States are selected as test data. The comparison results are shown in Table 5. The comparison of MAPE evaluation indicators of various loads under different sharing mechanisms is shown in Table 5. Fig.15 As shown in the figure, the model with the multi-task mechanism of the CNN-ECA module and the multi-task learning model MMoE performed best in all tasks, followed by the soft sharing mechanism based on the cross-stitch network and the hard sharing mechanism of multi-task learning. The prediction accuracy of the single-task learning model was the worst.
[0103] Table 5 Comparison results of MAPE evaluation indicators under different sharing mechanisms:
[0104] Compared with the hard sharing mechanism, the Cross-Stitch Network has improved the prediction accuracy of various loads to varying degrees. This is because the soft sharing mechanism avoids the interference of features caused by forced sharing of feature space. In particular, the dynamic adjustment characteristics of the cross weights in the Cross-Stitch Network allow each task to actively select features with high correlation. Therefore, this soft sharing mechanism is suitable for scenarios where there is a certain correlation between tasks but differentiated feature expressions are required. Compared with the hard sharing mechanism, the mean MAPE evaluation index of the Cross-Stitch Network based on the soft sharing mechanism has increased by 15%. However, the prediction effect of the traditional soft sharing mechanism represented by the Cross-Stitch Network is more dependent on the initial value of the feature sharing weight and the selection of some parameters. If the initialization is biased towards certain feature combinations, it may cause the model to over-focus on specific tasks, resulting in only some tasks being significantly improved. For example, the Cross-Stitch Network has achieved a significant improvement in the prediction of electric loads with lower prediction difficulty, but the improvement in the prediction of cold loads and heat loads with stronger coupling is relatively limited, so it lacks certain fine-grained adjustment capabilities and dynamic adaptability of features.
[0105] The multiple expert subnets and gating networks of the CNN-ECA-MMoE-LSTM multi-task learning model of the present invention can dynamically select the most suitable feature combination without relying on fixed initial feature sharing weights, thus avoiding the problem of feature bias and initial weight imbalance. In the prediction of cold load and heat load with stronger coupling and higher complexity, the CNN-ECA-MMoE-LSTM multi-task learning model of the present invention can also dynamically and flexibly call different expert features according to the adjustment of the gating network, so as to more fully capture the interactive relationship between complex loads. Compared with the cross-stitch network, the CNN-ECA-MMoE-LSTM multi-task learning model of the present invention improves the prediction accuracy of cold load and heat load by 40.2% and 38.2%, respectively, achieving satisfactory results.
Claims
1. A multi-load joint forecasting method integrating decomposition and reconstruction technology with multi-task learning is characterized by: The steps include: Step S1, obtaining the historical load sequence of electric load, cooling load and heating load in the integrated energy system, forming a historical multi-load sequence and forming a historical multi-load sequence library; at the same time, obtaining the data sequence of the characteristics of the influencing factors affecting the electric load, cooling load and heating load, forming an influencing factor characteristic library; Step S2, preliminarily decomposing the historical load sequences of electric load, cooling load and heating load in the historical multivariate load sequence by using the variational mode decomposition method, and obtaining K mode components for each load; Step S3, reconstructing the K modal components of each load obtained in step S2 from multiple angles by using a hybrid characteristic evaluation method, and converting the K modal components corresponding to each load into a set of three types of aggregate sequences; then updating the three types of aggregate sequences corresponding to each load into a historical multivariate load sequence library to obtain an updated historical multivariate load sequence library; Step S4, normalizing the influencing factor feature library obtained in step S1 and the updated historical multivariate load sequence library obtained in step S3 with a time step length M as the sequence length, and then merging them to obtain multiple groups of three-dimensional feature tensors and form a three-dimensional feature tensor library, and divide them into a training set and a test set in a ratio of 8:2; Step S5, constructing a CNN-ECA-MMoE multi-task learning model to output the results of feature sharing and feature extraction; The CNN-ECA-MMoE multi-task learning model includes multiple expert subnets and multiple gating networks. The expert subnets include a two-dimensional CNN layer, a pooling layer, and an efficient channel attention mechanism ECA connected in sequence; the expert subnets are used to learn specific feature representations of input data; the gating network is used to dynamically adjust the output weights of each expert subnet according to the input data, and combine the output features of each expert subnet; Step S6, construct three independent long short-term memory networks LSTM, embed them into the back end of the CNN-ECA-MMoE multi-task learning model as the task-specific sub-network of the model, and connect two fully connected layers after each long short-term memory network LSTM to obtain the CNN-ECA-MMoE-LSTM multi-task learning model, and train the model using the training set divided in step S4; The LSTM network is used to capture the temporal dependency of the results of feature sharing and feature extraction; the fully connected layer is used to perform feature transformation on the output of the LSTM network, output the prediction result of the multivariate load, and perform denormalization processing to obtain the prediction value of the same scale as the input data; Step S7, according to the method in step S1 to step S4, obtain the three-dimensional feature tensor corresponding to the time to be predicted; input the three-dimensional feature tensor corresponding to the time to be predicted into the trained CNN-ECA-MMoE-LSTM multi-task learning model for feature extraction and capture of timing dependencies, and obtain the prediction results of electric load, cooling load and heating load.
2. The multi-load joint forecasting method integrating decomposition and reconstruction technology with multi-task learning according to claim 1 is characterized in that: The historical multi-load sequence in step S1 includes a historical electric load sequence, a historical cooling load sequence, and a historical heating load sequence; the influencing factor characteristics in step S1 are meteorological factor characteristics associated with the electric load, cooling load, and heating load, and the meteorological factor characteristics include air temperature, cloud type, dew point, relative humidity, solar zenith angle, surface albedo, pressure, precipitation, wind direction, and wind speed; the meteorological factor characteristics are analyzed by a Pearson correlation analysis method, and the Pearson correlation coefficients between each meteorological factor characteristic and the electric load, cooling load, and heating load are calculated respectively, and the absolute values of the Pearson correlation coefficients between each meteorological factor characteristic and the three loads are taken and then averaged to obtain the correlation coefficients between each meteorological factor characteristic and the three loads, and the correlation coefficients are sorted in descending order, and the meteorological factor characteristics corresponding to the first three groups of correlation coefficients are selected to construct an influencing factor feature library.
3. The multi-load joint forecasting method integrating decomposition and reconstruction technology with multi-task learning according to claim 2 is characterized in that: The specific steps of step S3 are: evaluating the K modal components of each load obtained in step S2 from three perspectives: complex characteristics, coupling characteristics, and frequency characteristics, respectively, calculating the complex characteristics score, coupling characteristics score, and frequency characteristics score of each modal component, and integrating the three scores to obtain a hybrid characteristics evaluation score; then classifying each modal component according to the hybrid characteristics evaluation score, and aggregating the modal components belonging to the same component category corresponding to each load, so as to convert the K modal components corresponding to each load into a group of three types of aggregation sequences, and obtain three types of aggregation sequences corresponding to each load; the three types of aggregation sequences corresponding to each load contain three sequence components, namely, trend component, fluctuation component, and random component; then updating the three types of aggregation sequences corresponding to each load to the historical multivariate load sequence library obtained in step S1, replacing the historical multivariate load sequence before reconstruction, and obtaining an updated historical multivariate load sequence library.
4. The multi-load joint forecasting method integrating decomposition and reconstruction technology with multi-task learning according to claim 3 is characterized in that: In step S3, the specific steps of calculating the complex characteristic score of the modal component are as follows: Step S311, define the modal component as , x1 represents the first data point of the modal component, represents the Nth data point of the modal component, where N is the total number of time points in the historical load series; by embedding dimension m and time delay , rearrange the modal components according to time delay and construct the phase space matrix Y G , the phase space matrix Y G The representation is as follows: , Where G is the number of rows in the matrix, ; The first data point corresponding to the modal component, The corresponding modal component data points, The corresponding modal component data points, similarly, Corresponding to the Gth data point of the modal component, The corresponding modal component data points, The corresponding modal component The data point corresponds to the Nth data point of the modal component; Step S312, in the phase space matrix Y G In the above example, each row of data elements is regarded as an embedded vector. After arranging the data elements in the embedded vector in ascending order, the arrangement pattern corresponding to the embedded vector is obtained. ; Step S313, calculate the phase space matrix Y G All permutation modes The frequencies that appear in the modal components are denoted by , the total number of permutation patterns is recorded as , then based on the frequency distribution, the Shannon entropy formula is used to calculate the permutation entropy of the modal component. The calculation formula is as follows: , Where PE represents the permutation entropy of the modal component, It represents the natural logarithm of frequency and is used to quantify uncertainty or information content; Step S314, normalizing the permutation entropy of the modal component, the normalization formula is as follows: , In the formula, represents the normalized permutation entropy, representing the complex feature score; represents the normalization factor; where the normalized permutation entropy The range of is in [0,1]. Within the range, the larger the value of the normalized permutation entropy is, the more complex the corresponding modal component is, and the smaller the value of the normalized permutation entropy is, the more regular the corresponding modal component is.
5. The multi-load joint forecasting method integrating decomposition and reconstruction technology with multi-task learning according to claim 4 is characterized in that: In step S3, the specific steps of calculating the coupling characteristic score of the modal component are: The grey correlation analysis method is used to calculate the load nonlinear correlation coefficient between the modal components of each type of load and the historical load series of the other two types of loads, and obtain the coupling characteristic scores of the modal components of each type of load; among them, the calculation steps of the coupling characteristic scores of the modal components of the electric load are as follows: Step S321, calculating the i-th modal component of the electric load Pearson correlation coefficient with historical cooling load series , and calculate the i-th modal component of the electric load Pearson correlation coefficient with historical heat load series ; Step S322, using the grey correlation analysis method to calculate the i-th modal component of the electric load Grey correlation coefficient between the historical cooling load sequence , and the i-th modal component of the electric load Grey correlation coefficient between the historical heat load sequence , the calculation formula is as follows: , Where t represents the tth time point in the historical load series, represents the i-th modal component of the electric load The value at the tth time point in represents the value at the tth time point in the historical cooling load series, represents the value at the tth time point in the historical heat load sequence, It is expressed as the resolution coefficient; Indicates the minimum absolute value. Indicates taking the maximum absolute value; Step S323, calculate the i-th modal component of the electric load The coupling characteristic score , the specific calculation formula is as follows: 。 6. The multi-load joint forecasting method integrating decomposition and reconstruction technology with multi-task learning according to claim 5 is characterized in that: In step S3, the specific steps of calculating the frequency characteristic score of the modal component are as follows: Step S331, perform Hilbert transform on each modal component, obtain the analytical time series corresponding to the modal component, and convert the analytical time series into polar coordinate form; the calculation formula of Hilbert transform is: , in, represents the analytical signal of the ith modal component, represents the value of the i-th modal component at the t-th time point, represents the i-th modal component at time point The value at , j represents the imaginary unit, represents the time delay in the integration, Represents the time difference between two consecutive time points. is the instantaneous amplitude of the modal component, represents the complex exponential function, is the instantaneous phase of the modal component; Step S332, based on the instantaneous phase of the modal component Calculate instantaneous frequency , the specific calculation formula is: , in, Instantaneous phase The differential of, dt represents the differential at time point t; Step S333: calculate the instantaneous frequency of the entire time period N of the historical load sequence. Average to get the center frequency of the i-th modal component , the specific calculation formula is: ; Among them, the center frequency represents the frequency characteristic score, Indicates the instantaneous frequency Perform integral operation.
7. The multi-load joint forecasting method integrating decomposition and reconstruction technology with multi-task learning according to claim 6 is characterized in that: In step S3, the calculation formula of the mixed characteristic evaluation score is as follows: , In the formula, is the mixed characteristic evaluation score of the i-th modal component, Score the coupling characteristics; is the weight coefficient of complex characteristics, is the coupling characteristic weight coefficient, is the frequency characteristic weight coefficient.
8. The multi-load joint forecasting method integrating decomposition and reconstruction technology with multi-task learning according to claim 7 is characterized in that: In step S3, the specific method for classifying each modal component according to the mixed characteristic evaluation score is: classifying the modal components with a mixed characteristic evaluation score greater than or equal to 0 and less than 0.4 as trend components, classifying the modal components with a mixed characteristic evaluation score greater than or equal to 0.4 and less than 0.7 as fluctuation components, and classifying the modal components with a mixed characteristic evaluation score greater than or equal to 0.7 and less than 1 as random components.
9. The multi-load joint forecasting method integrating decomposition and reconstruction technology with multi-task learning according to claim 1 is characterized in that: The number of the expert subnetworks is 2 to 6, and the number of the gated networks is 3.
10. The multi-load joint forecasting method integrating decomposition and reconstruction technology with multi-task learning according to claim 1 is characterized in that: The efficient channel attention mechanism ECA captures the channel global information of the input feature map through the global maximum pooling layer GAP, and then adaptively learns the local interaction relationship between channels through a one-dimensional convolutional layer, generates attention weights using the Sigmoid function, and dynamically adjusts the response strength of each channel of the input feature map using the attention weights, thereby obtaining an output feature map containing key features.
Citation Information
Patent Citations
Comprehensive energy load prediction method based on multi-task learning strategy and deep learning
CN113822481A
Comprehensive energy system load prediction method considering multivariate load coupling characteristics
CN111950793A
Power load prediction method based on deep fusion of wavelet transform and multi-task learning
CN118249312A
Comprehensive energy multi-element load short-term prediction method based on secondary decoupling
CN119231529A
Bidirectional long short-term memory network-based power load prediction method and apparatus
WO2025059969A1
Cited By
Load twinborn modeling and predicting method and system based on multi-source heterogeneous feature fusion
CN120728589A
Load twin modeling and prediction method and system based on multi-source heterogeneous feature fusion
CN120728589B
Multi-element load hybrid prediction method, system and equipment of integrated energy system and medium
CN120855331A
Data-driven regional power grid load prediction method and system and medium
CN120879580A
Active load prediction method and device for low-voltage distribution transformer area
CN120955651A