Multivariate Load Joint Prediction Method Combining Decomposition-Reconstruction Technology and Multi-Task Learning
By using decomposition and reconstruction technology and multi-task learning methods to predict multi-energy loads in an integrated energy system, the problem of insufficient load prediction accuracy and stability in the prior art is solved, and higher accuracy and clearer feature representations are achieved.
Patent Information
- Application Number
- CN202510432615.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-04-08
AI Technical Summary
Existing load prediction methods are difficult to fully capture the dynamic coupling information in multi-energy loads in an integrated energy system, resulting in insufficient prediction accuracy and stability.
The combined multi-load prediction method of fusion decomposition and reconstruction technology and multi-task learning is adopted to decompose and reconstruct the load data through variational modal decomposition and mixed characteristic evaluation method, and combined with the CNN-ECA-MMoE multi-task learning model, the model's dynamic coupled feature capture capability of multi-load is improved.
It improves the accuracy and stability of load prediction, reduces the interference of noise on multi-task learning, provides clearer and more accurate feature representation, and enhances the training effect of the model.
Smart Images

Figure CN119965861B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of integrated energy load forecasting, and particularly relates to a multi-source load joint forecasting method that combines decomposition and reconstruction techniques with multi-task learning. Background Technique
[0002] Each subsystem of the traditional energy system operates relatively independently and in a closed manner, lacking a coordination mechanism. This decentralized structure restricts the efficient linkage between different energy forms, resulting in low energy utilization efficiency and difficulty in flexibly responding to the changes in modern energy demands. With the deep integration of multiple energy systems such as electricity, gas, and heat, an integrated energy system (IES) centered around the power system has been formed. This integrated energy system combines multiple energies through systematic and integrated methods, optimizes the production, transmission, storage, and use of energy, and promotes the development of renewable energy, laying a physical foundation for the construction of the energy Internet. Accurate load forecasting is crucial for optimizing energy allocation and ensuring the stable operation of the system. Load forecasting not only helps the energy sector formulate efficient supply and demand plans but also ensures the reasonable allocation and use of multiple energy forms, serving as an important foundation for maintaining the stability of the energy system and improving the overall energy efficiency.
[0003] With the rapid development of computer hardware and artificial intelligence technologies, machine learning, especially deep learning, has shown significant advantages in load forecasting. Deep learning models can automatically extract features from large amounts of data, capture complex non-linear relationships, and adapt to dynamic load changes. For example, the long short-term memory network (LSTM) can effectively capture time dependencies and long-term trends; the convolutional neural network (CNN) is good at extracting local features and spatio-temporal patterns; the Transformer is proficient in global dependency modeling and can efficiently capture long-distance relationships and process sequence data in parallel. These technologies can extract features from large amounts of historical data, capture complex non-linear relationships, and adapt to dynamic load patterns without the need to establish complex mathematical models. However, in an integrated energy system with high coupling and complex dynamic characteristics, these methods often ignore the dynamic information hidden in multi-energy loads, are unable to comprehensively capture the potential complex patterns in load data, and have insufficient ability to identify the fine structure retained after filtering out external disturbances within the load and the mutual relationship between the load and noise, showing obvious limitations in accurately identifying and separating the inherent characteristics of the load itself and being difficult to achieve ideal prediction accuracy and stability. Therefore, in recent years, many scholars have been committed to studying how to effectively capture the dynamic coupling information hidden in multi-energy loads and exploring the use of feature processing technologies to separate effective information to improve the prediction accuracy of the model.
[0004] In the research of feature processing technology in recent years, load decomposition has become one of the widely used technologies. By decomposing complex load time series into multiple modal components, it can effectively remove noise in the data and separate uncertain factors, improving the stationarity and processability of load data. However, the above decomposition-prediction method still faces challenges because a large number of modal components after decomposition will significantly increase the computational complexity of the model, resulting in a decline in computational efficiency and error accumulation. To solve this problem, some researchers have tried to incorporate the decomposition-reconstruction method into the prediction process. The decomposition-reconstruction-prediction strategy effectively reduces the computational complexity of the model while retaining the benefits of decomposition and enhances the ability of the prediction model to identify important patterns. However, the current reconstruction methods often rely on a single feature as a reference and are difficult to comprehensively consider the multi-dimensional information of modal components from multiple perspectives, resulting in insufficient comprehensiveness of the reconstructed data and difficulty in fully exerting the potential of load decomposition technology.
[0005] To improve the model's ability to capture the dynamic coupling characteristics in multi-source loads, some researchers have tried to apply multi-task learning (MTL) to multi-energy load forecasting. Multi-task learning is a typical inductive transfer mechanism. By training multiple related tasks simultaneously, it can make full use of the shared information between tasks, thereby enhancing the model's understanding ability of complex systems. Currently, most studies adopt the hard sharing mechanism in multi-task learning. This method fuses the independent information of sub-tasks through a shared layer and abstracts global features for representation to improve the generalization ability of the model for each task. However, the hard sharing mechanism allows all sub-tasks to share the underlying feature extraction layer, resulting in interference between sub-tasks due to their differences, making it difficult to extract specific information exclusive to each sub-task and lacking the flexibility to adjust the requirements of each task.
[0006] In summary, based on the inherent coupling characteristics of multi-source load sequences, leveraging efficient feature processing technology to enhance data representation ability and combining advanced deep learning networks and multi-task learning strategies are the keys to achieving joint prediction of integrated energy loads. However, the existing decomposition and reconstruction technology, as a feature processing means, does not integrate modal characteristics from multiple perspectives and still has deficiencies in the diversity of feature expression and the exploration of task relevance. At the same time, the existing multi-task learning method based on the parameter hard sharing mechanism is easily interfered by the weak relevance of sub-tasks, further limiting the learning effect of tasks. The patent application with the publication number CN113822481A discloses a comprehensive energy load forecasting method based on multi-task learning strategy and deep learning. It only adopts the structure of the existing multi-task learning model MMoE and does not improve the model. The feature extraction and identification ability of the model is limited. At the same time, it does not process the input data and is easily affected by data noise, resulting in insufficient prediction accuracy.
[0007] Therefore, how to avoid the interference of weakly relevant information to subtasks and achieve better prediction effects is a technical problem that those skilled in the art urgently need to solve. Summary of the Invention
[0008] In view of the deficiencies of existing research, the present invention proposes a multi-source load joint prediction method that combines decomposition and reconstruction technology with multi-task learning. This method combines decomposition and reconstruction technology with multi-task learning strategies. While eliminating redundant information and reducing the interference of noise on multi-task learning, it can provide clearer and more accurate feature representations for different subtasks, thereby achieving higher-precision multi-source load prediction.
[0009] To achieve the above object, the technical solutions adopted by the present invention are as follows.
[0010] A multi-source load joint prediction method that combines decomposition and reconstruction technology with multi-task learning includes the following steps:
[0011] Step S1: Obtain the historical load sequences of electrical load, cooling load, and heating load in the integrated energy system, form a historical multi-source load sequence, and form a historical multi-source load sequence library; at the same time, obtain the data sequences of the influencing factor features that affect the electrical load, cooling load, and heating load, and form an influencing factor feature library;
[0012] Step S2: Perform preliminary decomposition on the historical load sequences of electrical load, cooling load, and heating load in the historical multi-source load sequence respectively through the variational mode decomposition method, and each load obtains K modal components;
[0013] Step S3: Perform multi-angle reconstruction on the K modal components of each load obtained in Step S2 through the hybrid characteristic evaluation method, convert the K modal components corresponding to each load into a set of three-category aggregation sequences; then update the three-category aggregation sequences corresponding to each load to the historical multi-source load sequence library to obtain an updated historical multi-source load sequence library;
[0014] Step S4: Normalize and then merge the influencing factor feature library obtained in Step S1 and the updated historical multi-source load sequence library obtained in Step S3 with a time step M as the sequence length to obtain multiple groups of three-dimensional feature tensors and form a three-dimensional feature tensor library, and divide it into a training set and a test set in a ratio of 8:2;
[0015] Step S5: Construct a CNN-ECA-MMoE multi-task learning model for outputting the results of feature sharing and feature extraction;
[0016] The CNN-ECA-MMoE multi-task learning model includes multiple expert subnets and multiple gating networks. The expert subnets include a two-dimensional CNN layer, a pooling layer, and an efficient channel attention mechanism (ECA) connected in sequence. The expert subnets are used to learn specific feature representations of the input data. The gating networks are used to dynamically adjust the output weights of each expert subnet according to the input data and combine the output features of each expert subnet.
[0017] Step S6: Construct three independent long short-term memory networks (LSTMs), embed them at the backend of the CNN-ECA-MMoE multi-task learning model as task-specific subnets of the model, and connect two fully connected layers after each long short-term memory network (LSTM), so as to obtain the CNN-ECA-MMoE-LSTM multi-task learning model, and use the training set divided in step S4 to train the model.
[0018] The long short-term memory network (LSTM) is used to capture the temporal dependence relationship of the results of feature sharing and feature extraction. The fully connected layer is used to perform feature transformation on the output of the long short-term memory network (LSTM), output the prediction results of multiple loads, and perform inverse normalization processing to obtain prediction values with the same scale as the input data.
[0019] Step S7: Obtain the three-dimensional feature tensor corresponding to the prediction time according to the methods in steps S1 to S4. Input the three-dimensional feature tensor corresponding to the prediction time into the trained CNN-ECA-MMoE-LSTM multi-task learning model in step S6 for feature extraction and capture of the temporal dependence relationship, and obtain the prediction results of electric load, cooling load, and heating load.
[0020] Further, the historical multi-load sequence in step S1 includes a historical electric load sequence, a historical cooling load sequence, and a historical heating load sequence. The influencing factor features in step S1 are meteorological factor features associated with electric load, cooling load, and heating load. The meteorological factor features include temperature, cloud type, dew point, relative humidity, solar zenith angle, surface albedo, pressure, precipitation, wind direction, and wind speed. Analyze the meteorological factor features by the Pearson correlation analysis method, calculate the Pearson correlation coefficients between each meteorological factor feature and electric load, cooling load, and heating load respectively, take the absolute value of the Pearson correlation coefficients between each meteorological factor feature and the three loads and then calculate the average value to obtain the correlation degree coefficients between each meteorological factor feature and the three loads. Sort the correlation degree coefficients in descending order, and select the meteorological factor features corresponding to the top three groups of correlation degree coefficients to construct the influencing factor feature library.
[0021] Further, the meteorological factor features corresponding to the top three groups of correlation degree coefficients are temperature, precipitation, and dew point.
[0022] Further, the specific steps of step S3 are as follows: evaluate the K modal components of each load obtained in step S2 from three perspectives: complex characteristics, coupling characteristics, and frequency characteristics, calculate the complex characteristic score, coupling characteristic score, and frequency characteristic score of each modal component respectively, and integrate the three scores to obtain a hybrid characteristic evaluation score; then classify each modal component according to the hybrid characteristic evaluation score, and aggregate the modal components belonging to the same component category corresponding to each load, so as to transform the K modal components corresponding to each load into a set of three-category aggregation sequences, and obtain the three-category aggregation sequences corresponding to each load; each three-category aggregation sequence corresponding to each load contains three sequence components, namely a trend component, a fluctuation component, and a random component; then update the three-category aggregation sequences corresponding to each load to the historical multi-load sequence library obtained in step S1, replacing the historical multi-load sequence before reconstruction, to obtain an updated historical multi-load sequence library.
[0023] Further, in step S3, the specific method for classifying each modal component according to the hybrid characteristic evaluation score is as follows: classify the modal components with a hybrid characteristic evaluation score greater than or equal to 0 and less than 0.4 as trend components, classify the modal components with a hybrid characteristic evaluation score greater than or equal to 0.4 and less than 0.7 as fluctuation components, and classify the modal components with a hybrid characteristic evaluation score greater than or equal to 0.7 and less than 1 as random components.
[0024] Further, the number of expert subnets is 2 to 6, and the number of gating networks is 3.
[0025] Further, the efficient channel attention mechanism ECA captures the channel global information of the input feature map through the global maximum pooling layer GAP, then adaptively learns the local interaction relationship between channels through a one-dimensional convolutional layer, generates attention weights using the Sigmoid function, and dynamically adjusts the response intensity of each channel of the input feature map using the attention weights, so as to obtain an output feature map containing key features.
[0026] Advantages and beneficial effects of the present invention:
[0027] (1) The present invention comprehensively evaluates the complex characteristics, coupling characteristics, and frequency characteristics of each modal component, introduces the hybrid characteristic evaluation method (HFEM), comprehensively considers the information of modal components from multiple perspectives, reconstructs each modal component to obtain a trend component, a fluctuation component, and a random component, and transforms the modal components into three-category aggregation sequences. While eliminating redundant information and reducing noise, it further clarifies the representation of the effective information of the sequence, thereby reducing the interference caused by task differences in multi-task learning.
[0028] (2)The proposed method for joint prediction of multiple loads by integrating decomposition and reconstruction technology and multi-task learning in the present invention takes into account the coupling characteristics of multiple loads, can utilize the correlation between multiple loads to improve the training effect of the model, and enhance the load prediction accuracy.
[0029] (3)The present invention breaks through the limitations of the inherent multi-task learning with a hard sharing mechanism, improves the multi-task learning model MMoE, integrates a two-dimensional CNN layer and an efficient channel attention mechanism ECA to form a CNN-ECA module to replace the expert subnet composed of fully connected layers in the multi-task learning model MMoE, and forms an improved MMoE multi-task learning model. By combining the high-dimensional feature extraction ability of the two-dimensional CNN layer with the flexible task allocation mechanism of the multi-task learning model MMoE and introducing the efficient channel attention mechanism ECA, the inductive ability of the inherent differential features of the expert subnet is enhanced. Therefore, the model proposed in the present invention can consider the differences in the correlation degrees between multiple loads, ensure that each task can obtain the most effective information, and achieve a better prediction effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 It is a flowchart of the method for joint prediction of multiple loads by integrating decomposition and reconstruction technology and multi-task learning provided by an embodiment of the present invention;
[0031] Figure 2 It is a graph showing the correlation analysis results of multiple loads and meteorological factor features provided by an embodiment of the present invention;
[0032] Figure 3 It is three types of aggregated sequences after the reconstruction of the electrical load provided by an embodiment of the present invention;
[0033] Figure 4 It is three types of aggregated sequences after the reconstruction of the cooling load provided by an embodiment of the present invention;
[0034] Figure 5 It is three types of aggregated sequences after the reconstruction of the heating load provided by an embodiment of the present invention;
[0035] Figure 6 It is a schematic structural diagram of the multi-task learning model MMoE provided by an embodiment of the present invention;
[0036] Figure 7 It is a schematic structural diagram of the efficient channel attention mechanism ECA provided by an embodiment of the present invention;
[0037] Figure 8 It is a schematic structural diagram of the CNN-ECA-MMoE-LSTM multi-task learning model provided by an embodiment of the present invention;
[0038] Figure 9The predicted results of the electrical load under different reconstruction strategies provided by the embodiments of the present invention;
[0039] Figure 10 The predicted results of the cooling load under different reconstruction strategies provided by the embodiments of the present invention;
[0040] Figure 11 The predicted results of the heating load under different reconstruction strategies provided by the embodiments of the present invention;
[0041] Figure 12 The predicted results of the electrical load under the ablation comparison experiment provided by the embodiments of the present invention;
[0042] Figure 13 The predicted results of the cooling load under the ablation comparison experiment provided by the embodiments of the present invention;
[0043] Figure 14 The predicted results of the heating load under the ablation comparison experiment provided by the embodiments of the present invention;
[0044] Figure 15 The predicted results of the multi - load under different sharing mechanisms provided by the embodiments of the present invention. Detailed implementation manners
[0045] The following further details the embodiments of the present invention with reference to the accompanying drawings.
[0046] As Figure 1 shown, a multi - load joint prediction method integrating decomposition and reconstruction technology and multi - task learning of the present invention includes the following steps:
[0047] Step S1, obtain the historical load sequences of the electrical load, cooling load, and heating load in the integrated energy system, form a historical multi - load sequence and a historical multi - load sequence library; at the same time, obtain the data sequences of the influencing factor characteristics that affect the electrical load, cooling load, and heating load in the integrated energy system, and form an influencing factor characteristic library;
[0048] The historical multi - load sequences in the historical multi - load sequence library in step S1 include the historical electric load sequence, historical cooling load sequence, and historical heating load sequence for the M time steps before each moment to be predicted; the influencing factor features are meteorological factor features that are highly correlated with electric load, cooling load, and heating load. The meteorological factor features include air temperature, cloud type, dew point, relative humidity, solar zenith angle, surface albedo, pressure, precipitation, wind direction, and wind speed. The meteorological factor features are analyzed by the Pearson correlation analysis method, and the Pearson correlation coefficients between air temperature, cloud type, dew point, relative humidity, solar zenith angle, surface albedo, pressure, precipitation, wind direction, wind speed and electric load, cooling load, and heating load are calculated respectively. The average value of the absolute values of the Pearson correlation coefficients between each meteorological factor feature and the three types of loads is obtained to get the correlation degree coefficient between each meteorological factor feature and the three types of loads. The correlation degree coefficients are sorted in descending order (i.e., from large to small), and the meteorological factor features corresponding to the top three groups of correlation degree coefficients are selected to construct the influencing factor feature library. In this embodiment, the meteorological factor features corresponding to the top three groups of correlation degree coefficients are air temperature, precipitation, and dew point.
[0049] The energy - using behaviors of users in the integrated energy system are related to multiple factors, mainly including meteorological factors, economic factors, social factors, etc. Considering the ease of data acquisition, only meteorological factors are selected for variable correlation analysis in this embodiment. The Pearson correlation coefficient is used to analyze the internal correlation of the electric, cooling, and heating load sequences in the integrated energy system and the correlation between the electric, cooling, and heating load sequences in the integrated energy system and meteorological factor features. The specific calculation formula of the Pearson correlation coefficient is:
[0050] ,
[0051] where, X n represents the data of the feature variable X at the nth time point. The feature variable X is electric load, cooling load, or heating load. Y n represents the data of the feature variable Y at the nth time point. The feature variable Y is electric load, cooling load, heating load, air temperature, cloud type, dew point, relative humidity, solar zenith angle, surface albedo, pressure, precipitation, wind direction, or wind speed; represents the mean value of the data of the feature variable X at all time points, represents the mean value of the data of the feature variable Y at all time points, r represents the Pearson correlation coefficient, and the value range of r is [-1, 1]. When the value of r is -1, it means that the feature variable X and the feature variable Y are completely negatively correlated. When the value of r is +1, it means that the feature variable X and the feature variable Y are completely positively correlated. When the value of r is 0, it means that there is no correlation between the feature variable X and the feature variable Y; N represents the total number of time points in the historical load sequence, and n represents the nth time point in the historical load sequence.
[0052] This embodiment uses the data from January 1, 2022 to December 31, 2022 in the integrated energy system (IES) dataset of Arizona State University in Tempe, Arizona, USA for simulation. The data is provided at 1-hour intervals and includes data on electrical load, cooling load, heating load, and meteorological factor characteristic data. The meteorological factor characteristic data is obtained from the weather station at Phoenix International Airport, which is the closest to the Tempe campus. The collected meteorological factor characteristics include air temperature, cloud type, dew point, relative humidity, solar zenith angle, surface albedo, pressure, precipitation, wind direction, and wind speed. Calculate the Pearson correlation coefficients between the data of electrical load, cooling load, heating load, and meteorological factor characteristic data, and obtain the correlation analysis result graph as Figure 2 shown.
[0053] Figure 2 The correlation analysis results shown indicate that any two of the three loads of electrical load, cooling load, and heating load show a significant correlation; there is a strong mutual relationship between air temperature and the three loads; precipitation is strongly correlated with electrical load and cooling load, and moderately correlated with heating load; dew point and pressure show a moderate correlation with the three loads. This embodiment considers the strength of the correlation between each meteorological factor characteristic and the three loads, calculates the average value of the absolute values of the Pearson correlation coefficients between each meteorological factor characteristic and the three loads, obtains the correlation degree coefficients between each meteorological factor characteristic and the three load sequences, and sorts the correlation degree coefficients between all meteorological factor characteristics and the three loads in descending order. The correlation degree coefficients from high to low are air temperature, precipitation, dew point, pressure, relative humidity, surface albedo, solar zenith angle, wind direction, wind speed, and cloud type. In prediction, considering more feature variables does not necessarily mean higher accuracy. Inputting feature variables with a low degree of correlation will increase the computational complexity and deteriorate the training effect of the model. Based on this consideration, this embodiment only selects the meteorological factor characteristics that are strongly correlated with the multi-load sequence: air temperature, precipitation, and dew point.
[0054] The influence factor feature library in step S1 includes the meteorological factor characteristics corresponding to the previous M time steps before the prediction time, and the meteorological factor characteristics do not need to be decomposed and reconstructed. Due to the high correlation between the same time of adjacent days, considering the daily periodicity, the value of the interval time step M for the prediction time is selected as 24 (unit: hour).
[0055] Step S2, preliminary decomposition based on the variational mode decomposition (VMD) method. The historical load sequences of electrical load, cooling load, and heating load (i.e., historical electrical load sequence, historical cooling load sequence, and historical heating load sequence) in the historical multi-load sequence library obtained in step S1 are respectively preliminarily decomposed by the variational mode decomposition (VMD) method, and each load obtains K modal components (IMFs).
[0056] Variational mode decomposition (VMD) is an adaptive sequence decomposition method. By means of Hilbert transform and minimum frequency bandwidth, it solves the problem that empirical mode decomposition (EMD) is vulnerable to noise interference and causes mode aliasing. The variational mode decomposition method adaptively decomposes the sequence into multiple discrete subsequences, reducing the complexity of the sequence to a certain extent. The solution process of variational mode decomposition is as follows:
[0057] ,
[0058] where, represents the Lagrangian function corresponding to the solution problem, represents the k-th modal component, represents the central frequency of the k-th modal component, v is the penalty factor, K represents the total number of set modal components, represents the partial derivative with respect to time, represents the Dirac function, j represents the imaginary unit, t represents the t-th time point in the historical load sequence, represents the complex exponential function, f(t) is the historical load sequence, represents the value of the Lagrange multiplier at the t-th time point;
[0059] After aggregating the K modal components of each load, the modal aggregation sequence of the corresponding load is formed. Subtracting the modal aggregation sequence from the historical load sequence of the corresponding load can obtain the residual component (Res). The calculation formula of the residual component is: ,
[0060] where, R(t) is the residual component obtained by variational mode decomposition, K is the total number of set modal components, f(t) is the historical load sequence, represents the k-th modal component.
[0061] The residual component (Res) is the part of the historical load sequence that cannot be further decomposed, which contains a large amount of noise and is not included in the data processing of subsequent prediction tasks, but can be used to evaluate the stability of the division of the number of modal components; the smaller the residual component, the more sufficient the effective information retained after removing noise during the initial decomposition. The selection of the total number K of modal components directly affects the quality of the decomposition result and the performance of subsequent tasks. In this embodiment, the mean square error is used as an evaluation index to determine the value of the total number K of modal components, so as to obtain K modal components with stable and periodic trends. Let the total number K of modal components be a positive integer between [K min , K max , and calculate in [K min , K maxAmong all the values between, the mean square error between the historical load sequence and the modal aggregation sequence is taken, and the value corresponding to the minimum mean square error is set as the total number K of modal components. In this embodiment, K min = 1, K max = 10. The calculation formula for the mean square error is: ,
[0062] where R MSE represents the mean square error between the historical load sequence and the modal aggregation sequence, R(t) is the residual component obtained by variational mode decomposition, N represents the total number of time points in the historical load sequence, and t represents the t-th time point in the historical load sequence.
[0063] Tests and experiments of the present invention show that when using the variational mode decomposition method to decompose the historical electric load sequence, historical cooling load sequence, and historical heating load sequence, when the value of the number K of modal components is 8, the mean square error between the modal aggregation sequence and the historical multi-load sequence is the smallest.
[0064] Step S3, based on the reconstruction strategy of the Hybrid Feature Evaluation Method (HFEM), multi-angle reconstruction is performed on the K modal components of each type of load obtained in step S2 through the Hybrid Feature Evaluation Method (HFEM), and the K modal components corresponding to each type of load are respectively transformed into a set of three-category aggregation sequences; then the three-category aggregation sequences corresponding to each type of load are updated to the historical multi-load sequence library obtained in step S1 to obtain an updated historical multi-load sequence library;
[0065] The specific steps of step S3 are as follows: The K modal components of each type of load obtained in step S2 are evaluated from three angles of complex characteristics, coupling characteristics, and frequency characteristics respectively, the complex characteristic scores, coupling characteristic scores, and frequency characteristic scores of each modal component are calculated respectively, and the three scores are integrated to obtain the hybrid feature evaluation score (i.e., the HFEM score); then, taking the hybrid feature evaluation score as the key index for classification and reconstruction, each modal component is classified into a trend component, a fluctuation component, and a random component, and the modal components belonging to the same component category corresponding to each type of load are aggregated, so as to obtain a reconstruction sequence considering multi-modal characteristics, that is, three-category aggregation sequences corresponding to each type of load are obtained, and each three-category aggregation sequence corresponding to each type of load contains a trend component, a fluctuation component, and a random component; then the three-category aggregation sequences corresponding to each type of load are updated to the historical multi-load sequence library obtained in step S1, replacing the historical electric load sequence, historical cooling load sequence, and historical heating load sequence in the historical multi-load sequence before reconstruction, to obtain an updated historical multi-load sequence library.
[0066] This embodiment adopts a reconstruction strategy based on the Hybrid Feature Evaluation Method (HFEM), which considers the characteristics of the load sequence from multiple perspectives. The modal components preliminarily decomposed in step S2 are reconstructed from three aspects: complex characteristics, coupling characteristics, and frequency characteristics, aiming to improve the reconstruction quality of the load sequence and enable the reconstructed load sequence to more accurately reflect the physical characteristics and internal patterns of the original load sequence. The specific steps are as follows:
[0067] Step S31, calculate the complex characteristic score of the modal component:
[0068] The historical load sequences of each type of load in the historical multi-source load sequence library of the integrated energy system are characterized by high non-stationarity, non-linearity, and complexity. Among them, complexity can help distinguish useful load sequences from random components, so it is extremely important for the reconstruction of load sequences. High complexity means that the load sequence has higher uncertainty and disorder, with rich detailed fluctuations, while low complexity means that the load sequence is more regular and stable, which may reflect the key patterns or main trends in the load sequence. For the measurement of complex characteristics, this embodiment uses sample entropy for calculation, and the specific method is as follows:
[0069] Step S311, define the modal component as , x 1 represents the first data point of the modal component, x N represents the Nth data point of the modal component, where N is the total number of time points in the historical load sequence; through the embedding dimension m and the time delay , the modal component is rearranged according to the time delay to construct the phase space matrix Y G , and the representation form of the phase space matrix Y G is as follows:
[0070] ,
[0071] where m is the embedding dimension, is the time delay, G is the number of rows of the matrix, ; corresponds to the first data point of the modal component, corresponds to the th data point of the modal component, corresponds to the th data point of the modal component, and similarly, corresponds to the Gth data point of the modal component, corresponds to the th data point of the modal component, corresponds to the th data point of the modal component, that is, it corresponds to the Nth data point of the modal component. This step ensures that the embedding vectors extracted from the modal component can appropriately reflect the dynamic behavior of the modal component.
[0072] Step S312, in the phase space matrix Y G each row of data elements is regarded as an embedding vector. After arranging the data elements in the embedding vector in ascending order, the corresponding arrangement pattern of the embedding vector is obtained .
[0073] Step S313, calculate the frequency of occurrence of all possible arrangement patterns G in the modal component in the phase space matrix Y, denoted as , the total number of arrangement patterns is denoted as , and then based on the frequency distribution, use the Shannon entropy formula to calculate the permutation entropy of the modal component. The calculation formula is: , ,
[0074] where PE represents the permutation entropy of the modal component, represents the frequency of occurrence of the arrangement pattern in the modal component, represents taking the natural logarithm of the frequency, which is used to quantify uncertainty or information content. This step quantifies the complexity of the sequence and can reflect the degree of disorder of the local structure in the sequence.
[0075] Step S314, perform normalization processing on the permutation entropy of the modal component. The formula for normalization processing is expressed as: ,
[0076] where represents the normalized permutation entropy, representing the complex characteristic score; represents the normalization factor; among them, the normalized permutation entropy ranges from [0, 1]. The larger the value of the normalized permutation entropy within the range, the more complex the corresponding modal component is, and the smaller the value of the normalized permutation entropy, the more regular the corresponding modal component is.
[0077] Step S32, calculate the coupling characteristic score of the modal component:
[0078] In this embodiment, the linear and non-linear correlation relationships between the modal components of various loads and the historical load sequences of the other two loads are considered as the basis for coupling characteristic analysis. On the basis of calculating the Pearson correlation coefficient, the grey relational analysis method (GRA) is used to calculate the load non-linear correlation coefficient between the modal components of various loads and the historical load sequences of the other two loads, and the coupling characteristic scores of the modal components of various loads are obtained; taking the electrical load as an example, the calculation steps of the coupling characteristic score of the modal component of the electrical load are as follows:
[0079] Step S321, calculate the i-th modal component of the electrical load Pearson correlation coefficient between the historical cooling load series Meanwhile, calculate the i-th modal component of the electrical load Pearson correlation coefficient between the historical heating load series .
[0080] Step S322: Use the grey relational analysis method (GRA) to calculate the grey relational degree coefficients between the i-th modal component of the electrical load and the historical cooling load series , and the grey relational degree coefficients between the i-th modal component of the electrical load and the historical heating load series . Grey relational degree is used to measure the development trend and similarity between multiple systems or multiple factors, so it is more suitable for analyzing non-stationary sequences. The calculation formula of the grey relational degree coefficient is as follows:
[0081] ,
[0082] where represents the grey relational degree coefficient between the i-th modal component of the electrical load and the historical cooling load series, represents the grey relational degree coefficient between the i-th modal component of the electrical load and the historical heating load series; N represents the total number of time points in the historical load series, t represents the t-th time point in the historical load series, represents the value of the i-th modal component of the electrical load at the t-th time point, represents the value of the historical cooling load series at the t-th time point, represents the value of the historical heating load series at the t-th time point, represents the discrimination coefficient. In this embodiment, the discrimination coefficient takes a value of 0.5; represents taking the minimum absolute value, represents taking the maximum absolute value.
[0083] Step S323: For the Pearson correlation coefficient between the i-th modal component of the electrical load and the historical cooling load series obtained in Step S321 and Step S322 , the Pearson correlation coefficient between the i-th modal component of the electrical load and the historical heating load series , the grey relational degree coefficient between the i-th modal component of the electrical load and the historical cooling load series , the grey relational degree coefficient between the i-th modal component of the electrical load Grey correlation degree coefficient with the historical heat load sequence The average value of these four indicators is taken to calculate the i-th modal component of the electrical load Coupling characteristic score of . The specific calculation formula is as follows:
[0084] S e , i = [ 1 − 1 4 ( ρ e , c , i + ρ e , h , i + γ e , c o o l , i + γ e , h e a t , i ) ] .
[0085] The calculation method of the coupling characteristic score of the modal component of the cooling load is the same as that of the modal component of the heat load
[0086] Step S33, calculate the frequency characteristic score of the modal component:
[0087] The frequency characteristics of the load sequence directly reflect the changes of the load sequence at different time scales. Only analyzing the complex characteristics and coupling characteristics may not be able to capture the time resolution of the load sequence. The high-frequency components in the load sequence may represent emergencies, noises or short-term fluctuations, and the low-frequency components in the load sequence represent the long-term trend of the load sequence. In order to consider the frequency magnitude of the modal component as a whole, the center frequency is used as the key index in this embodiment, and the Hilbert transform widely used in analyzing non-stationary load sequences is used for calculation. The specific steps are as follows:
[0088] Step S331, perform the Hilbert transform on each modal component to obtain the corresponding analytical time series of the modal component, and convert the analytical time series into polar coordinate form. The calculation formula of the Hilbert transform is:
[0089] ,
[0090] where represents the analytical signal of the i-th modal component, represents the value of the i-th modal component at the t-th time point, represents the value of the i-th modal component at the time point , j represents the imaginary unit, represents the time delay in the integral, represents the time difference between two consecutive time points, is the instantaneous amplitude of the modal component, represents the complex exponential function, is the instantaneous phase of the modal component
[0091] Step S332, calculate the instantaneous frequency according to the instantaneous phase of the modal component, and the specific calculation formula is:
[0092] ,
[0093] Among them, represents the differential of the instantaneous phase , and dt represents the differential with respect to the time point t.
[0094] Step S333: Average the instantaneous frequencies within the entire time period N in which the historical load sequence is located to obtain the center frequency of the i-th modal component. The specific calculation formula is:
[0095] ;
[0096] Among them, the center frequency represents the frequency characteristic score, N represents the total number of time points in the historical load sequence, represents the integral operation of the instantaneous frequency , and dt represents the differential with respect to the time point t.
[0097] Step S34: Calculate the hybrid characteristic evaluation score:
[0098] Calculate the hybrid characteristic evaluation score (i.e., HFEM score) of each modal component according to the complex characteristic score, coupling characteristic score, and frequency characteristic score of each modal component calculated in Steps S31 to S33. The calculation formula for the hybrid characteristic evaluation score is as follows:
[0099] ,
[0100] In the formula, is the hybrid characteristic evaluation score of the i-th modal component, is the complex characteristic score, is the coupling characteristic score, is the frequency characteristic score; is the complex characteristic weight coefficient, is the coupling characteristic weight coefficient, is the frequency characteristic weight coefficient; in this embodiment, .
[0101] Step S35: Reconstruct the modal component. The specific steps are as follows:
[0102] Classify each modal component according to the mixed characteristic evaluation scores of each modal component obtained in step S34. Classify the modal components with mixed characteristic evaluation scores in the interval [0, 0.4) (i.e., greater than or equal to 0 and less than 0.4) as trend components, classify the modal components with mixed characteristic evaluation scores in the interval [0.4, 0.7) (i.e., greater than or equal to 0.4 and less than 0.7) as fluctuation components, classify the modal components with mixed characteristic evaluation scores in the interval [0.7, 1) (i.e., greater than or equal to 0.7 and less than 1) as random components, and aggregate the modal components belonging to the same component category corresponding to each load (i.e., add the modal components) to respectively convert the K modal components corresponding to each load into a set of three-category aggregation sequences. Each of the three-category aggregation sequences corresponding to each load contains three sequence components, namely trend components, fluctuation components, and random components.
[0103] In this embodiment, the data from January 1, 2022 to December 31, 2022 in the integrated energy system (IES) dataset of Arizona State University, Tempe campus, USA is used for testing experiments. Taking the electrical load as an example, the mixed characteristic evaluation scores of the modal components 1 (IMF1) to modal components 8 (IMF8) of the electrical load are calculated, and the results are shown in Table 1 below.
[0104] Table 1 Calculation results of the mixed characteristic evaluation scores of the modal components of the electrical load:
[0105]
[0106] According to steps S31 to S35, the reconstruction of the historical multi-load sequence is completed, and the three-category aggregation sequences of each load are obtained respectively. Then, the historical load sequences of each load in the historical multi-load sequence library in step S1 are updated to the three-category aggregation sequences corresponding to each type of load, and the updated historical multi-load sequence library is obtained.
[0107] The three-category aggregation sequences after the reconstruction of the historical electrical load sequence obtained from the testing experiment in this embodiment are as Figure 3 shown, the three-category aggregation sequences after the reconstruction of the historical cooling load sequence are as Figure 4 shown, and the three-category aggregation sequences after the reconstruction of the historical heating load sequence are as Figure 5 shown.
[0108] Step S4, data normalization and reshaping. Normalize and then merge the influencing factor feature library obtained in step S1 and the updated historical multi-load sequence library obtained in step S3 with the sequence length of the time step M to obtain multiple groups of three-dimensional feature tensors and form a three-dimensional feature tensor library; then divide the data in the three-dimensional feature tensor library into a training set and a test set for model training and testing.
[0109] The specific process of step S4 is as follows: First, normalize the sequence components in the three types of aggregated sequences in the updated historical multi-load sequence library according to the component category, and sequentially extract data with a sequence length of M (corresponding to the time step) from each normalized sequence component to construct a historical multi-load feature matrix, obtaining multiple historical multi-load feature matrices composed of multiple sequences with a length of M. The dimension of the historical multi-load feature matrix is , , , where H r represents the number of reconstructed components, with a value of 3; M represents the time step, with a value of 24 and the unit of hour; represents the number of types of multi-loads, with a value of 3; Then, normalize the influencing factor features in the influencing factor feature library in the same way, and obtain multiple influencing factor feature matrices composed of multiple sequences with a length of M. The dimension of the influencing factor feature matrix is , , , where, represents the number of variables of the influencing factor, with a value of 3; represents the added dimension for matching the data format of the historical multi-load feature matrix, with a value of 1; Finally, merge each historical multi-load feature matrix and its corresponding influencing factor feature matrix in the third dimension to obtain a group of three-dimensional feature tensors with a dimension of , , . Combine all the three-dimensional feature tensors constructed according to the data in the updated historical multi-load sequence library to form a three-dimensional feature tensor library; then divide the data in the three-dimensional feature tensor library into a training set and a test set according to a ratio of 8:2 for model training and testing.
[0110] Step S5, construct a CNN-ECA-MMoE multi-task learning model for outputting the results of feature sharing and feature extraction.
[0111] The multi-task learning model MMoE consists of multiple expert subnets, multiple gating networks (Gate) composed of trainable parameters, and multiple task-specific subnets. The specific structure is as Figure 6 shown. Among them, the role of the expert subnet is to be responsible for learning the specific feature representation of the input data and providing shared feature extraction capabilities for different tasks; the role of the gating network is to dynamically adjust the output weights of each expert subnet according to the input data to adapt to the requirements of different tasks; the role of the task-specific subnet is to optimize for a specific task and convert the output of the expert subnet into a task-specific prediction result.
[0112] The CNN-ECA-MMoE multi-task learning model constructed in the present invention replaces the only fully connected layer inside the expert subnet in the multi-task learning model MMoE with a two-dimensional CNN layer (i.e., a two-dimensional convolutional layer) to extract the spatial local features of the input data; and a pooling layer and an efficient channel attention mechanism ECA are sequentially connected after the two-dimensional CNN layer to form a CNN-ECA module as the expert subnet of the multi-task learning model, which is used to enhance the feature attention ability of the multi-task learning model MMoE; at the same time, combined with the gating network in the multi-task learning model MMoE, an improved MMoE multi-task learning model based on CNN and ECA is formed, named the CNN-ECA-MMoE multi-task learning model. That is, the CNN-ECA-MMoE multi-task learning model constructed in the present invention includes multiple expert subnets and multiple gating networks, and the expert subnet includes a two-dimensional CNN layer, a pooling layer, and an efficient channel attention mechanism ECA connected in sequence.
[0113] The introduction of the efficient channel attention mechanism ECA aims to capture the dependencies between adjacent channels while strengthening the feature selection ability between channels, making the expert subnet in the multi-task learning model MMoE more targeted and accurate when processing features. In addition, the incorporation of the efficient channel attention mechanism ECA does not change the dimension of the shared feature data output by the multi-task learning model, so it has high adaptability. Fusing the efficient channel attention mechanism ECA can strengthen the multi-task learning model and enhance the ability to generalize different features.
[0114] The channel attention mechanism has been proven to have great potential in improving the performance of deep convolutional neural networks. Compared with the widely used squeeze-and-excitation network (SENet), the efficient channel attention mechanism ECA can bring obvious performance improvement to the model while adding a small number of parameters. Due to its lightweight design, it is feasible to use in the multi-task learning model MMoE with parallelization and a certain computational complexity. Compared with the global channel interaction of the squeeze-and-excitation network (SENet), the efficient channel attention mechanism ECA adopts a local cross-channel interaction strategy based on one-dimensional convolution, which can capture the dependencies between adjacent channels and adaptively select an appropriate convolutional kernel size to ensure strong adaptability in different depths of the network hierarchy, so as to selectively emphasize effective features and suppress useless features. Based on this feature, the efficient channel attention mechanism ECA has certain similarity with the expert subnet in the multi-task learning model MMoE in realizing the selective processing of information. Therefore, in this embodiment, the efficient channel attention mechanism ECA is embedded in the expert subnet of the multi-task learning model MMoE to further enhance its inherent feature generalization ability. As Figure 7As shown, the efficient channel attention mechanism ECA captures the global channel information of the input feature map through the global max pooling layer GAP, then adaptively learns the local interaction relationships between channels through a one-dimensional convolutional layer, generates attention weights using the Sigmoid function, and dynamically adjusts the response intensity of each channel of the input feature map using the attention weights, so as to significantly improve the model's expression ability for key features with lightweight calculations and obtain an output feature map containing key features. The specific calculation expression of the efficient channel attention mechanism ECA is as follows:
[0115] ,
[0116] In the formula, represents the output channel descriptor after being processed by the global max pooling layer, represents the input feature map, W represents the width of the input feature map, H represents the height of the input feature map, and c represents the c-th channel; represents the input feature map at the position (o, q) and the value on the c-th channel; represents the weight vector of the c-th channel, represents the Sigmoid function, represents a one-dimensional convolutional operation used to calculate the channel weights, represents the convolutional kernel size; represents the output feature map weighted by the efficient channel attention mechanism ECA, represents the output feature map at the position (o, q) and the value on the c-th channel; represents a function for calculating the convolutional kernel size, C represents the number of channels of the input feature map, represents the logarithmic function with base 2, represents taking the odd number of the result, is the scaling ratio coefficient of the convolutional kernel size. In this embodiment, the scaling ratio coefficient of the convolutional kernel size is set to 2, b is the bias term coefficient, and in this embodiment, the bias term coefficient b is set to 1.
[0117] Step S6, construct a CNN-ECA-MMoE-LSTM multi-task learning model; the specific steps are as follows: construct three independent long short-term memory networks LSTM, embed them into the back end of the CNN-ECA-MMoE multi-task learning model as the task-specific sub-networks of the model, and connect two fully connected layers after each long short-term memory network LSTM, so as to obtain a CNN-ECA-MMoE-LSTM multi-task learning model, and use the training set divided in step S4 to train the model.
[0118] To enable the long short-term memory network (LSTM) to fully utilize the spatial features extracted by the two-dimensional CNN layer and the efficient channel attention mechanism (ECA) in the CNN-ECA-MMoE multi-task learning model, the present invention does not flatten the output of the CNN-ECA-MMoE multi-task learning model or add a fully connected layer. Instead, it maximally retains the original structure of the time steps, multiplying the dimension representing the number of feature maps by the height of the feature map as the feature dimension of the time steps. Therefore, the dimension of the data input to the long short-term memory network (LSTM) is , where represents the number of feature maps output after the CNN-ECA-MMoE multi-task learning model performs feature sharing and feature extraction, represents the height of the feature map input to the long short-term memory network (LSTM), represents the width of the feature map input to the long short-term memory network (LSTM); this detailed processing can effectively maintain the integrity of the spatial information extracted by the two-dimensional CNN layer and avoid the loss of time step information caused by flattening.
[0119] Such as Figure 8As shown, the CNN-ECA-MMoE-LSTM multi-task learning model includes an input layer, a CNN-ECA-MMoE multi-task learning model, task-specific sub-networks, and two fully-connected layers; among them, the number of expert sub-networks in the CNN-ECA-MMoE multi-task learning model is 2 to 6; the number of two-dimensional CNN layers in each expert sub-network is 1 to 3. The specific working process of the CNN-ECA-MMoE-LSTM multi-task learning model is as follows: The three-dimensional feature tensor (dimension: height × width × number of channels = 3 × 24 × 4, where the height corresponds to the number of reconstructed components, the width corresponds to the time step M, and the number of channels corresponds to the number of feature maps) is used as input data and input into the CNN-ECA-MMoE-LSTM multi-task learning model. Multiple expert sub-networks in the CNN-ECA-MMoE multi-task learning model simultaneously obtain the input data and perform feature extraction and induction. In each expert sub-network, the two-dimensional CNN layer extracts the spatial local features of the input data and condenses the common features of multiple channels to obtain a set of tensors with a dimension of 4 × 25 × 32. Then, the pooling layer is used to extract the significant features of the tensors and reduce the computational complexity to obtain tensors with a dimension of 3 × 25 × 32. Finally, the efficient channel attention mechanism ECA further enhances the feature attention ability of the expert sub-network to form a unique expression of the expert sub-network. At the same time, the gating network dynamically adjusts the output weights of each expert sub-network and flexibly combines the output features of each expert sub-network to obtain three results of feature sharing and feature extraction (i.e., a feature tensor with a dimension of 25 × 96); subsequently, three long short-term memory networks LSTM respectively capture the temporal dependence relationships of each result of feature sharing and feature extraction. Finally, the two fully-connected layers perform feature transformation on the output of the long short-term memory network LSTM, output the prediction results of multiple loads, and perform anti-normalization processing to obtain prediction values with the same data scale as the input data.
[0120] The selection of hyperparameters of the CNN-ECA-MMoE-LSTM multi-task learning model has a crucial impact on performance. In this embodiment, the control variable method is used to find the optimal hyperparameters. First, the number of epochs, batch size, and learning rate are set to fixed values, and the rest of the model is kept at a fixed number of layers. First, the number of two-dimensional CNN layers is set to one layer, and the convolutional kernel size in the two-dimensional CNN layer is continuously adjusted. After determining the optimal convolutional kernel size, it is fixed and the convolutional stride is adjusted. After determining the optimal convolutional stride, the number of two-dimensional CNN layers is adjusted. Through test experiments, when the number of two-dimensional CNN layers is 3, the convolutional stride is 1, the convolutional kernel size in the first two-dimensional CNN layer is 2×2, the convolutional kernel size in the second two-dimensional CNN layer is 3×3, and the convolutional kernel size in the third two-dimensional CNN layer is 3×3, the test prediction results are the most accurate. Next, with the optimal parameters of the two-dimensional CNN layer fixed, the number of neurons in the long short-term memory network LSTM and the fully connected layer are gradually adjusted, and the rest of the hyperparameters are determined in the same way. In addition, since the framework structure of the multi-task learning model MMoE is adopted, the number of expert subnets needs to be determined. Through experimental verification, when the number of expert subnets is 4, the prediction accuracy of the cooling load and heating load reaches the highest, while the prediction of the electrical load reaches the optimal effect when the number of expert subnets is 5, but it only increases by 2% compared with when the number of expert subnets is 4, and the computational complexity increases by 10%. To achieve a balance between computational efficiency and prediction accuracy as much as possible, this embodiment uses 4 expert subnets for prediction.
[0121] In this embodiment, the number of tasks set in the CNN-ECA-MMoE-LSTM multi-task learning model constructed in step S6 is 3, namely 1 electric load prediction task, 1 cooling load prediction task, and 1 heating load prediction task. The number of gating networks corresponds to the number of tasks. In this embodiment, the number of gating networks is 3. The number of long short-term memory networks (LSTMs) is 3, corresponding to the prediction tasks of electric load, cooling load, and heating load respectively. The number of layers of the LSTM selected for the task-specific sub-network of each task is 1 layer. The number of fully connected layers is 2. The first fully connected layer is used to perform feature transformation on the output of the last time step of the LSTM, and the second fully connected layer is used to output the prediction results of multiple loads according to the results of the feature transformation. When training the model, the weight ratio of the multi-task loss function is set to 1:1:1, that is, the proportion of information transmitted by the 3 tasks is the same; the activation function used when training the model is the Relu function, the optimizer is the Adam optimizer, and the loss function is the mean absolute error (MAE). The construction and training of the CNN-ECA-MMoE-LSTM multi-task learning model both use the general deep learning toolkit (Pytorch) in the Python programming language. The hyperparameters of the CNN-ECA-MMoE-LSTM multi-task learning model adopted in this embodiment are shown in Table 2.
[0122] Table 2 Hyperparameters of the CNN-ECA-MMoE-LSTM multi-task learning model:
[0123]
[0124] When training the model, the three-dimensional feature tensors in the training set of the three-dimensional feature tensor library obtained in step S4 are taken out in batches of 24 and input into the CNN-ECA-MMoE-LSTM multi-task learning model constructed in step S6 for iterative training.
[0125] Step S7: Obtain the three-dimensional feature tensor corresponding to the prediction moment according to the methods in steps S1 to S4; input the three-dimensional feature tensor corresponding to the prediction moment into the trained CNN-ECA-MMoE-LSTM multi-task learning model in step S6 for feature extraction and capture of temporal dependence relationships to obtain the prediction results of electric load, cooling load, and heating load.
[0126] The working process of the present invention is as follows: A multi-source load joint prediction method integrating decomposition and reconstruction technology and multi-task learning of the present invention first obtains the historical electric load sequence, historical cooling load sequence, and historical heating load sequence of the integrated energy system, and at the same time obtains the data sequences of the influencing factor characteristics that affect the electric load, cooling load, and heating load of the integrated energy system, forming a historical multi-source load sequence library and an influencing factor characteristic library; then, the historical electric load sequence, historical cooling load sequence, and historical heating load sequence are preliminarily decomposed by the variational mode decomposition method and reconstructed based on the hybrid feature evaluation method (HFEM) to obtain three types of aggregated sequences for each load and update the historical multi-source load sequence library; then, the influencing factor characteristic library and the updated historical multi-source load sequence library are normalized to obtain a historical multi-source load feature matrix and an influencing factor feature matrix, and the historical multi-source load feature matrix and the influencing factor feature matrix are merged in the third dimension to obtain a three-dimensional feature tensor; then, a CNN-ECA-MMoE multi-task learning model is constructed to output the results of feature sharing and feature extraction; then, three long short-term memory networks LSTM are constructed and embedded in the back end of the CNN-ECA-MMoE multi-task learning model as task-specific sub-networks, and two layers of fully connected layers are connected after each long short-term memory network LSTM, so as to obtain a CNN-ECA-MMoE-LSTM multi-task learning model. Finally, the finally constructed CNN-ECA-MMoE-LSTM multi-task learning model is trained, and the three-dimensional feature tensor corresponding to the moment to be predicted is input into the CNN-ECA-MMoE-LSTM multi-task learning model to obtain the prediction results of the electric load, cooling load, and heating load respectively.
[0127] In this embodiment, the data from January 1, 2022 to December 31, 2022 in the integrated energy system (IES) dataset of Arizona State University, Tempe Campus, USA is used for simulation. The data is provided at 1-hour intervals and includes data on electric load, cooling load, heating load, and meteorological factor characteristic data. The meteorological factor characteristic data is obtained from the meteorological station at Phoenix Sky Harbor International Airport, which is the closest to the Tempe Campus. According to the obtained data, a three-dimensional feature tensor library is constructed and divided into a training set and a test set in a ratio of 8:2, and input into the CNN-ECA-MMoE-LSTM multi-task learning model constructed by the present invention for training and prediction. The load prediction values corresponding to the corresponding dates output by the CNN-ECA-MMoE-LSTM multi-task learning model are compared with the actual load power values on the same day, and two evaluation indexes, root mean square error (RMSE) and mean absolute percentage error (MAPE), are calculated. The following are the detailed experimental results.
[0128] (1) Effectiveness analysis of the reconstruction strategy based on the hybrid feature evaluation method (HFEM):
[0129] To verify the contribution of the reconstruction strategy based on the Hybrid Feature Evaluation Method (HFEM) proposed in the present invention to the reconstruction of load data, the reconstruction strategy based on the Hybrid Feature Evaluation Method of the present invention is compared with the methods of reconstruction based on complex features, reconstruction based on coupling features, and reconstruction based on frequency features that only consider a single feature, and the influence of data reconstruction on the accuracy of energy load prediction under different reconstruction strategies is tested. Figure 9 , Figure 10 , Figure 11 The prediction results of electric load, cooling load, and heating load under different reconstruction strategies are respectively shown; the specific results of two evaluation indexes, Root Mean Square Error (RMSE) and Mean Absolute Percentage Error (MAPE), predicted under different reconstruction strategies are shown in Table 3. It can be seen that the reconstruction strategy based on the Hybrid Feature Evaluation Method proposed in the present invention has the smallest prediction error, followed by the reconstruction method based on complex features and the reconstruction method based on frequency features. The overall prediction result of the model using the reconstruction method based on coupling features is the worst, but its accuracy in predicting electric load is second only to the reconstruction strategy based on the Hybrid Feature Evaluation Method proposed in the present invention. This is because the reconstruction calculation based on coupling features takes into account the similarity between all subsequences and the original sequence, thus classifying the modal components other than modal component 1 in Table 1 into the mid-frequency fluctuation components, thereby simplifying the representation of the trend component, so it is only applicable to electric loads with strong regularity.
[0130] Table 3 Comparison results of prediction errors under different reconstruction strategies:
[0131]
[0132] The experimental results show that in the reconstruction stage after preliminary decomposition, the reconstruction strategy based on the Hybrid Feature Evaluation Method can make more comprehensive use of the different inherent attributes of each modal component. Specifically, the coupling feature evaluation is more applicable to sequences with strong regularity, while the advantages of complex feature and frequency feature evaluation lie in distinguishing and capturing mid-frequency and high-frequency features and being good at processing sequences with higher complexity. The present invention combines the advantages of the three feature evaluation methods, can deeply evaluate the attributes of each modal component from multiple dimensions, not only breaks through the limitations of a single index, but also realizes more refined classification in the reconstruction stage, thereby improving the accuracy of the overall analysis.
[0133] (2) Model ablation experiment analysis:
[0134] To verify the contributions of the two-dimensional CNN layer, the efficient channel attention mechanism ECA, and the multi-task learning model MMoE to the model proposed in the present invention, in this embodiment, under the condition of adopting the same decomposition and reconstruction strategy, the LSTM model, the CNN-LSTM model, the CNN-MMoE-LSTM model, and the CNN-ECA-MMoE-LSTM multi-task learning model were selected for prediction, and the hyperparameters of the above four models were adjusted according to the control variable method to compare the prediction results of the models under the optimal parameters. Table 4 summarizes the evaluation results of the above four models. It can be seen from Table 4 that the improvement of each model from top to bottom is significant. Among all the models, the prediction accuracy of each load of the LSTM model is the lowest, indicating that it has certain limitations in independently processing different load types. The two-dimensional CNN layer captures the long-term load pattern in the time dimension and captures the spatial correlation information of different aggregation modes in the space at the same time, extracts the variation law and internal connection of the long-term load trend and local fluctuation, and further strengthens the representation of the multi-load coupling relationship through the fusion of different channels. Therefore, compared with the LSTM model, the CNN-LSTM model has a significant improvement in the single-task case, especially in the prediction of electrical load, and the MAPE evaluation index has increased by 54.1%. By integrating the multi-task learning mechanism, the prediction accuracy of the CNN-MMoE-LSTM model is further improved on the basis of the CNN-LSTM model, especially in the prediction of cooling load and heating load, with an increase of 49.0% and 68.5% respectively compared with the prediction accuracy of the CNN-LSTM model. In the integrated energy system (IES), the cooling and heating load demands of users are supplied by energy conversion devices, and the coupling between energy conversion devices is extremely strong. Traditional single-task prediction cannot extract the coupling characteristics. Multi-task learning can accurately capture the characteristics of cooling load and heating load in seasonal and time laws by sharing features and capturing the mutual correlation between loads, thus significantly improving its prediction accuracy. In the optimal scheduling of the integrated energy system, accurate prediction of peak and valley values is particularly important. Accurate prediction of peak and valley values can optimize the power scheduling during high-load periods and the energy storage scheduling strategy during low-load periods, thus avoiding economic losses caused by phenomena such as wind curtailment and photovoltaic curtailment. However, at peak and valley moments, the prediction target changes violently and has poor regularity, which brings difficulties to prediction. Figure 12 , Figure 13 , Figure 14It shows that after integrating the multi-task learning model MMoE, compared with the CNN-LSTM model, the predicted values of the CNN-MMoE-LSTM model in each time period are closer to the true values of the load data, especially the peak and valley values of the cold load and heat load predictions are closer to the true values. This is because under multi-task learning, the data of different tasks complement each other, and the model can understand the reasons for load changes from multiple perspectives, thus more accurately predicting the peak and valley values of the load. The addition of the multi-task learning model MMoE based on the multi-task architecture has the greatest impact on improving the prediction accuracy of the model. The average MAPE evaluation index has increased by 53.3%, which is much greater than the 36.4% improvement brought by only adding the two-dimensional CNN layer. Finally, in this embodiment, the efficient channel attention mechanism ECA is further integrated on the basis of the CNN-MMoE-LSTM model, enabling the expert sub-network to more accurately focus on key features. The design concept of the expert sub-network is to perform in-depth modeling for specific tasks or specific features, and the efficient channel attention mechanism ECA enhances the ability of the expert sub-network in key feature extraction by focusing on local information and important channel features. The experimental results show that after introducing the efficient channel attention mechanism ECA, the prediction accuracies of the electric load, cold load, and heat load have increased by 26.7%, 27.1%, and 10.1% respectively, and the prediction accuracy of the model for the peak and valley values of the load has also been further improved.
[0135] Table 4 Comparison results of model ablation experiments:
[0136]
[0137] (3)Comparison of multi-task learning sharing mechanisms:
[0138] To further verify the prediction ability of the multi-task learning model architecture adopted in the present invention, in this embodiment, the CNN-ECA-MMoE-LSTM multi-task learning model of the present invention is compared with the traditional multi-task learning hard sharing mechanism, the single-task learning model (the single-task learning model selects the CNN-ECA-LSTM model), and a relatively typical soft sharing mechanism (cross-stitch network), so as to compare the prediction accuracy between the CNN-ECA-MMoE-LSTM multi-task learning model of the present invention and the conventional multi-task learning model and the single-task learning model; at the same time, to ensure the rigor of the experiment, in this embodiment, the CNN-ECA module applied in the expert subnet of the CNN-ECA-MMoE-LSTM multi-task learning model of the present invention is used as the feature sharing learner for all multi-task models, and the setting of the task-specific subnet is still retained as the same as that of the traditional multi-task learning model MMoE. The data from November 23, 2022 to December 31, 2022 in the Integrated Energy System (IES) dataset of Arizona State University, Tempe Campus, USA is selected as the test data. The comparison results are shown in Table 5, and the comparison chart of the MAPE evaluation index of various loads under different sharing mechanisms is as Figure 15 shown. Among them, the model applying the multi-task mechanism of the CNN-ECA module and the multi-task learning model MMoE performs the best in all tasks, followed by the soft sharing mechanism based on the cross-stitch network and the multi-task learning hard sharing mechanism, and the single-task learning model has the worst prediction accuracy.
[0139] Table 5 Comparison results of MAPE evaluation index under different sharing mechanisms:
[0140]
[0141] Compared with the hard sharing mechanism, the Cross-Stitch Network has varying degrees of improvement in the prediction accuracy of multiple loads. This is because the soft sharing mechanism avoids feature interference caused by forced sharing of the feature space. In particular, the cross-weight dynamic adjustment characteristic in the Cross-Stitch Network enables each task to actively select highly relevant features. Therefore, this soft sharing mechanism is applicable to scenarios where there is a certain relationship between tasks but differential feature expression is required. Compared with the hard sharing mechanism, the average value of the MAPE evaluation index of the Cross-Stitch Network based on the soft sharing mechanism has increased by 15%. However, the prediction effect of the traditional soft sharing mechanism represented by the Cross-Stitch Network is relatively dependent on the selection of the initial value of the feature sharing weight and some parameters. If the initialization biases towards certain feature combination directions, it may cause the model to overly focus on specific tasks, resulting in only some tasks having significant improvement. For example, the Cross-Stitch Network has achieved a large improvement in the prediction of electrical load with relatively low prediction difficulty, but the improvement in the prediction of the more coupled cooling load and heating load is relatively limited. Therefore, it lacks a certain fine-grained adjustment ability and feature dynamic adaptability.
[0142] The multiple expert subnets and gating networks of the CNN-ECA-MMoE-LSTM multi-task learning model of the present invention can dynamically select the most suitable feature combination without relying on fixed initial feature sharing weights, avoiding the problems of feature bias and initial weight imbalance. In the prediction of the more coupled and complex cooling load and heating load, the CNN-ECA-MMoE-LSTM multi-task learning model of the present invention can also dynamically and flexibly call different expert features according to the adjustment of the gating network, so as to more fully capture the interaction relationship between complex loads. Compared with the Cross-Stitch Network, the prediction accuracy of the CNN-ECA-MMoE-LSTM multi-task learning model of the present invention for the cooling load and heating load has increased by 40.2% and 38.2% respectively, achieving satisfactory results.
Claims
1. A multi-load joint forecasting method integrating decomposition and reconstruction technology with multi-task learning is characterized by: The steps include: Step S1, obtaining the historical load sequence of electric load, cooling load and heating load in the integrated energy system, forming a historical multi-load sequence and forming a historical multi-load sequence library; at the same time, obtaining the data sequence of the characteristics of the influencing factors affecting the electric load, cooling load and heating load, forming an influencing factor characteristic library; Step S2, preliminarily decomposing the historical load sequences of electric load, cooling load and heating load in the historical multivariate load sequence by using the variational mode decomposition method, and obtaining K mode components for each load; Step S3, reconstructing the K modal components of each load obtained in step S2 from multiple angles by using a hybrid characteristic evaluation method, and converting the K modal components corresponding to each load into a set of three types of aggregate sequences; then updating the three types of aggregate sequences corresponding to each load into a historical multivariate load sequence library to obtain an updated historical multivariate load sequence library; Step S4, normalizing the influencing factor feature library obtained in step S1 and the updated historical multivariate load sequence library obtained in step S3 with a time step length M as the sequence length, and then merging them to obtain multiple groups of three-dimensional feature tensors and form a three-dimensional feature tensor library, and divide them into a training set and a test set in a ratio of 8:2; Step S5, constructing a CNN-ECA-MMoE multi-task learning model to output the results of feature sharing and feature extraction; The CNN-ECA-MMoE multi-task learning model includes multiple expert subnets and multiple gating networks. The expert subnets include a two-dimensional CNN layer, a pooling layer, and an efficient channel attention mechanism ECA connected in sequence; the expert subnets are used to learn specific feature representations of input data; the gating network is used to dynamically adjust the output weights of each expert subnet according to the input data, and combine the output features of each expert subnet; Step S6, construct three independent long short-term memory networks LSTM, embed them into the back end of the CNN-ECA-MMoE multi-task learning model as the task-specific sub-network of the model, and connect two fully connected layers after each long short-term memory network LSTM to obtain the CNN-ECA-MMoE-LSTM multi-task learning model, and train the model using the training set divided in step S4; The LSTM network is used to capture the temporal dependency of the results of feature sharing and feature extraction; the fully connected layer is used to perform feature transformation on the output of the LSTM network, output the prediction result of the multivariate load, and perform denormalization processing to obtain the prediction value of the same scale as the input data; Step S7, according to the method in step S1 to step S4, obtain the three-dimensional feature tensor corresponding to the time to be predicted; input the three-dimensional feature tensor corresponding to the time to be predicted into the trained CNN-ECA-MMoE-LSTM multi-task learning model for feature extraction and capture of timing dependencies, and obtain the prediction results of electric load, cooling load and heating load.
2. The multi-load joint forecasting method integrating decomposition and reconstruction technology with multi-task learning according to claim 1 is characterized in that: The historical multi-load sequence in step S1 includes a historical electric load sequence, a historical cooling load sequence, and a historical heating load sequence; the influencing factor characteristics in step S1 are meteorological factor characteristics associated with the electric load, cooling load, and heating load, and the meteorological factor characteristics include air temperature, cloud type, dew point, relative humidity, solar zenith angle, surface albedo, pressure, precipitation, wind direction, and wind speed; the meteorological factor characteristics are analyzed by a Pearson correlation analysis method, and the Pearson correlation coefficients between each meteorological factor characteristic and the electric load, cooling load, and heating load are calculated respectively, and the absolute values of the Pearson correlation coefficients between each meteorological factor characteristic and the three loads are taken and then averaged to obtain the correlation coefficients between each meteorological factor characteristic and the three loads, and the correlation coefficients are sorted in descending order, and the meteorological factor characteristics corresponding to the first three groups of correlation coefficients are selected to construct an influencing factor feature library.
3. The multi-load joint forecasting method integrating decomposition and reconstruction technology with multi-task learning according to claim 2 is characterized in that: The specific steps of step S3 are: evaluating the K modal components of each load obtained in step S2 from three perspectives: complex characteristics, coupling characteristics, and frequency characteristics, respectively, calculating the complex characteristics score, coupling characteristics score, and frequency characteristics score of each modal component, and integrating the three scores to obtain a hybrid characteristics evaluation score; then classifying each modal component according to the hybrid characteristics evaluation score, and aggregating the modal components belonging to the same component category corresponding to each load, so as to convert the K modal components corresponding to each load into a group of three types of aggregation sequences, and obtain three types of aggregation sequences corresponding to each load; the three types of aggregation sequences corresponding to each load contain three sequence components, namely, trend component, fluctuation component, and random component; then updating the three types of aggregation sequences corresponding to each load to the historical multivariate load sequence library obtained in step S1, replacing the historical multivariate load sequence before reconstruction, and obtaining an updated historical multivariate load sequence library.
4. The multi-load joint forecasting method integrating decomposition and reconstruction technology with multi-task learning according to claim 3 is characterized in that: In step S3, the specific steps of calculating the complex characteristic score of the modal component are as follows: Step S311, define the modal component as , x1 represents the first data point of the modal component, represents the Nth data point of the modal component, where N is the total number of time points in the historical load series; by embedding dimension m and time delay , rearrange the modal components according to time delay and construct the phase space matrix Y G , the phase space matrix Y G The representation is as follows: , Where G is the number of rows in the matrix, ; The first data point corresponding to the modal component, The corresponding modal component data points, The corresponding modal component Data points, similarly, Corresponding to the Gth data point of the modal component, The corresponding modal component data points, The corresponding modal component The data point corresponds to the Nth data point of the modal component; Step S312, in the phase space matrix Y G In the above example, each row of data elements is regarded as an embedded vector. After arranging the data elements in the embedded vector in ascending order, the arrangement pattern corresponding to the embedded vector is obtained. ; Step S313, calculate the phase space matrix Y G All permutation modes The frequencies that appear in the modal components are denoted by , the total number of permutation patterns is recorded as , then based on the frequency distribution, the Shannon entropy formula is used to calculate the permutation entropy of the modal component. The calculation formula is as follows: , Where PE represents the permutation entropy of the modal component, It represents the natural logarithm of frequency and is used to quantify uncertainty or information content; Step S314, normalizing the permutation entropy of the modal component. The normalization formula is as follows: , In the formula, represents the normalized permutation entropy, representing the complex feature score; represents the normalization factor; where the normalized permutation entropy The range of is in [0,1]. Within the range, the larger the value of the normalized permutation entropy is, the more complex the corresponding modal component is, and the smaller the value of the normalized permutation entropy is, the more regular the corresponding modal component is.
5. The multi-load joint forecasting method integrating decomposition and reconstruction technology with multi-task learning according to claim 4 is characterized in that: In step S3, the specific steps of calculating the coupling characteristic score of the modal component are: The grey correlation analysis method is used to calculate the load nonlinear correlation coefficient between the modal components of each type of load and the historical load series of the other two types of loads, and obtain the coupling characteristic scores of the modal components of each type of load; among them, the calculation steps of the coupling characteristic scores of the modal components of the electric load are as follows: Step S321, calculating the i-th modal component of the electric load Pearson correlation coefficient with historical cooling load series , and calculate the i-th modal component of the electric load Pearson correlation coefficient with historical heat load series ; Step S322, using the grey correlation analysis method to calculate the i-th modal component of the electric load Grey correlation coefficient between the historical cooling load sequence , and the i-th modal component of the electric load Grey correlation coefficient between the historical heat load sequence , the calculation formula is as follows: , Where t represents the tth time point in the historical load series, represents the i-th modal component of the electric load The value at the tth time point in represents the value at the tth time point in the historical cooling load series, represents the value at the tth time point in the historical heat load sequence, It is expressed as the resolution coefficient; Indicates the minimum absolute value. Indicates taking the maximum absolute value; Step S323, calculate the i-th modal component of the electric load The coupling characteristic score , the specific calculation formula is as follows: 。 6. The multi-load joint forecasting method integrating decomposition and reconstruction technology with multi-task learning according to claim 5 is characterized in that: In step S3, the specific steps of calculating the frequency characteristic score of the modal component are as follows: Step S331, perform Hilbert transform on each modal component, obtain the analytical time series corresponding to the modal component, and convert the analytical time series into polar coordinate form; the calculation formula of Hilbert transform is: , in, represents the analytical signal of the ith modal component, represents the value of the i-th modal component at the t-th time point, represents the i-th modal component at time point The value at , j represents the imaginary unit, represents the time delay in the integration, Represents the time difference between two consecutive time points. is the instantaneous amplitude of the modal component, represents the complex exponential function, is the instantaneous phase of the modal component; Step S332, based on the instantaneous phase of the modal component Calculate instantaneous frequency , the specific calculation formula is: , in, Instantaneous phase The differential of, dt represents the differential at time point t; Step S333: calculate the instantaneous frequency of the entire time period N of the historical load sequence. Average to get the center frequency of the i-th modal component , the specific calculation formula is: ; Among them, the center frequency represents the frequency characteristic score, Instantaneous frequency Perform integral operation.
7. The multi-load joint forecasting method integrating decomposition and reconstruction technology with multi-task learning according to claim 6 is characterized in that: In step S3, the calculation formula of the mixed characteristic evaluation score is as follows: , In the formula, is the mixed characteristic evaluation score of the i-th modal component, Score the coupling characteristics; is the weight coefficient of complex characteristics, is the coupling characteristic weight coefficient, is the frequency characteristic weight coefficient.
8. The multi-load joint forecasting method integrating decomposition and reconstruction technology with multi-task learning according to claim 7 is characterized in that: In step S3, the specific method for classifying each modal component according to the mixed characteristic evaluation score is: classifying the modal components with a mixed characteristic evaluation score greater than or equal to 0 and less than 0.4 as trend components, classifying the modal components with a mixed characteristic evaluation score greater than or equal to 0.4 and less than 0.7 as fluctuation components, and classifying the modal components with a mixed characteristic evaluation score greater than or equal to 0.7 and less than 1 as random components.
9. The multi-load joint forecasting method integrating decomposition and reconstruction technology with multi-task learning according to claim 1 is characterized in that: The number of the expert subnetworks is 2 to 6, and the number of the gated networks is 3.
10. The multi-load joint forecasting method integrating decomposition and reconstruction technology with multi-task learning according to claim 1 is characterized in that: The efficient channel attention mechanism ECA captures the channel global information of the input feature map through the global maximum pooling layer GAP, and then adaptively learns the local interaction relationship between channels through a one-dimensional convolutional layer, generates attention weights using the Sigmoid function, and dynamically adjusts the response strength of each channel of the input feature map using the attention weights, thereby obtaining an output feature map containing key features.
Citation Information
Patent Citations
Comprehensive energy load prediction method based on multi-task learning strategy and deep learning
CN113822481A
Comprehensive energy system load prediction method considering multivariate load coupling characteristics
CN111950793A
Power load prediction method based on deep fusion of wavelet transform and multi-task learning
CN118249312A