Multi-energy Load Forecasting Method and System for Integrated Energy System with Zero Historical Data
Through the Tnet algorithm and improved meta-learning strategy, the source domain data is selected, and feature extraction is combined with the encoding and decoding module, the multi-energy load prediction problem under zero historical data in the newly built park in the integrated energy system is solved, and high-precision load prediction is achieved.
Patent Information
- Application Number
- CN202510412447.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-03
AI Technical Summary
In the integrated energy system, under the conditions of building new parks or lack of historical data, existing transfer learning technologies face difficulties in source domain selection and insufficient cross-domain generalization capabilities, resulting in low prediction accuracy of multi-energy loads.
The Tnet algorithm and improved meta-learning strategy are adopted to build a multi-energy load prediction model, select appropriate source domain data through campus cross-correlation and generalization ability analysis, adjust gradient weights using Metas training strategy, and combine encoder and decoder modules for feature extraction and prediction.
It realizes long-term accurate prediction of cold, heat, electricity and gas loads under zero historical data conditions, improves prediction accuracy and generalization performance of the model, and is suitable for planning and scheduling of new parks or comprehensive energy system lacking historical data.
Smart Images

Figure CN119939395B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of load forecasting, and particularly to a multi-energy load forecasting method and system under zero historical data in an integrated energy system. Background Art
[0002] The statements in this part merely provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] With the rapid development of the integrated energy system (IES), accurate forecasting of multi-energy loads has become the core requirement for the optimal operation of the system. However, in practical applications, obtaining comprehensive historical data of multi-energy loads often requires a large amount of time and capital costs, and due to reasons such as data privacy, some information cannot be made public, resulting in a limited scale of available datasets. Traditional deep learning methods face a serious data dependence bottleneck, bringing huge challenges to IES load forecasting. In response to this, many scholars at home and abroad have adopted the method of transfer learning to address this problem.
[0004] Auxiliary selection of a suitable source domain by performing correlation analysis on the load sequence is one of the most important key steps before transfer learning. Some scholars consider the complex linear and nonlinear characteristics of load data and have improved the correlation analysis, further reducing the amount of target domain data required. However, for the extreme case where the park is not yet built, the target system cannot obtain any historical load data. Traditional source domain selection relies on target domain correlation analysis, and the zero-data condition leads to the complete failure of feature comparison. Some scholars have proposed transfer learning schemes based on generative adversarial networks (GANs), fine-tuning strategies, and meta-learning frameworks, but still require a small amount of target domain data and have obvious defects in complex time series feature extraction and multi-source domain collaborative optimization.
[0005] In summary, for high-precision multi-energy load forecasting under extremely scarce load data conditions, scholars at home and abroad have given various solutions. However, under the condition of zero historical load data in a newly built park, the existing transfer learning technologies still face two bottlenecks:
[0006] (1) Difficulty in source domain selection: The similarity analysis between the source domain and the target domain in transfer learning strongly depends on the target domain data, and the traditional correlation analysis method fails under the zero-data condition.
[0007] (2) Insufficient cross-domain generalization ability: The cross-domain parameter migration does not fully consider the dynamic differences of multiple source domains, resulting in low parameter migration efficiency and large prediction errors. Summary of the Invention
[0008] To solve the above problems, the present invention proposes a multi - energy load prediction method and system under zero historical data of an integrated energy system. Based on the Tnet algorithm and an improved meta - learning strategy, a multi - energy load prediction model based on transfer learning is constructed, realizing long - term accurate prediction of multi - energy loads in the target integrated energy system under the condition of zero historical load data.
[0009] To achieve the above object, the present invention adopts the following technical solutions:
[0010] In the first aspect, the present invention provides a multi - energy load prediction method under zero historical data of an integrated energy system, including the following steps:
[0011] Obtain the meteorological characteristics of the target park and the historical data of cooling, heating, electricity, and gas of the source - domain group of parks, and pre - process the obtained data;
[0012] Perform park cross - correlation and generalization ability analysis on the pre - processed historical data of cooling, heating, electricity, and gas of the source - domain group of parks to determine appropriate source - domain data;
[0013] Construct a multi - energy load prediction model, and use the Metas training strategy to train the model based on the source - domain data. Adjust the gradient weights according to the validation loss of the inner - loop task and the source - domain generalization probability to obtain a trained prediction model;
[0014] Input the pre - processed meteorological characteristics of the target park into the prediction model to obtain a prediction result.
[0015] As an alternative implementation, the Metas training strategy is specifically:
[0016] Considering the parameter update of different source domains, the loss value and weight value of each training task in the inner layer are fully applied to the update function in the outer layer.
[0017] As an alternative implementation, the multi - energy load prediction model is composed of an encoder module and a decoder module.
[0018] As an alternative implementation, the encoder uses CNN layers to perform feature extraction and dimensionality reduction through convolution operations, and transfers the compressed data to LSTM to further capture the long - term dependence relationship in the time series.
[0019] As an alternative implementation, the multi - head attention mechanism is used to dynamically weight the relationships of different time steps, assign different hidden - state probability weights to LSTM, and focus on the influence of important information related to cooling, heating, electricity, and gas loads.
[0020] As an alternative implementation, the park cross - correlation and generalization ability analysis is specifically:
[0021] Using the Time2vec algorithm, multiple load time series data of each source domain are transformed into periodic and linear parts, processed using the MIC algorithm and the MD algorithm respectively, constructing a K-means clustering analysis to stratify the generalization potential of the source domain, removing low-potential clusters to reduce noise interference, and performing weighted fusion of MD and MIC, using grid search to find the optimal parameter range, obtaining the cross-correlation scores of the source domain group and a reasonable source domain range.
[0022] In a second aspect, the present invention provides a multi-energy load prediction system under zero historical data for an integrated energy system, including:
[0023] A data acquisition and preprocessing module, configured to: acquire the meteorological characteristics of the target park and the historical data of cooling, heating, electricity, and gas of the source domain group parks, and preprocess the acquired data;
[0024] A source domain data determination module, configured to: perform cross-correlation and generalization ability analysis on the historical data of cooling, heating, electricity, and gas of the preprocessed source domain group parks to determine appropriate source domain data;
[0025] A model construction and training module, configured to: construct a multi-energy load prediction model, use the Metas training strategy, train the model based on the source domain data, and adjust the gradient weights according to the validation loss of the inner loop task and the source domain generalization probability to obtain a trained prediction model;
[0026] A model output module, configured to: input the meteorological characteristics of the preprocessed target park into the prediction model to obtain a prediction result.
[0027] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in the first aspect is completed.
[0028] In a fourth aspect, the present invention provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method described in the first aspect is completed.
[0029] In a fifth aspect, the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, the method described in the first aspect is implemented.
[0030] Compared with the prior art, the beneficial effects of the present invention are:
[0031] The multi - energy load prediction method and system for the integrated energy system of the present invention under zero historical data, based on the Tnet algorithm and an improved meta - learning strategy, constructs a transfer - learning - driven multi - head attention optimized encoder - decoder model, achieves prediction breakthroughs through three - level cascade optimization, realizes long - term prediction of cooling, heating, electricity, and gas loads in the target domain, and significantly improves the load prediction accuracy in the zero - data scenario.
[0032] The multi - energy load prediction method and system for the integrated energy system of the present invention proposes a source - domain selection algorithm (Tnet) for the zero - historical load data scenario of the target domain, transforms the traditional source - domain similarity analysis into a probability evaluation problem of model generalization ability, selects the most suitable source - domain set through Tnet without any target - domain historical data, and realizes load prediction, breaking through the initialization problem of transfer learning under zero - data constraints.
[0033] The multi - energy load prediction method and system for the integrated energy system of the present invention designs an improved meta - learning training algorithm (Metas), introduces the loss value and weight information of each inner - loop task during the update of the outer - loop of meta - learning, makes full use of the multi - source - domain knowledge provided by Tnet, breaks through the parameter transfer barrier in time - series prediction, and improves the generalization performance of the prediction model.
[0034] Advantages of additional aspects of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.
[0036] Figure 1 It is a flowchart of the multi - energy load prediction method for the integrated energy system under zero historical data provided in Embodiment 1 of the present invention;
[0037] Figure 2 It is a flowchart of the Tnet algorithm provided in Embodiment 1 of the present invention;
[0038] Figure 3 It is a schematic diagram of the Metas algorithm provided in Embodiment 1 of the present invention;
[0039] Figure 4 It is a framework diagram of the multi - energy load prediction model provided in Embodiment 1 of the present invention, where, Figure 4 (a) is the overall framework diagram of the model, Figure 4 (b) is the framework diagram of the encoder - decoder of the model. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0041] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0042] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0043] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0044] Embodiment 1
[0045] As Figure 1 shown, this embodiment provides a multi-energy load prediction method under zero historical data of an integrated energy system, including the following steps:
[0046] Obtain the meteorological characteristics of the target park and the historical data of cooling, heating, electricity, and gas of the source domain group parks, and preprocess the obtained data;
[0047] Perform park cross-correlation and generalization ability analysis on the preprocessed historical data of cooling, heating, electricity, and gas of the source domain group parks to determine appropriate source domain data;
[0048] Construct a multi-energy load prediction model, use the Metas training strategy, train the model based on the source domain data, adjust the gradient weights according to the verification loss of the inner loop task and the source domain generalization probability, and obtain a trained prediction model;
[0049] Input the preprocessed meteorological characteristics of the target park into the prediction model to obtain a prediction result.
[0050] To solve the extreme problem of no historical load data in the target domain, this application proposes a brand-new multi-energy load forecasting method. This method includes a new source domain selection algorithm, an improved meta-learning training strategy, and an encoder-decoder prediction model integrating the multi-head attention mechanism. Through three-level cascade optimization, prediction breakthroughs are achieved, so as to realize long-term accurate forecasting of various loads of cooling, heating, electricity, and gas in the target integrated energy system. It is applicable to the planning and scheduling of newly built parks or integrated energy systems lacking historical data.
[0051] As Figure 1 shown. First, preprocess the original data of the source domain group and the target park, including outlier detection, data normalization, feature contribution judgment, etc. Secondly, use the Tnet algorithm to analyze the cross-correlation and generalization ability of the source domain group, convert the similarity analysis problem into a classification problem, and find the parks suitable as the source domain. Then, combined with the data preprocessing of the target park, construct an autocorrelation analysis to determine the type and period of the input data. Subsequently, input the data of multiple parks into the prediction model, and successively pass through the encoding and decoding module for feature extraction and sharing composed of a Convolutional Neural Network (CNN), a Long Short-Term Memory (LSTM), and a Multi-Head Attention (MHA) to decompose and extract the coupled information, and obtain the joint prediction result. Rely on Metas to apply the loss value and weight value of each inner task to the outer function to update the model parameters. Finally, bring the weather characteristics of the target park into the updated model to achieve multi-energy forecasting with zero historical data volume of cooling, heating, electricity, and gas load data.
[0052] Among them, use the Tnet algorithm to analyze the cross-correlation and generalization ability of the source domain group to find the parks suitable as the source domain. Specifically:
[0053] The Tnet algorithm reconstructs the source domain selection paradigm through a five-stage probability inference framework, and uses five main steps to find the appropriate source domain data: Time to Vector (Time2vec), Mahalanobis Distance (MD), Maximal Information Coefficient (MIC), K-means Clustering (K-means), and Bayesian Weighted Probability Averaging Method (BWPA). Figure 1 In Respectively referring to the best source domain groups selected by the Tnet algorithm, the best source domain groups are sets of industrial parks, and each industrial park has its own historical load data and weather data. The data of these best source domain groups are brought into the prediction model to train the parameters of the model, and the parameters of the model are updated according to the loss between the data predicted by the prediction model and the real data. However, the loss cannot be directly used, and the Metas algorithm needs to be used to convert the loss into parameters. represents the loss after weighting, which is only part of the Metas training strategy. The other part is described in Figure 3 (that is, gradient update is performed through the loss). After the model parameters are updated, through transfer learning, the model is transferred to the target industrial park to be predicted, and the load data of the industrial park is predicted through the weather data of the target industrial park. Figure 1 in refers to predicting 24h of data each time, that is, predicting 24h in the first round and also 24h in the second round, adding up to 48h, which is a stacked graph (indicating that the predicted load is increasing continuously).
[0054] Figure 2 The flowchart is given. Since the historical data of the cold, heat, electricity, and gas loads of the target industrial park are unknown, it is crucial to find a similar industrial park group in the source domain group as the source domain. Given the periodic and linear characteristics of the load data used, simply considering linear and non-linear factors and directly using linear analysis and non-linear analysis will inevitably be affected by opposing factors. Therefore, the Time2vec algorithm is used to transform the multiple load time series data of each source domain into two parts: periodic and linear, and map the time features to a high-dimensional space. The periodic part mainly considers the seasonal changes in load prediction, while the linear feature captures the baseline trend of the load data. As shown in formula (1). Among them, represents the original time series, represents the time step, represents the embedding dimension, the 0 dimension represents the linear feature, and the 1 to k dimensions represent the non-linear features.
[0055] (1)
[0056] To conduct the joint linear and non - linear analysis of high - dimensional space, the MD algorithm is used to introduce the covariance matrix to adjust the weights of different features, so that features with different scales and correlations can be treated fairly during calculation. The MIC algorithm is used to divide different non - linear variables into different grids, and through different discretization strategies, the maximum amount of information is found, as shown in formulas (2) and (3). Furthermore, MD and MIC are used to construct K - means clustering analysis to stratify the generalization potential of the source domain, and low - potential clusters are removed to reduce noise interference. If multiple parks are included in a cluster and used uniformly, the exact generalization ability of each park for the target park cannot be considered, and the prediction error will be large. Therefore, MD and MIC are weighted and fused using formula (4), and grid search is used to find the optimal parameter range to obtain the cross - correlation score of the source domain group in this cluster and a reasonable source domain range.
[0057] (2)
[0058] (3)
[0059] (4)
[0060] (5)
[0061] Finally, in order to make full use of the contribution degree of different source domains to the target domain in the prediction model, weighted mutual information entropy is introduced to construct a Bayesian weighted model. The reasonable source domain is used as the label, and the in formula (4) is used to construct a parameter model and substitute it into formula (5), considering the "uncertainty" of each data point to assign weights to it. In this way, the traditional probability prediction problem is converted into a classification problem to find the most suitable park group. Among them, and are the data of cooling, heating, electricity, and gas of two different parks, both of which are 4 - dimensional. is expressed as the inverse matrix of, representing the correlation between features of data points, which is a matrix. is to divide the data into grids, is the mutual information estimate value under the grid , and the maximum value is selected from all grids that satisfy (usually , is the sample size), is the prior probability of the model , is the prediction result, is the input feature.
[0062] Therefore, in this application, this algorithm can be regarded as the problem of maximizing the generalization ability after standardizing data of four different scales: heat, electricity, cold, and gas. At the same time, the Tnet algorithm is used to analyze the source domain group, and the probability that the source domain can represent the target domain is obtained through BWPA (Bayesian weighted probability), which also provides valuable reference for the optimization method in the following prediction model (the inner layer of the Metas algorithm below needs to use this probability value for weighting).
[0063] An improved meta-learning training strategy, specifically:
[0064] Meta-learning updates by performing gradient descent on each task and uses a meta-optimizer to minimize the test loss of all tasks, thereby optimizing the initialization parameters of the model, enabling it to quickly adapt to new tasks and improving the generalization ability of the prediction model. Although Tnet effectively solves the key problem of source domain selection in meta-learning and meets the stringent requirements of meta-learning for high-quality prior tasks, there are still limitations in the parameter update mechanism. Specifically, the meta-layer update optimizes the initialization parameter state between tasks through inner-layer gradient descent. This design may lead to two contradictions in time series prediction tasks: on the one hand, the inner-layer update fails to capture the dynamic correlation of time series features, resulting in difficulty for model parameters to fully fit complex patterns in the time dimension; on the other hand, the initialization parameters shared between tasks lack adaptive adjustment to source domain differences during cross-domain migration, ultimately leading to a significant decline in the prediction accuracy of the model. Inspired by the meta-learning idea, an improved meta-learning parameter optimization method is proposed. As Figure 3 shown, three source domains are selected as representatives, represents the initial parameters in the prediction model, represents the final parameters obtained after the prediction model is updated m + 1 times, where x 1, x 2 to x 3 represents the probability that a certain park is the source domain, 1 represents the loss corresponding to Park 1, 1 is for gradient update, 1 represents the inner update step size. Similarly, 1 and 1 are the parameters corresponding to Park 2 and Park 3 respectively. The Metas algorithm updates through two layers (including the inner layer and the outer layer), making both layer updates act on the hyperparameters of the model. The inner layer is the loss value of each task and the source domain weight information calculated by Tnet, which is dynamically incorporated into the outer layer parameter update to achieve differential gradient adjustment. By using the loss value and weight value of each task in the inner layer to fully act on the update function of the outer layer, the "inner layer loss - outer layer update" is realized through the dynamic gradient weighting mechanism, so that the parameter updates of different source domains can be fully considered to improve the prediction accuracy. As shown in formula (6). Among them is the model vector parameter, is the task training loss parameter, is the task validation loss parameter, is the task weight calculated by the Tnet algorithm, is the internal gradient update step size, is the outer layer learning rate, is the batch size, is the input model prediction output, task Hessian matrix of the training loss.
[0065] (6)
[0066] Based on the above analysis, this application proposes a multi-energy load prediction model based on transfer learning (Metas-Tnet-CNN-MHA-LSTM), as Figure 4 shown, this model is an m2m multi-step load prediction model. Figure 4 In i (1) to i (n) These represent the inputs, o (1) to o (n) represent the prediction outputs, y represents the true value, which is used to calculate the loss with the predicted value. The final output of the LSTM (Long Short-Term Memory Network) , t represents the time step, represents the output vector of the long short-term time network at a certain moment, and this is passed to the multi-head attention mechanism (MHA). The principle of MHA is to analyze the local features of the time series, focus on the features to find the truly weighted content. is divided into two parts ( ) for the reason that: to perform MHA analysis on , the above uses for weighting, and the following For calculation How much. Here is the head of MHA, that is, the head of the multi-head attention mechanism.
[0067] Among them, Figure 4 (a) is the overall model framework, the number of source domain parks is set to , the number of features is , and the number of output features is . The prediction model is of the m2m structure, as shown in Figure 4 (b). It consists of an encoder module and a decoder module. The input is different weather features, and the output is the predicted values of heat, cold, electricity, and gas. In the encoder, CNN layers are used to perform feature extraction and dimensionality reduction through convolution operations, and the compressed data is passed to LSTM to further capture the long-term dependencies in the time series. The multi-head attention mechanism is used to dynamically weight the relationships of different time steps, giving different hidden state probability weights to LSTM, and focusing on the influence of important information related to the heat, cold, electricity, and gas loads. The global modeling ability of the model is further optimized. Finally, the hidden layer relationships obtained by the encoder module are input into the decoder module with LSTM as the core to output the final load prediction values.
[0068] So far, a multi-energy load prediction model based on transfer learning for zero historical load data has been established. To verify the effectiveness of the proposed method, three case experiments are carried out for illustration.
[0069] Case 1: Prediction for commercial parks.
[0070] The source domain data selects the electricity, heat, gas, and cold load data of 9 parks (sampling interval of 1 hour) and 6 meteorological features (temperature, humidity, wind speed, etc.), and performs normalization and outlier processing. The target domain data only inputs 6 meteorological data without historical load records. The Tnet algorithm is executed to select the appropriate source domains as 1, 7, and 8. Then, comparative experiments are carried out using other source domains, and the results are shown in Table 1:
[0071] Table 1 Comparison of prediction accuracies for different source domains in commercial parks;
[0072]
[0073] From the experimental data in Table 1, it can be seen that the Tnet source domain selection algorithm shows significant advantages when selecting the source domain. By comparing the Metas-CNN-MHA-LSTM model with different source domains, the performance differences presented by each source domain under different target domains can be clearly observed. Especially for target park 0, the Tnet algorithm successfully selects source domain groups 1, 7, and 8 as the best source domain groups, and this selection significantly improves the model's fitting degree and prediction accuracy. Further analysis shows that the prediction effect is the worst in source domains 5 and 9, which is exactly consistent with the result that the probability of Tnet selecting 5 and 9 is low. Secondly, the prediction effect of source domains 1, 4, 7, and 8 is slightly worse, which is due to the decrease in accuracy caused by the probability of including source domain 4 being <80%. The two source domains with the highest probabilities are not selected because it is hoped to fully utilize the three source domains with the highest probabilities for prediction, and their probabilities are similar, and the final prediction results are also very similar, thus proving the reliability of the probability values output by Tnet. Through these experimental analyses, the advantage of the Tnet algorithm lies in its ability to accurately select the best source domain group in different combinations of source domains and target domains, helping the model obtain better performance. Especially in multi-source domain and multi-target domain tasks, the Tnet algorithm obviously has strong adaptability and effect optimization ability.
[0074] Case 2: Office park prediction.
[0075] To further verify the generalization ability of the proposed prediction model in different scenarios, this experiment conducts cross-regional and cross-application scenario tests using the same model architecture: Metas-Tnet-CNN-MHA-LSTM. It focuses on analyzing the adaptive ability of the model in different scenarios to verify its cross-scenario robustness. An office park is selected as the new target domain. Its cooling load is significantly affected by the weekday periodicity, in sharp contrast to the commercial park in Section 3.4. The Tnet algorithm selects a high-probability source domain group (1, 5, 9) for the office park. The results in Table 2 show that when using the preferred source domain group (5, 9), the MAPE values of the electric, heat, and cooling loads are reduced to 10.166%, 10.341%, and 9.342% respectively, and the R² values all exceed 85%, significantly better than other combinations. It is worth noting that when including low-probability source domains (such as 2, 4), the R² value of the cooling load prediction is even negative (-45.2%), indicating that the interference of heterogeneous features leads to model failure. Tnet excludes such interfering source domains (probability < 60%) through Bayesian weighting, ensuring that the model focuses on subgroups with high generalization ability. In contrast, the MAPE of the cooling load for traditional single-source domain migration (such as only using park 5) is 12.557%, and the error is 34.4% higher than that of the Tnet preferred combination, highlighting the necessity of multi-source domain collaboration. In addition, source domain 3 is given a higher weight (probability 70% - 80%) by the Tnet algorithm because it highly matches the target domain in terms of the time periodic component and local nonlinear trend. Its heat load prediction contributes key information (MAPE = 15.585%, R² = 67.7%), indicating that Tnet can adaptively adjust the source domain priority based on multi-dimensional features.
[0076] The Metas-Tnet-CNN-MHA-LSTM model dynamically captures cross-domain spatio-temporal dependencies through the multi-head attention mechanism, alleviating the distribution shift problem. For example, in gas load prediction, the linear correlation between the target domain and source domains 5 and 9 is relatively low compared to other load predictions. However, the model further extracts cross-domain non-linear dependencies (such as load fluctuation patterns) through attention weights, resulting in a significantly lower MAPE (16.421%) compared to the full source domain combination (19.372%). Moreover, during the gradient update process, the Metas optimization strategy considers the adaptability of different source domain features to the target domain, enabling the model to learn the optimal parameter configuration faster and further improving the prediction accuracy. In similar scenarios, the Tnet algorithm effectively identifies source domain groups with heterogeneous distributions but complementary functions through spatio-temporal feature decoupling and probabilistic optimization. Metas-Tnet-CNN-MHA-LSTM, on the other hand, achieves directional enhancement of cross-domain non-linear features through the attention mechanism. The two work together to keep the model at a high accuracy level (MAPE < 10.15%, R² > 83.7%) in similar generalization tasks. Meanwhile, through the experiments in this section, it is verified that the proposed model is not only applicable to specific target domains but also can be extended to different application scenarios, providing a theoretical basis for the standardized deployment of campus-level integrated energy systems.
[0077] Table 2 Predicted values for the office park;
[0078]
[0079] Case 3: Prediction optimization with few-shot fine-tuning.
[0080] Based on the above zero-data migration, to further explore the fine-tuning effect of a small amount of target domain data on the prediction model, this section introduces a few-shot transfer learning strategy and analyzes the impact of different fine-tuning data ratios on the model accuracy. The experiments use 3.3% (12h), 10% (36h), and 16.7% (60h) of the target domain data for model fine-tuning and adopt the layer-wise freezing strategy (LFS), that is, freezing the first two layers of the CNN and LSTM and only unfreezing the last layer of the LSTM and the multi-head attention layer. This can not only retain the general features learned from the source domain but also adapt to the specific patterns of the target domain, improving the prediction accuracy. Table 3 shows the prediction results under different amounts of transferred data.
[0081] Table 3 shows that after introducing 3.33% of the target domain data (12 hours), the cold load MAPE decreased from 9.329% to 7.094%, and the R² increased by 6.9 percentage points to 85.5%. When the fine-tuning data increased to 16.7% (60 hours), the power load MAPE was optimized to 7.916% (R² = 89.6%), and the error was reduced by 13.2% compared to the zero-sample case. This improvement stems from the dynamic adjustment of the attention layer: the model strengthens the feature extraction of key time steps in the target domain (such as weekday peaks) by updating the multi-head attention weights, effectively alleviating the local distribution differences. The R² value of the gas load prediction increased to 86.1% after fine-tuning, indicating that the model can quickly correct the feature mapping relationship with a small amount of data. It is worth noting that although a relatively small amount of data (such as 3.3%) can significantly improve the prediction accuracy, when the data volume further increases (such as 16.7%), the improvement amplitude tends to converge.
[0082] This shows that the prediction model based on the Tnet-Metas framework already has strong generalization ability. Even when the amount of fine-tuning data is small, it can make full use of the source domain knowledge for accurate prediction. In addition, the experimental results also verify the effectiveness of the hierarchical freezing strategy in capturing the unique laws of the target domain by retaining the global features of the source domain and adjusting local parameters. It avoids the overfitting problem caused by excessive adjustment of the underlying feature extraction layer during the fine-tuning process. This method has important application value in scenarios where data can be incrementally obtained.
[0083] Table 3 Small sample fine-tuning prediction results;
[0084]
[0085] Example 2
[0086] This example provides a multi-energy load prediction system for a comprehensive energy system with zero historical data, including:
[0087] A data acquisition and preprocessing module, configured to: acquire the meteorological characteristics of the target park and the historical data of cooling, heating, electricity, and gas of the source domain group parks, and preprocess the acquired data;
[0088] A source domain data determination module, configured to: perform park cross-correlation and generalization ability analysis on the preprocessed historical data of cooling, heating, electricity, and gas of the source domain group parks, and determine appropriate source domain data;
[0089] A model construction and training module, configured to: construct a multi-energy load prediction model, use the Metas training strategy, train the model based on the source domain data, and adjust the gradient weights according to the validation loss of the inner loop task and the source domain generalization probability to obtain a trained prediction model;
[0090] A model output module, configured to: input the preprocessed meteorological features of the target park into a prediction model to obtain a prediction result.
[0091] It should be noted here that the above modules correspond to the steps described in Embodiment 1. The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.
[0092] In more embodiments, there is also provided:
[0093] An electronic device includes a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in Embodiment 1 is completed. For the sake of brevity, it will not be elaborated here.
[0094] It should be understood that in this embodiment, the processor may be a central processing unit CPU, and the processor may also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0095] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.
[0096] A computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by the processor, the method described in Embodiment 1 is completed.
[0097] The method in Embodiment 1 can be directly embodied as being executed by a hardware processor, or completed by a combination of hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0098] A computer program product includes a computer program. When the computer program is executed by the processor, the method described in Embodiment 1 is implemented and completed.
[0099] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which are executed in a device on a target real or virtual processor to perform the processes / methods described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functions of program modules can be combined or divided among program modules as needed. The machine-executable instructions for program modules can be executed within local or distributed devices. In a distributed device, program modules can be located in local and remote storage media.
[0100] The computer program code for implementing the method of the present invention can be written in one or more programming languages. This computer program code can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program code is executed by the computer or other programmable data processing device, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the computer, partially on the computer, as a stand-alone software package, partially on the computer and partially on a remote computer, or entirely on a remote computer or server.
[0101] In the context of the present invention, the computer program code or related data can be carried by any suitable carrier so that a device, apparatus, or processor can perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, and the like. Examples of signals can include electrical, optical, radio, acoustic, or other forms of propagated signals, such as carrier waves, infrared signals, etc.
[0102] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0103] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation to the protection scope of the present invention. Those skilled in the art should understand that, based on the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.
Claims
1. A multi-energy load forecasting method for integrated energy systems under zero historical data, characterized in that It includes the following steps: Obtain the meteorological characteristics of the target park and the historical data of cooling, heating, electricity, and gas of the source domain group of parks, and preprocess the obtained data; Perform park cross-correlation and generalization ability analysis on the preprocessed historical data of cooling, heating, electricity, and gas of the source domain group of parks to determine appropriate source domain data; The park cross-correlation and generalization ability analysis is specifically as follows: Use the Time2vec algorithm to transform the multiple load time series data of each source domain into two parts: periodic and linear. Use the MIC algorithm and the MD algorithm for processing respectively. Construct a K-means clustering analysis to stratify the generalization potential of the source domain, eliminate low-potential clusters to reduce noise interference, and perform weighted fusion of MD and MIC. Use grid search to find the optimal parameter range, obtain the cross-correlation score of the source domain group and a reasonable source domain range, and obtain the probability that the source domain can represent the target domain through Bayesian weighted probability; Construct a multi-energy load prediction model, use the Metas training strategy, train the model based on the source domain data, and adjust the gradient weight according to the validation loss of the inner loop task and the source domain generalization probability to obtain a trained prediction model; Input the preprocessed meteorological characteristics of the target park into the prediction model to obtain the prediction result.
2. The multi-energy load forecasting method under zero historical data of the integrated energy system according to claim 1, wherein The Metas training strategy is specifically as follows: Consider the parameter update of different source domains, and fully apply the loss value and weight value of each training task in the inner layer to the update function in the outer layer.
3. The multi-energy load forecasting method under zero historical data of the integrated energy system according to claim 1, wherein The multi-energy load prediction model is composed of an encoder module and a decoder module.
4. The multi-energy load forecasting method under zero historical data of the integrated energy system according to claim 3, characterized in that The encoder uses a CNN layer to perform feature extraction and dimensionality reduction through convolution operations, and transfers the compressed data to the LSTM to further capture the long-term dependence relationship in the time series.
5. The multi-energy load forecasting method under zero historical data of the integrated energy system according to claim 4, wherein Use the multi-head attention mechanism to dynamically weight the relationships of different time steps, assign different hidden state probability weights to the LSTM, and focus on the influence of important information related to the cooling, heating, electricity, and gas loads.
6. A multi-energy load forecasting system for an integrated energy system with zero historical data, characterized in that, It includes: A data acquisition and preprocessing module, configured to: obtain the meteorological characteristics of the target park and the historical data of cooling, heating, electricity, and gas of the source domain group of parks, and preprocess the obtained data; A source domain data determination module, configured to: perform park cross-correlation and generalization ability analysis on the preprocessed historical data of cooling, heating, electricity, and gas of the source domain group of parks to determine appropriate source domain data; The park cross-correlation and generalization ability analysis is specifically as follows: Use the Time2vec algorithm to transform the multiple load time series data of each source domain into two parts: periodic and linear. Use the MIC algorithm and the MD algorithm for processing respectively. Construct a K-means clustering analysis to stratify the generalization potential of the source domain, eliminate low-potential clusters to reduce noise interference, and perform weighted fusion of MD and MIC. Use grid search to find the optimal parameter range, obtain the cross-correlation score of the source domain group and a reasonable source domain range, and obtain the probability that the source domain can represent the target domain through Bayesian weighted probability; A model construction and training module, configured to: construct a multi-energy load prediction model, use the Metas training strategy, train the model based on the source domain data, and adjust the gradient weight according to the validation loss of the inner loop task and the source domain generalization probability to obtain a trained prediction model; The model output module is configured to: input the meteorological characteristics of the target park after preprocessing into the prediction model to obtain a prediction result.
7. An electronic device, characterized in that, It includes a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method according to any one of claims 1-5 is completed.
8. A computer-readable storage medium, characterized in that, For storing computer instructions, when the computer instructions are executed by the processor, the method according to any one of claims 1-5 is completed.
9. A computer program product, characterized in that, It includes a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1-5 is implemented.
Citation Information
Patent Citations
Short-term load prediction method under small sample set based on transfer learning
CN114169416A
Festival and holiday load prediction method and system based on deep transfer learning
CN118676910A