Comprehensive energy load day-ahead probability prediction model establishment method and device
Through the combination of Copula entropy correlation coefficient and multi-scale feature-time Transformer network, the problem of the coupling relationship of multi-energy loads being insufficiently modeled is solved, and a more accurate prediction of the pre-date probability of comprehensive energy loads is achieved.
Patent Information
- Application Number
- CN202510505271.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-08-05
AI Technical Summary
The prior art fails to fully consider the coupling relationship between multi-energy loads and its uncertainty, resulting in insufficient accuracy and credibility of the prediction of the previous probability of comprehensive energy loads.
The Copula entropy correlation coefficient was used to perform multi-dimensional influencing factors correlation analysis, combined with the multi-scale feature-time Transformer network and conformal prediction, a comprehensive energy load pre-probability prediction model was established, and the load prediction interval was generated through a multi-level attention mechanism and a multi-scale aggregator.
It significantly improves the accuracy and credibility of the prediction of the previous days of comprehensive energy load, can capture multi-scale timing structural characteristics and global dependencies, and generate load prediction intervals with confidence levels.
Smart Images

Figure CN120429638A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power big data analysis, and in particular to a method and device for establishing a comprehensive energy load day-ahead probability prediction model. Background Art
[0002] As the penetration of renewable energy in integrated energy systems continues to increase, the volatility and uncertainty of their output pose unprecedented challenges to the safe and stable operation of energy systems. Against this backdrop, accurate day-ahead probabilistic forecasts of integrated energy load not only provide a scientific basis for system scheduling but also provide data support for energy market transactions. Compared to traditional deterministic forecasting methods, probabilistic load forecasts can quantify the magnitude and uncertainty of load fluctuations, providing more comprehensive support for scheduling decisions. Probabilistic interval forecasts are particularly important in day-ahead load forecasting. They focus on the range of load fluctuations and, by providing the range within which load values may fall, help decision makers make more effective risk management decisions during scheduling. For example, in the allocation of reserve capacity, traditional forecasting methods often fail to fully account for the extreme range of load fluctuations. Probabilistic interval forecasts, however, can more accurately allocate reserve resources based on a defined confidence level, thereby reducing resource waste or capacity shortages caused by forecast bias. Furthermore, probabilistic load forecasts can provide effective data support for the formulation of demand response strategies, the design of energy price regulation mechanisms, and the coordinated optimization of multiple energy sources in integrated energy systems, further improving the economic efficiency and flexibility of energy system operations.
[0003] However, in the process of researching and practicing the existing technologies, the inventors found that most of the current research focuses on the deterministic prediction of the integrated energy load, and fails to fully consider the coupling relationship between multiple energy loads and the uncertainty caused by it, thus affecting the accuracy of the prediction. Although some studies have begun to focus on the probability interval prediction of a single energy load, there are still the following deficiencies in the day-ahead probability prediction of the integrated energy load: on the one hand, the existing correlation analysis method has limitations in revealing the complex correlations between loads; on the other hand, there is a lack of comprehensive modeling of the characteristics of multiple energy loads and the uncertainty of their interactions. Therefore, it is urgent to propose a method and device for establishing a day-ahead probability prediction model for the integrated energy load, which can fully explore the coupling structure and uncertainty characteristics between multiple energy loads, combine external influencing factors, and improve the prediction accuracy and credibility to meet the high reliability prediction needs of the integrated energy system in scheduling optimization and market transactions. Summary of the Invention
[0004] The purpose of the present invention is to provide a method and device for establishing a comprehensive energy load day-ahead probability forecasting model to solve the problem that the existing technology fails to fully model the complex coupling relationship and its uncertainty among multiple energy loads, thereby improving the accuracy and credibility of the forecast.
[0005] The present invention provides a method for establishing a comprehensive energy load day-ahead probability forecasting model, comprising:
[0006] Preprocess the collected comprehensive energy load historical data;
[0007] Based on the Copula entropy correlation coefficient, the pre-processed data is subjected to multi-dimensional influencing factor correlation analysis to obtain the analysis results;
[0008] According to the analysis results, a comprehensive energy load day-ahead probabilistic forecasting model is established by integrating multi-scale feature-time Transformer network and shape-preserving forecasting.
[0009] Furthermore, the multi-dimensional influencing factor correlation analysis is performed on the pre-processed data based on the Copula entropy correlation coefficient to obtain analysis results, including:
[0010] In order to solve the problem that the original Copula entropy has inconsistent dimensions and an unfixed numerical range, and is difficult to be directly used for horizontal comparison and screening of the correlation strengths between different external influencing factors, a correlation coefficient based on Copula entropy is constructed, which is defined as follows:
[0011]
[0012] Where C represents the Copula entropy between the integrated energy load and external influencing factors, and M represents the Copula entropy between different types of integrated energy loads. Based on the correlation coefficient, external influencing factors are screened by setting a threshold. When the correlation coefficient is greater than the threshold, the external influencing factor is determined to have a significant impact on the integrated energy load and is used as an input feature of the prediction model to improve the model's prediction accuracy and robustness.
[0013] Furthermore, based on the analysis results, a comprehensive energy load day-ahead probability forecasting model is established that integrates multi-scale feature-time Transformer network and shape-preserving forecasting, including:
[0014] Feed the input features of the prediction model into the embedding layer to generate an embedded representation of the time series;
[0015] The multi-scale selector extracts information of different scales through time series decomposition, constructs multi-scale features, and generates weights of corresponding scales for subsequent predictive modeling;
[0016] Input the output of the multi-scale selector into a multi-level attention mechanism, which includes an intra-segment feature-time joint attention mechanism and an inter-segment attention mechanism, and fuse the above attention mechanisms to obtain the output result of the multi-level attention mechanism;
[0017] The multi-scale aggregator performs weighted aggregation on the outputs of each scale to generate a representation of multi-scale information aggregation;
[0018] Based on the shape-preserving prediction method, the representation of the above multi-scale information aggregation is processed to finally obtain the day-ahead prediction probability interval result of the comprehensive energy load cell.
[0019] Furthermore, the multi-scale selector extracts information of different scales through time series decomposition, constructs multi-scale features, and generates weights of corresponding scales for subsequent predictive modeling, including:
[0020] The multi-scale selector divides the time series X into M different segment scales through the expert system, which is defined as the set V = {V 1 ,V 2 ,...,V M The multi-scale selector uses the time decomposition module to calculate the weight of each scale and uses Fourier transform to convert the time series from the time domain to the frequency domain. In order to maintain the sparsity of the frequency domain, the first K frequency components with the largest amplitude {f1,f2,...,f K}, the specific process is as follows:
[0021] X season =IDFT({f1,f2,...,f K},φ,A)
[0022] Where IDFT stands for inverse Fourier transform, φ and A represent the phase and amplitude of each frequency, respectively. The trend component can be expressed as follows:
[0023]
[0024] Where, X res =XX season . is the pooling function with the i-th kernel, and Softmax(L(·)) is the weight of the results of different kernels. The transformed results are as follows:
[0025] X Trans =X season +X trend +X
[0026] In the process of weight generation, a noise term is introduced to increase randomness. The formula is as follows:
[0027] R(X Trans )=Softmax(X Trans W r +ε·Softplus(X Trans W noise )),ε~N(0,1)
[0028] Where R(·) represents the weight generation function, W r and W noise Parameters generated for weights.
[0029] Furthermore, the output of the multi-scale selector is input into a multi-level attention mechanism, which includes an intra-segment feature-time joint attention mechanism and an inter-segment attention mechanism, and the above attention mechanisms are integrated to obtain the output result of the multi-level attention mechanism, including:
[0030] The intra-segment temporal attention mechanism aims to capture the local temporal pattern on a single time scale within segment i. The specific formula is:
[0031]
[0032] Where, and is the query, key, value and corresponding weight matrix in the attention mechanism corresponding to time t in segment i, d model is the dimension of the embedding layer. represents the temporal attention mechanism of all objects in segment i, and Softmax is the activation function.
[0033] The intra-segment feature attention mechanism focuses on the correlation between different features within a single time scale within segment i. The specific formula is:
[0034]
[0035] Where, and is the query, key, value and corresponding weight matrix in the attention mechanism corresponding to the feature f in the segment i. represents the feature attention weights of all time steps within segment i.
[0036] Combining the intra-fragment feature attention mechanism and the temporal attention mechanism, we get the intra-fragment feature-temporal attention mechanism. The specific formula is as follows:
[0037]
[0038] Exchange A′ intra[i]_ft The first two dimensions of For V intra[i] Each dimension of A intra[i]_ft and V intra[i] The product of the feature of segment i is the output of the temporal attention mechanism O intra[i] .
[0039] Ointra[i] =A intra[i]_ft ·V intra[i]
[0040] Then the output of the attention mechanism in all segments is
[0041] The inter-fragment attention mechanism is used to capture global correlation at a single time scale, and its calculation formula is:
[0042]
[0043] Where, and The inputs of the attention mechanism between fragments are query, key and value, respectively, where d′ model =P·d model ,
[0044] The output of the j-th scale multi-level attention mechanism is obtained by fusing the intra-fragment attention mechanism with the inter-fragment attention mechanism:
[0045]
[0046] Furthermore, the multi-scale aggregator performs weighted aggregation on the outputs of each scale to generate a representation of the multi-scale information aggregation, including:
[0047] The multi-scale aggregator generates an aggregated representation of multi-scale information by weighted aggregation of outputs of different scales. The specific formula is:
[0048]
[0049] Where M represents the total number of segments of different scales, R(·) represents the weight generation function, and T j (·) represents the function for converting different scales.
[0050] Furthermore, the shape-preserving prediction method processes the representation of the above multi-scale information aggregation to finally obtain the day-ahead prediction probability interval result of the integrated energy load unit, including:
[0051] Assuming historical load data Where H is the historical data step size, and d is the dimension of multi-source load data. The prediction interval for each future step size can be expressed as:
[0052]
[0053] Where, is the non-conformance score, and L is the future prediction time range. Through shape-preserving prediction, we ensure that for each prediction step l, the true value y t+lIt falls within the corresponding prediction interval with a probability of at least (1-α).
[0054] The present invention provides a device for establishing a comprehensive energy load day-ahead probability forecasting model, comprising:
[0055] A preprocessing module is used to preprocess the collected comprehensive energy load historical data;
[0056] The multi-dimensional influencing factor correlation analysis module performs multi-dimensional influencing factor correlation analysis on the pre-processed data based on the Copula entropy correlation coefficient to obtain the analysis results;
[0057] A forecasting model module was established. Based on the analysis results, a comprehensive energy load day-ahead probability forecasting model was built that integrated a multi-scale feature-time Transformer network with shape-preserving forecasting. The multi-scale feature-time Transformer network was used to extract multi-scale features and model temporal dependencies, while the shape-preserving forecasting was used to generate load probability interval forecast results.
[0058] The present invention provides an electronic device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the steps of the method for establishing a day-ahead probability prediction model of a comprehensive energy load are implemented.
[0059] The present invention also provides a computer-readable storage medium, on which an information transmission implementation program is stored. When the program is executed by a processor, the steps of the method for establishing the above-mentioned comprehensive energy load day-ahead probability prediction model are implemented.
[0060] Compared with the prior art, the present invention has the following beneficial effects:
[0061] By introducing the Copula entropy correlation coefficient, the present invention effectively reveals the nonlinear coupling relationship between multiple energy loads and between them and external influencing factors, thereby improving the scientific nature of feature selection and the representativeness of the prediction model input; at the same time, a model that integrates multi-scale features-time Transformer network and shape-preserving prediction is proposed, which can not only capture the multi-scale temporal structure characteristics and global dependencies in load data, but also generate load prediction intervals with confidence levels, thereby significantly improving the accuracy and credibility of the comprehensive energy load day-ahead forecast.
[0062] The above content is only an overview of the technical solution of the present invention. In order to make the technical means of the present invention clearer and more understandable and to enable implementation thereof, the present invention will be further described in detail below in conjunction with specific implementation methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0064] Figure 1 Flowchart of a method for establishing a comprehensive energy load day-ahead probability forecasting model according to an embodiment of the present invention;
[0065] Figure 2 Schematic diagram of a device for establishing a comprehensive energy load day-ahead probability prediction model according to a first embodiment of the present invention;
[0066] Figure 3 It is a schematic diagram of a device for establishing a comprehensive energy load day-ahead probability prediction model according to a second embodiment of the present invention. DETAILED DESCRIPTION
[0067] To overcome the shortcomings of existing technologies, the present invention proposes a method and apparatus for establishing a comprehensive energy load day-ahead probabilistic forecasting model. This method analyzes the multidimensional factors influencing comprehensive energy load day-ahead forecasting and, in combination with comprehensive energy coupling relationships, establishes a comprehensive energy load day-ahead probabilistic forecasting model that integrates a multi-scale feature-time Transformer network with shape-preserving forecasting.
[0068] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0069] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise" and the like to indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as limiting the present invention.
[0070] In addition, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the said features. In the description of the present invention, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined. In addition, the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be a communication between the two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0071] The technical solutions of the embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0072] The embodiment of the method of the present invention provides a comprehensive energy load day-ahead probability forecasting model based on multi-scale feature-time Transformer network and shape-preserving prediction. The flow chart of the establishment method is as follows: Figure 1 As shown, the specific processing includes the following:
[0073] Step 1 (S101) pre-processes the collected comprehensive energy load historical data, specifically including:
[0074] Step 1.1: Identify and remove outliers from the collected comprehensive energy load data to improve the quality and usability of the original data;
[0075] In step 1.2, the cleaned data is normalized to eliminate the dimensionality effect, thereby improving the stability and convergence speed of the model training process.
[0076] Step 2 (S102) performs a multi-dimensional influencing factor correlation analysis on the pre-processed data based on the Copula entropy correlation coefficient to obtain the analysis results, which specifically include:
[0077] In step 2.1, the Copula entropy between the comprehensive energy load and different external influencing factors, as well as the Copula entropy between different types of comprehensive energy loads, are calculated respectively. The calculation formula is as follows:
[0078]
[0079] Where, c(u1,u2,…,u N ) is the density of the Copula function, [0,1] NThe above formula represents the negative logarithmic expectation of the Copula density function, which is used to characterize the complexity of the dependencies between variables. The larger the Copula entropy, the more complex the coupling relationship between variables and the higher the uncertainty of the system.
[0080] Step 2.2: Based on the above two types of Copula entropy, construct the correlation coefficient based on Copula entropy. Specifically, let C be the Copula entropy between the comprehensive energy load and a certain external influencing factor, and M be the Copula entropy between different types of comprehensive energy loads. Then the correlation coefficient S based on Copula entropy is CEC The calculation formula is:
[0081]
[0082] Step 2.3, by setting appropriate thresholds, screen external influencing factors; when S CEC When it is greater than the threshold, it indicates that the considered external influencing factors have a significant impact on the comprehensive energy load and can be used as the key feature of the prediction model input.
[0083] In other words, the specific analysis process of the correlation analysis of multi-dimensional influencing factors of the day-ahead probability forecast of comprehensive energy load is as follows:
[0084] The Copula entropy between the comprehensive energy load and meteorological and temporal factors was calculated. A threshold of 0.1 was set to select features significantly correlated with the comprehensive energy load. Relevant meteorological factors included air pressure, dew point temperature, air temperature, and wind direction, while temporal factors included month, day, and hour. The temporal factor used data at the time of prediction, while the historical step size for meteorological factors was set to one week to fully utilize periodic information and characterize the temporal evolution of the load, thereby improving forecast accuracy.
[0085] Step 3 (S103) is to establish a comprehensive energy load day-ahead probability forecasting model based on the analysis results, which integrates multi-scale feature-time Transformer network and shape-preserving forecasting. Specifically, the model includes:
[0086] In step 3.1, the input features of the prediction model are fed into the embedding layer to generate an embedded representation of the time series. The specific process is as follows:
[0087] Assume the time series is Where L represents the length of the time series and D represents the dimension of the feature. Divide X into segments of size P, and get S = L / P segments, then we have After the embedding layer, the time series embedding representation X = {X 1 ,X 2 ,...,X S}, where d model is the dimension of the embedding layer.
[0088] In step 3.2, the multi-scale selector extracts information of different scales through time series decomposition, constructs multi-scale features, and generates weights of corresponding scales for subsequent predictive modeling. The specific process is as follows:
[0089] The multi-scale selector divides the time series X into M different segment scales through the expert system, which is defined as the set V = {V 1 ,V 2 ,...,V M The multi-scale selector uses the time decomposition module to calculate the weight of each scale and uses Fourier transform to convert the time series from the time domain to the frequency domain. In order to maintain the sparsity of the frequency domain, the first K frequency components {f1,f2,...,f K}, the specific process is as follows:
[0090] X season =IDFT({f1,f2,...,f K},φ,A)
[0091] Where IDFT stands for inverse Fourier transform, φ and A represent the phase and amplitude of each frequency, respectively. The trend component can be expressed as follows:
[0092]
[0093] Where, X res =XX season . is the pooling function with the i-th kernel, and Softmax(L(·)) is the weight of the results of different kernels. The transformed results are as follows:
[0094] X Trans =X season +X trend +X
[0095] In the process of weight generation, a noise term is introduced to increase randomness. The formula is as follows:
[0096] R(X Trans )=Softmax(X Trans W r +ε·Softplus(X Trans W noise )),ε~N(0,1)
[0097] Where R(·) represents the weight generation function, W r and W noise Parameters generated for weights.
[0098] In step 3.3, the output of the multi-scale selector is input into a multi-level attention mechanism, which includes an intra-segment feature-time joint attention mechanism and an inter-segment attention mechanism. The above attention mechanisms are combined to obtain the output of the multi-level attention mechanism. The specific process is as follows:
[0099] The intra-segment temporal attention mechanism aims to capture the local temporal pattern on a single time scale within segment i. The specific formula is:
[0100]
[0101] Where, and is the query, key, value and corresponding weight matrix in the attention mechanism corresponding to time t in segment i, d model is the dimension of the embedding layer. represents the temporal attention mechanism of all objects in segment i, and Softmax is the activation function.
[0102] The intra-segment feature attention mechanism focuses on the correlation between different features within a single time scale within segment i. The specific formula is:
[0103]
[0104] Where, and is the query, key, value and corresponding weight matrix in the attention mechanism corresponding to the feature f in the segment i. represents the feature attention weights of all time steps within segment i.
[0105] Combining the intra-fragment feature attention mechanism and the temporal attention mechanism, we get the intra-fragment feature-temporal attention mechanism. The specific formula is as follows:
[0106]
[0107] Exchange A′ intra[i]_ft The first two dimensions of For V intra[i] Each dimension of A intra[i]_ft and V intra[i] The product of the feature of segment i is the output of the temporal attention mechanism O intra[i] .
[0108] O intra[i] =A intra[i]_ft ·V intra[i]
[0109] Then the output of the attention mechanism in all segments is
[0110] The inter-fragment attention mechanism is used to capture global correlation at a single time scale, and its calculation formula is:
[0111]
[0112] Where, and The inputs of the attention mechanism between fragments are query, key and value, respectively, where d′ model =P·d model ,
[0113] The output of the j-th scale multi-level attention mechanism is obtained by fusing the intra-fragment attention mechanism with the inter-fragment attention mechanism:
[0114]
[0115] In step 3.4, the multi-scale aggregator performs weighted aggregation on the outputs of each scale to generate a representation of the multi-scale information aggregation. The specific process is as follows:
[0116] Based on multi-scale feature extraction, the multi-scale aggregator fuses the outputs of each scale using a weighted aggregation strategy. Each scale output is assigned a different weight based on its importance, effectively capturing the load fluctuation characteristics at different time scales.
[0117] The calculation formula of the weighted aggregation is as follows:
[0118]
[0119] Where M represents the total number of segments of different scales, R(·) represents the weight generation function, and T j (·) represents the function for converting different scales.
[0120] In step 3.5, the representation of the multi-scale information aggregation is processed based on the shape-preserving prediction method to obtain the day-ahead prediction probability interval of the integrated energy load cell. The specific process is as follows:
[0121] Assuming historical load data Where H is the historical data step size, and d is the dimension of multi-source load data. The predicted value of the future L step size can be expressed as:
[0122]
[0123] Where t is the current time point and L is the future prediction time range. The prediction interval of l∈{1,...,L}, where l∈{1,...,L}, makes the true value yt+l The confidence level (1-α) is set so that the prediction interval satisfies the following conditions:
[0124]
[0125] Furthermore, for regression problems, the commonly used formula for calculating the non-conformity score is:
[0126] R i =Δ(f(x (i) ),y (i) )
[0127] Where Δ is the distance metric, f is the prediction function, and x (i) is the input of the i-th sample, y (i) is the true value of the i-th sample. For the test set x (s+1) , the prediction interval can be expressed as:
[0128]
[0129] Where, Represents the predicted value of the test set. By expanding the one-dimensional non-conformity score to L dimensions, the interval of each prediction step is obtained, and the prediction interval set is expressed as:
[0130]
[0131] Where, for each step size l, the prediction interval is:
[0132]
[0133] Ensure that for each prediction step l, the true value y t+l It falls within the corresponding prediction interval with a probability of at least (1-α).
[0134] That is to say, the load forecasting process is as follows:
[0135] (1) Experimental preparation
[0136] The experimental dataset is derived from historical data on electricity, cooling, and heating loads for a building on Arizona State University's Tempe campus. Electricity load refers to the electricity consumption of electrical equipment within the building, heating load represents the demand for heating (e.g., steam / hot water systems), and cooling load reflects the demand for cooling energy (e.g., chilled water systems). The campus' energy supply comes from a central power plant, a combined cooling, heating, and power (CCHP) system, and renewable energy sources. The CCHP system uses fuels (e.g., natural gas) to generate electricity and uses the resulting waste heat to heat hot water or steam, while also providing chilled water for cooling needs. Furthermore, the building is equipped with a heating, ventilation, and air conditioning (HVAC) system that uses water cooling to regulate indoor temperature. Outdoor meteorological data for the area, including atmospheric temperature, atmospheric pressure, dew point temperature, wind speed, and direction, can be obtained from the National Climatic Data Center website. Data was collected from 01:00 on January 1, 2017, to 23:00 on January 31, 2021, with a sampling frequency of once per hour, totaling 35,808 data sets. The dataset was divided into training, validation, and test sets in a 6:2:2 ratio. The prediction model parameters used a batch size of 80, an iteration count of 100, a mean squared error loss function, a loss rate of 0.01, and an initial learning rate of 0.0001.
[0137] (2) Prediction performance evaluation
[0138] In order to evaluate the prediction performance, Pinball loss function and Winkler score were selected to evaluate the effect of the probability prediction model.
[0139] The Pinball loss calculation formula is as follows:
[0140]
[0141] Where, represents the qth quantile load forecast value of the i-th sample, y i is the actual observed load value. The smaller the pinball loss score, the better the prediction effect.
[0142] The Winkler score calculation formula is as follows:
[0143]
[0144] Where, δ i =U i -L i Indicates the prediction interval width of the i-th sample, Ui and L i are the upper and lower bounds, respectively. When the actual value falls within the prediction interval, the Winkler score depends only on the response width; otherwise, a penalty term is added to reflect the deviation of the uncovered actual value. The smaller the Winkler score, the better the prediction effect.
[0145] In this embodiment, six prediction models are used to predict the load for the next day, including the method in this paper (denoted as MSBFTformer-CP) and the multi-scale feature-time Transformer and quantile method (denoted as MSBFTformer-MQ), the multi-scale feature-time Transformer and maximum likelihood estimation method (denoted as MSBFTformer-MLF), the multi-scale feature-time Transformer and maximum likelihood estimation method (denoted as MSBFTformer-Bayes), Transformer and shape-preserving prediction (denoted as Transformer-CP), Convformer and shape-preserving prediction (denoted as Convformer-CP), and Autoformer and shape-preserving prediction (denoted as Autoformer-CP). The prediction results are compared in Table 1.
[0146] Table 1
[0147]
[0148]
[0149] In summary, the present embodiment proposes a comprehensive energy load day-ahead probabilistic forecasting model. First, the Copula entropy correlation coefficient is used to analyze the correlations between multidimensional influencing factors, fully exploring the coupling relationships between multidimensional data. Then, based on this, a forecasting model is proposed that integrates a multi-scale feature-time Transformer network with shape-preserving prediction to improve forecasting accuracy. This embodiment of the present invention can effectively improve the accuracy of comprehensive energy load day-ahead probabilistic forecasting.
[0150] Device Example 1
[0151] According to an embodiment of the present invention, a device for establishing a comprehensive energy load day-ahead probability prediction model is provided. Figure 2 Schematic diagram of a device for establishing a comprehensive energy load day-ahead probability forecasting model according to an embodiment of the present invention. Figure 2 As shown, the apparatus for establishing a comprehensive energy load day-ahead probability forecasting model according to an embodiment of the present invention specifically includes: a pre-processing module, an influencing factor analysis module, and a forecasting model building module, thereby obtaining the result of comprehensive energy load forecasting. Specifically:
[0152] The pre-processing module 60 is used to pre-process the collected comprehensive energy load historical data. The pre-processing module 60 is specifically used to:
[0153] The collected historical data of integrated energy system load is cleaned and normalized.
[0154] The multi-dimensional influencing factor correlation analysis module 62 is used to perform multi-dimensional influencing factor correlation analysis on the pre-processed data based on the Copula entropy correlation coefficient to obtain analysis results. The multi-dimensional influencing factor correlation analysis module 62 is specifically used to:
[0155] The Copula entropy between the comprehensive energy load and different external influencing factors, as well as the Copula entropy between different types of comprehensive energy loads, are calculated separately. The calculation formula is as follows:
[0156]
[0157] Where, c(u1,u2,...,u N ) is the density of the Copula function, [0,1] N The above formula represents the negative logarithmic expectation of the Copula density function, which is used to characterize the complexity of the dependencies between variables. The larger the Copula entropy, the more complex the coupling relationship between variables and the higher the uncertainty of the system.
[0158] Based on the above two types of Copula entropy, a correlation coefficient based on Copula entropy is constructed. Specifically, let C be the Copula entropy between the comprehensive energy load and a certain external influencing factor, and M be the Copula entropy between different types of comprehensive energy loads. Then the correlation coefficient S based on Copula entropy is CEC The calculation formula is:
[0159]
[0160] By setting appropriate thresholds, external influencing factors are screened; when S CEC When it is greater than the threshold, it indicates that the considered external influencing factors have a significant impact on the comprehensive energy load and can be used as the key feature of the prediction model input.
[0161] The prediction model building module 64 is used to build a comprehensive energy load day-ahead probability prediction model that integrates multi-scale feature-time Transformer network and shape-preserving prediction based on the analysis results. The prediction model building module 64 is specifically used to:
[0162] Feed the input features of the prediction model into the embedding layer to generate an embedded representation of the time series;
[0163] The multi-scale selector extracts information of different scales through time series decomposition, constructs multi-scale features, and generates weights of corresponding scales for subsequent predictive modeling;
[0164] Input the output of the multi-scale selector into a multi-level attention mechanism, which includes an intra-segment feature-time joint attention mechanism and an inter-segment attention mechanism, and fuse the above attention mechanisms to obtain the output result of the multi-level attention mechanism;
[0165] The multi-scale aggregator performs weighted aggregation on the outputs of each scale to generate a representation of multi-scale information aggregation;
[0166] Based on the shape-preserving prediction method, the representation of the above multi-scale information aggregation is processed to finally obtain the day-ahead prediction probability interval result of the comprehensive energy load cell.
[0167] Furthermore, the input features of the prediction model are fed into the embedding layer to generate an embedded representation of the time series, specifically including:
[0168] Assume the time series is Where L represents the length of the time series and D represents the dimension of the feature. Divide X into segments of size P, and get S = L / P segments, then we have After the embedding layer, the time series embedding representation X = {X 1 ,X 2 ,...,X S}, where d model is the dimension of the embedding layer.
[0169] Furthermore, the multi-scale selector extracts information of different scales through time series decomposition, constructs multi-scale features, and generates weights of corresponding scales for subsequent predictive modeling, specifically including:
[0170] The multi-scale selector divides the time series X into M different segment scales through the expert system, which is defined as the set V = {V 1 ,V 2 ,...,V M The multi-scale selector uses the time decomposition module to calculate the weight of each scale and uses Fourier transform to convert the time series from the time domain to the frequency domain. In order to maintain the sparsity of the frequency domain, the first K frequency components {f1,f2,...,f K}, the specific process is as follows:
[0171] X season =IDFT({f1,f2,...,f K},φ,A)
[0172] Where IDFT stands for inverse Fourier transform, φ and A represent the phase and amplitude of each frequency, respectively. The trend component can be expressed as follows:
[0173]
[0174] Where, X res =XX season . is the pooling function with the i-th kernel, and Softmax(L(·)) is the weight of the results of different kernels. The transformed results are as follows:
[0175] X Trans =X season +X trend +X
[0176] In the process of weight generation, a noise term is introduced to increase randomness. The formula is as follows:
[0177] R(X Trans )=Softmax(X Trans W r +ε·Softplus(X Trans W noise )),ε~N(0,1)
[0178] Where R(·) represents the weight generation function, W r and W noise Parameters generated for weights.
[0179] Furthermore, the output of the multi-scale selector is input into a multi-level attention mechanism, which includes an intra-segment feature-time joint attention mechanism and an inter-segment attention mechanism, and the above attention mechanisms are integrated to obtain the output result of the multi-level attention mechanism, specifically including:
[0180] The intra-segment temporal attention mechanism aims to capture the local temporal pattern on a single time scale within segment i. The specific formula is:
[0181]
[0182] Where, and is the query, key, value and corresponding weight matrix in the attention mechanism corresponding to time t in segment i, d model is the dimension of the embedding layer. represents the temporal attention mechanism of all objects in segment i, and Softmax is the activation function.
[0183] The intra-segment feature attention mechanism focuses on the correlation between different features within a single time scale within segment i. The specific formula is:
[0184]
[0185] Where, and is the query, key, value and corresponding weight matrix in the attention mechanism corresponding to the feature f in the segment i. represents the feature attention weights of all time steps within segment i.
[0186] Combining the intra-fragment feature attention mechanism and the temporal attention mechanism, we get the intra-fragment feature-temporal attention mechanism. The specific formula is as follows:
[0187]
[0188] Exchange A′ intra[i]_ft The first two dimensions of For V intra[i] Each dimension of A intra[i]_ft and V intra[i] The product of the feature of segment i is the output of the temporal attention mechanism O intra[i] .
[0189] O intra[i] =A intra[i]_ft ·V intra[i]
[0190] Then the output of the attention mechanism in all segments is
[0191] The inter-fragment attention mechanism is used to capture global correlation at a single time scale, and its calculation formula is:
[0192]
[0193] Where, and The inputs of the attention mechanism between fragments are query, key and value, respectively, where d′ model =P·d model ,
[0194] The output of the j-th scale multi-level attention mechanism is obtained by fusing the intra-fragment attention mechanism with the inter-fragment attention mechanism:
[0195]
[0196] Furthermore, the multi-scale aggregator performs weighted aggregation on the outputs of each scale to generate a representation of the multi-scale information aggregation, specifically including:
[0197] Based on multi-scale feature extraction, the multi-scale aggregator fuses the outputs of each scale using a weighted aggregation strategy. Each scale output is assigned a different weight based on its importance, effectively capturing the load fluctuation characteristics at different time scales.
[0198] The calculation formula of the weighted aggregation is as follows:
[0199]
[0200] Where M represents the total number of segments of different scales, R(·) represents the weight generation function, and T j (·) represents the function for converting different scales.
[0201] Furthermore, the shape-preserving prediction method processes the representation of the above multi-scale information aggregation to finally obtain the day-ahead prediction probability interval result of the integrated energy load unit, which specifically includes:
[0202] Assuming historical load data Where H is the historical data step size, and d is the dimension of multi-source load data. The predicted value of the future L step size can be expressed as:
[0203]
[0204] Where t is the current time point and L is the future prediction time range. The prediction interval of l∈{1,...,L}, where l∈{1,...,L}, makes the true value y t+l The confidence level (1-α) is set so that the prediction interval satisfies the following conditions:
[0205]
[0206] Furthermore, for regression problems, the commonly used formula for calculating the non-conformity score is:
[0207] R i =Δ(f(x (i) ),y (i) )
[0208] Where Δ is the distance metric, f is the prediction function, and x (i) is the input of the i-th sample, y (i) is the true value of the i-th sample. For the test set x (s+1) , the prediction interval can be expressed as:
[0209]
[0210] Where, Represents the predicted value of the test set. By expanding the one-dimensional non-conformity score to L dimensions, the interval of each prediction step is obtained, and the prediction interval set is expressed as:
[0211]
[0212] Where, for each step size l, the prediction interval is:
[0213]
[0214] Ensure that for each prediction step l, the true value y t+l It falls within the corresponding prediction interval with a probability of at least (1-α).
[0215] The embodiment of the present invention is an apparatus embodiment corresponding to the above-mentioned method embodiment. The specific operations of each module can be understood by referring to the description of the method embodiment, which will not be repeated here.
[0216] Device Example 2
[0217] An embodiment of the present invention provides an electronic device, such as Figure 3 As shown, it includes: a memory 70, a processor 72 and a computer program stored in the memory 70 and executable on the processor 72. When the computer program is executed by the processor 72, the steps described in the method embodiment are implemented.
[0218] Device Example 3
[0219] An embodiment of the present invention provides a computer-readable storage medium, on which a program for implementing information transmission is stored. When the program is executed by the processor 72, the steps described in the method embodiment are implemented.
[0220] The computer-readable storage medium in this embodiment includes, but is not limited to, ROM, RAM, magnetic disk, or optical disk.
[0221] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0222] In the 1930s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly done using "logic compiler" software. This is similar to the software compiler used when developing programs. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.
[0223] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules that implement the method and structures within the hardware component.
[0224] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0225] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing the embodiments of this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0226] Those skilled in the art will appreciate that one or more embodiments of this specification may be provided as a method, system, or computer program product. Thus, one or more embodiments of this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0227] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0228] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0229] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0230] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0231] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0232] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0233] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0234] One or more embodiments of this specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0235] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.
[0236] The foregoing description is merely an example of the present invention and is not intended to limit the present invention. Persons skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims herein.
Claims
1. A method for establishing a comprehensive energy load day-ahead probability forecasting model, characterized in that: include: S101 pre-processes the collected comprehensive energy load historical data; S102 performs a multi-dimensional influencing factor correlation analysis on the pre-processed data based on the Copula entropy correlation coefficient to obtain an analysis result; S103 establishes a comprehensive energy load day-ahead probability forecasting model based on the analysis results, which integrates a multi-scale feature-time Transformer network and shape-preserving prediction; wherein the multi-scale feature-time Transformer network is used to extract multi-scale features and model time dependencies, and the shape-preserving prediction is used to generate load probability interval prediction results.
2. The method for establishing a comprehensive energy load day-ahead probability forecasting model according to claim 1, characterized in that: Based on the Copula entropy correlation coefficient, the pre-processed data is subjected to a multi-dimensional correlation analysis of influencing factors, and the analysis results are obtained, including: In order to solve the problem that the original Copula entropy has inconsistent dimensions and an unfixed numerical range, and is difficult to be directly used for horizontal comparison and screening of the correlation strengths between different external influencing factors, a correlation coefficient based on Copula entropy is constructed, which is defined as follows: Where C represents the Copula entropy between the comprehensive energy load and the external influencing factors, and M represents the Copula entropy between different types of comprehensive energy loads. Based on the correlation coefficient, external influencing factors are screened by setting a threshold. When the correlation coefficient is greater than the threshold, it is determined that the external influencing factor has a significant impact on the comprehensive energy load and is used as an input feature of the prediction model to improve the prediction accuracy and robustness of the model.
3. The method for establishing a comprehensive energy load day-ahead probability forecasting model according to claim 1, characterized in that: Based on the analysis results, a comprehensive energy load day-ahead probabilistic forecasting model is established by integrating multi-scale feature-time Transformer network and shape-preserving forecasting. Specifically, it includes: Feed the input features of the prediction model into the embedding layer to generate an embedded representation of the time series; The multi-scale selector extracts information of different scales through time series decomposition, constructs multi-scale features, and generates weights of corresponding scales for subsequent predictive modeling; Input the output of the multi-scale selector into a multi-level attention mechanism, which includes an intra-segment feature-time joint attention mechanism and an inter-segment attention mechanism, and fuse the above attention mechanisms to obtain the output result of the multi-level attention mechanism; The multi-scale aggregator performs weighted aggregation on the outputs of each scale to generate a representation of multi-scale information aggregation; Based on the shape-preserving prediction method, the representation of the above multi-scale information aggregation is processed to finally obtain the day-ahead prediction probability interval result of the comprehensive energy load cell.
4. The method for establishing a comprehensive energy load day-ahead probability forecasting model according to claim 3, characterized in that: The multi-scale selector extracts information of different scales through time series decomposition, constructs multi-scale features, and generates weights of corresponding scales for subsequent predictive modeling. Specifically, it includes: The multi-scale selector divides the time series X into M different segment scales through the expert system, which is defined as the set V = {V 1 ,V 2 ,...,V M The multi-scale selector uses the time decomposition module to calculate the weight of each scale and uses Fourier transform to convert the time series from the time domain to the frequency domain. In order to maintain the sparsity of the frequency domain, the first K frequency components with the largest amplitude {f1,f2,...,f K }, the specific process is as follows: X season =IDFT({f1,f2,...,f K },φ,A) Where IDFT stands for inverse Fourier transform, φ and A represent the phase and amplitude of each frequency, respectively, and the trend component is expressed as follows: Where, X res =XX season ; is the pooling function with the i-th kernel, Softmax(L(·)) is the weight of the results of different kernels, and the transformed results are as follows: X Trans =X season +X trend +X In the process of weight generation, a noise term is introduced to increase randomness. The formula is as follows: R(X Trans )=Softmax(X Trans ·W r +ε·Softplus(X Trans ·W noise )),ε~N(0,1) Where R(·) represents the weight generation function, W r and W noise Parameters generated for weights.
5. The method for establishing a comprehensive energy load day-ahead probability forecasting model according to claim 3, characterized in that: The output of the multi-scale selector is input into a multi-level attention mechanism, which includes an intra-segment feature-time joint attention mechanism and an inter-segment attention mechanism. The above attention mechanisms are combined to obtain the output of the multi-level attention mechanism, which specifically includes: The intra-segment temporal attention mechanism aims to capture the local temporal pattern within segment i at a single time scale. The specific formula is: Where, and is the query, key, value and corresponding weight matrix in the attention mechanism corresponding to time t in segment i, d model is the dimension of the embedding layer; represents the temporal attention mechanism of all objects in segment i, and Softmax is the activation function; The intra-segment feature attention mechanism focuses on the correlation between different features within a single time scale within segment i. The specific formula is: Where, and is the query, key, value and corresponding weight matrix in the attention mechanism corresponding to the feature f in the segment i; represents the feature attention weights of all time steps within segment i; Combining the intra-fragment feature attention mechanism and the temporal attention mechanism, we get the intra-fragment feature-temporal attention mechanism. The specific formula is as follows: Exchange A′ intra[i]_ft The first two dimensions of For V intra[i] Each dimension of A intra[i]_ft and V intra[i] The product of the feature of segment i is the output of the temporal attention mechanism O intra[i] ; About intra[i] =A intra[i]_ft ·In intra[i] Then the output of the attention mechanism in all segments is The inter-fragment attention mechanism is used to capture global correlation at a single time scale, and its calculation formula is: Where, and The input of the attention mechanism between fragments is query, key and value, respectively, where d′ model =P·d model , The output of the j-th scale multi-level attention mechanism is obtained by fusing the intra-fragment attention mechanism with the inter-fragment attention mechanism:
6. The method for establishing a comprehensive energy load day-ahead probability forecasting model according to claim 3, characterized in that: The multi-scale aggregator performs weighted aggregation on the outputs of each scale to generate a representation of the multi-scale information aggregation, which includes: The multi-scale aggregator generates an aggregated representation of multi-scale information by weighted aggregation of outputs of different scales. The specific formula is: Where M represents the total number of segments of different scales, R(·) represents the weight generation function, and T j (·) represents the function for converting different scales.
7. The method for establishing a comprehensive energy load day-ahead probability forecasting model according to claim 3, characterized in that: Based on the shape-preserving prediction method, the above multi-scale information aggregation representation is processed to finally obtain the day-ahead prediction probability interval results of the integrated energy load cell, including: Assuming historical load data Where H is the historical data step size, d is the dimension of multi-source load data, and the prediction interval for each future step size is expressed as: Where, is the score that does not meet the requirements, L is the future prediction time range; by shape-preserving prediction, we ensure that for each prediction step l, the true value y t+l It falls within the corresponding prediction interval with a probability of at least (1-α).
8. A device for establishing a comprehensive energy load day-ahead probability forecasting model, characterized in that: include: A preprocessing module is used to preprocess the collected comprehensive energy load historical data; The multi-dimensional influencing factor correlation analysis module is used to perform multi-dimensional influencing factor correlation analysis on the pre-processed data based on the Copula entropy correlation coefficient to obtain the analysis results; A prediction model module is established to establish a comprehensive energy load day-ahead probability prediction model that integrates a multi-scale feature-time Transformer network and shape-preserving prediction based on the analysis results. The multi-scale feature-time Transformer network is used to extract multi-scale features and model time dependencies, and the shape-preserving prediction is used to generate load probability interval prediction results.
9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, the steps of the method for establishing a comprehensive energy load day-ahead probability prediction model are implemented as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores an implementation program for information transmission, and when the program is executed by the processor, the steps of the method for establishing a comprehensive energy load day-ahead probability prediction model as described in any one of claims 1 to 7 are implemented.