Regional energy system multi-element load prediction method and device based on big data
By screening the target influencing factors with high correlation and using LSTM neural network and Attention mechanism to train the model, the problem of insufficient prediction accuracy in the prediction of multi-load prediction in regional energy systems is solved, and accurate prediction of electrical load, thermal load and cold load is achieved.
Patent Information
- Application Number
- CN202510287829.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art fails to fully consider the coupling characteristics of multi-energy systems and the dynamic correlation analysis of multi-dimensional data in the multi-load prediction of regional energy systems, resulting in insufficient prediction accuracy and adaptability, and it is difficult to deal with the complex coupling relationship of multi-energy systems.
By obtaining candidate influencing factors, calculating the correlation coefficient and weight coefficient, filtering out the target influencing factors with high correlation, and using LSTM neural network and Attention mechanism to train a multivariate load prediction model to improve prediction accuracy.
It realizes accurate prediction of electric load, thermal load and cooling load, solves the complexity problem of multi-energy coupling relationship, and improves prediction accuracy and reliability.
Smart Images

Figure CN120258202A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of load forecasting, and particularly to a multi-energy load forecasting method and device for a regional energy system based on big data. Background Art
[0002] Multi-energy load forecasting for a regional energy system based on big data is a key link in the optimal scheduling and stable operation of the regional energy system. Especially in a system where energy forms such as electricity, heat, and cold are closely coupled, its accuracy is directly related to the economy and safety of the system. Through accurate short-term load forecasting, reliable basis can be provided for energy demand planning, capacity configuration, and real-time scheduling, thereby improving the operating efficiency of the energy system. However, due to the complex coupling relationship between loads in the regional energy system, existing load forecasting technologies are difficult to comprehensively reflect the dynamic correlation between multi-energy systems, posing higher requirements for forecasting accuracy and response speed. With the development of big data technology and artificial intelligence, load forecasting methods have gradually evolved from traditional classical methods such as time series analysis and regression analysis to directions such as neural networks, support vector machines, and fuzzy inference systems based on machine learning. Although these technologies show good effects in dealing with single load forecasting, there are still many limitations when dealing with the scenario of multi-energy coupling in the regional energy system.
[0003] There are obvious deficiencies in the existing technology for multi-energy load forecasting of a regional energy system based on big data. The most prominent problem is that the forecasting models are mostly designed for a single energy type, and they are unable to cope with the complex coupling relationship of the multi-energy system, making it difficult to fully explore the correlation characteristics between energy loads. The root cause of these problems is that the existing technology fails to comprehensively consider the coupling characteristics of the multi-energy system and the dynamic correlation analysis of multi-dimensional data during model design, lacking effective screening and optimization of key influencing factors, resulting in the forecasting accuracy and adaptability being difficult to meet the actual requirements. Summary of the Invention
[0004] Embodiments of the present invention provide a multi-energy load forecasting method and device for a regional energy system based on big data. By implementing the present invention, the accuracy of multi-energy load forecasting of a regional energy system based on big data can be improved.
[0005] An embodiment of the present invention provides a multi-energy load forecasting method for a regional energy system based on big data, including: obtaining candidate influencing factors of the area to be forecast; wherein, the candidate influencing factors include historical electric load, historical heat load, historical cold load, humidity, temperature, calendar information, air pressure, and wind speed; Calculating and generating correlation coefficients between the candidate influencing factors according to the candidate influencing factors; Calculating and generating weight coefficients of the candidate influencing factors according to the candidate influencing factors; Generate the correlation degrees among the candidate influencing factors according to the correlation coefficients and the weight coefficients of the candidate influencing factors; Use the candidate influencing factors with the correlation degrees greater than a preset threshold as the target influencing factors; Input the target influencing factors into a preset multivariate load prediction model, so that the multivariate load prediction model outputs the electrical load, heat load, and cooling load of the target day in the area to be predicted according to the target influencing factors.
[0006] Further, calculating and generating the correlation coefficients among the candidate influencing factors according to the candidate influencing factors includes: Based on the candidate influencing factors, calculate and generate the marginal cumulative distributions of the candidate influencing factors based on the kernel density estimation algorithm; according to the marginal cumulative distributions of the candidate influencing factors, calculate the correlation function parameter estimation values of the candidate influencing factors based on the distribution maximum likelihood estimation algorithm; Based on the correlation function parameter estimation values of the candidate influencing factors, determine the optimal correlation function model of the candidate influencing factors based on the Euclidean norm minimization algorithm; Calculate and generate the correlation coefficients among the candidate influencing factors according to the optimal correlation function model of the candidate influencing factors.
[0007] Further, calculating and generating the weight coefficients of the candidate influencing factors according to the candidate influencing factors includes: constructing a comparison matrix among the candidate influencing factors based on a preset relative importance scale according to the candidate influencing factors; calculating and generating the weight coefficients of the candidate influencing factors based on the geometric mean algorithm according to the comparison matrix.
[0008] Further, calculate and generate the marginal cumulative distribution of each candidate influencing factor through the following formula: where f(x) is the probability density function of the candidate influencing factor x; n is the total number of sample points of the candidate influencing factor x; h is the window frame constant coefficient; K(·) is the kernel function; x is the candidate influencing factor; x i is the value of the i-th sample point among the candidate influencing factors.
[0009] Further, train the preset multivariate load prediction model in the following manner: Obtain the target influencing factors and the corresponding historical multivariate load prediction labels of historical typical days; the historical multivariate load prediction labels are the electrical load, heat load, and cooling load of historical typical days; Construct a training set according to the target influencing factors and the corresponding historical multivariate load prediction labels of historical typical days; Randomly divide the training set into several batches of training samples according to a preset batch size; Input the training samples of each batch into the multi-load prediction model in sequence to train the multi-load prediction model until the preset number of training times is reached; among them, when the multi-load prediction model receives a batch of training samples, it outputs the predicted electrical load, predicted heat load, and predicted cooling load corresponding to the training samples; according to the predicted electrical load, predicted heat load, and predicted cooling load and the corresponding historical multi-load prediction labels, calculate the loss function value through the loss function; perform Based on the above method item embodiments, the present invention correspondingly provides device item embodiments.
[0010] An embodiment of the present invention provides a multi-load prediction device for a regional energy system based on big data, including: a data acquisition module, a correlation coefficient calculation module, a weight coefficient calculation module, a target influencing factor determination module, and a multi-load prediction module; The data acquisition module is used to acquire candidate influencing factors of the area to be predicted; among them, the candidate influencing factors include historical electrical load, historical heat load, historical cooling load, humidity, temperature, calendar information, air pressure, and wind speed; The correlation coefficient calculation module is used to calculate and generate the correlation coefficients between the candidate influencing factors according to the candidate influencing factors; The weight coefficient calculation module is used to calculate and generate the weight coefficients of the candidate influencing factors according to the candidate influencing factors; the target influencing factor determination module is used to generate the association degree between the candidate influencing factors according to the correlation coefficients between the candidate influencing factors and the weight coefficients of the candidate influencing factors; use the candidate influencing factors with the association degree greater than the preset threshold as the target influencing factors; The multi-load prediction module is used to input the target influencing factors into a preset multi-load prediction model, so that the multi-load prediction model outputs the electrical load, heat load, and cooling load of the target day of the area to be predicted according to the target influencing factors.
[0011] Further, for the multi-load prediction device of the regional energy system based on big data, the correlation coefficient calculation module includes: an edge cumulative distribution calculation unit, a correlation function parameter estimate calculation unit, an optimal correlation function model determination unit, and a correlation coefficient determination unit; The edge cumulative distribution calculation unit is used to calculate and generate the edge cumulative distributions of the candidate influencing factors based on the kernel density estimation algorithm according to the candidate influencing factors; The correlation function parameter estimation value calculation unit is used to calculate the estimated values of the correlation function parameters of each candidate influencing factor based on the marginal cumulative distribution of each candidate influencing factor and the distribution maximum likelihood estimation algorithm; The optimal correlation function model is used to determine the optimal correlation function model of each candidate influencing factor based on the Euclidean norm minimization algorithm according to the estimated values of the correlation function parameters of each candidate influencing factor; The correlation coefficient determination unit is used to calculate and generate the correlation coefficients between each candidate influencing factor according to the optimal correlation function model of each candidate influencing factor.
[0012] Further, for the device for multi - load prediction of regional energy systems based on big data, the weight coefficient calculation module includes: a comparison matrix determination unit and a weight coefficient determination unit; The comparison matrix determination unit is used to construct a comparison matrix between candidate influencing factors based on a preset relative importance scale according to the candidate influencing factors; The weight coefficient determination unit is used to calculate and generate the weight coefficients of each candidate influencing factor based on the geometric mean algorithm according to the comparison matrix.
[0013] Further, the marginal cumulative distribution of each candidate influencing factor is calculated through the following formula: where f(x) is the probability density function of the candidate influencing factor x; n is the total number of sample points of the candidate influencing factor x; h is the window frame constant coefficient; K(·) is the kernel function; x is the candidate influencing factor; x i is the value of the i - th sample point among the candidate influencing factors.
[0014] Further, the preset multi - load prediction model is trained in the following manner: Obtain the target influencing factors and corresponding historical multi - load prediction labels of historical typical days; the historical multi - load prediction labels are the electrical load, heat load, and cooling load of historical typical days; Construct a training set according to the target influencing factors and corresponding historical multi - load prediction labels of historical typical days; Randomly divide the training set into several batches of training samples according to a preset batch size; Input the training samples of each batch into the multivariate load prediction model in sequence to train the multivariate load prediction model until the preset number of training times is reached; wherein, when the multivariate load prediction model receives each batch of training samples, it outputs the predicted electrical load, predicted heat load, and predicted cooling load corresponding to the training samples; calculate the loss function value through the loss function according to the predicted electrical load, predicted heat load, predicted cooling load, and the corresponding historical multivariate load prediction labels; update the multivariate load prediction model according to the loss function value.
[0015] Compared with the prior art, the present invention has the following beneficial effects: The embodiment of the present invention provides a method and device for multivariate load prediction of a regional energy system based on big data. The method realizes accurate prediction of electrical load, heat load, and cooling load by determining the target influencing factors of the area to be predicted and inputting them into the multivariate load prediction model. The process of determining the target influencing factors includes screening out the key influencing factors with high correlation from the candidate influencing factors, thus effectively solving the problems of the complexity of the multi-energy coupling relationship and insufficient prediction accuracy. The present invention calculates the correlation and weight coefficients of the candidate influencing factor data, screens out the key factors that have a significant impact on load prediction, solves the problem that the prior art does not fully consider the dynamic correlation analysis of multi-dimensional data, and uses the key factors for load prediction, thereby improving the accuracy and reliability of the multivariate load prediction of the regional energy system based on big data. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 FIG. is a flowchart of a method for multivariate load prediction of a regional energy system based on big data provided by an embodiment of the present invention.
[0017] Figure 2 FIG. is a structural diagram of an LSTM system provided by an embodiment of the present invention.
[0018] Figure 3 FIG. is a structural diagram of an Attention mechanism provided by an embodiment of the present invention.
[0019] Figure 4 FIG. is a structural schematic diagram of a device for multivariate load prediction of a regional energy system based on big data provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0021] As Figure 1 shown, an embodiment of the present invention provides a method for multi - load prediction of a regional energy system based on big data, which at least includes the following steps: Step S1, obtain candidate influencing factors of the area to be predicted.
[0022] Specifically, the candidate influencing factors include historical electricity load, historical heat load, historical cooling load, humidity, temperature, calendar information, air pressure, and wind speed. The historical electricity load, historical heat load, and historical cooling load can be load data of a certain past time period or a certain typical day in the past, and can be adjusted according to actual needs.
[0023] Optionally, the candidate influencing factor data can also include solar radiation, rainfall, cloud cover, population density, etc., and can be flexibly adjusted according to the actual situation.
[0024] Step S2, calculate the correlation coefficients between the candidate influencing factors according to the candidate influencing factors.
[0025] In a preferred embodiment, calculating the correlation coefficients between the candidate influencing factors according to the candidate influencing factors includes: Based on the candidate influencing factors, calculate the marginal cumulative distribution of each candidate influencing factor based on the kernel density estimation algorithm; based on the marginal cumulative distribution of each candidate influencing factor, calculate the parameter estimation value of the correlation function of each candidate influencing factor based on the distribution maximum likelihood estimation algorithm; Based on the parameter estimation value of the correlation function of each candidate influencing factor, determine the optimal correlation function model of each candidate influencing factor based on the Euclidean norm minimization algorithm; Calculate the correlation coefficients between the candidate influencing factors according to the optimal correlation function model of each candidate influencing factor.
[0026] Optionally, the marginal cumulative distribution of each candidate influencing factor is calculated by the following formula: where f(x) is the probability density function of the candidate influencing factor x; n is the total number of sample points of the candidate influencing factor x; h is the window frame constant coefficient; K(·) is the kernel function; x is the candidate influencing factor; x i is the value of the i - th sample point in the candidate influencing factors.
[0027] It should be noted here that an embodiment of the present invention measures the correlation of each candidate influencing factor based on the Copula theory.
[0028] Use the correlation analysis method based on Copula theory to calculate the correlation coefficients between various types of loads and each candidate influencing factor. Although the widely used Pearson correlation analysis method can effectively describe the linear correlation characteristics between variables, its ability to describe non-linear correlation characteristics is weak and it is not suitable for describing the correlation relationships between candidate influencing factors in a district integrated energy system (DIES). The correlation analysis method based on Copula theory has strong processing ability for non-linear correlation characteristics and can more flexibly describe the correlations between candidate influencing factors in a district integrated energy system.
[0029] First, the present invention uses a non-parametric method based on kernel density estimation to determine the marginal cumulative distributions of each candidate influencing factor. Then comes the parameter estimation of the Copula function model: Given that the marginal cumulative distributions of each candidate influencing factor have been obtained in the previous step, the present invention selects the stepwise maximum likelihood estimation to solve the parameters to be determined in the marginal cumulative distribution and each Copula function model respectively. Specifically, first estimate the parameters θ1 and θ2 of the marginal cumulative distribution: Among them, are the maximum likelihood values of θ1 and θ2; f(x, θ1) and g(y, θ2) are the probability density functions of the marginal cumulative distribution functions F(x, θ1) and G(y, θ2) of the random variables X and Y respectively. Then substitute into the following formula to obtain the parameter α to be determined in the Copula function: Among them, is the Copula function for which parameter estimation is to be performed.
[0030] After that is to determine the optimal Copula function model: Since there are many types of Copula function models and different Copula function models have different description results for the correlation characteristics of the above coupling relationship, the correlation analysis based on Copula theory needs to select the optimal Copula function model according to the fitting results of various Copula function models. In this section, the Euclidean norm minimum method is used to test the goodness of fit of the sample data of each Copula function model to determine the optimal Copula function. Specifically, first calculate the empirical Copula function of each type of load data sample and the above candidate influencing factor data sample as follows: Among them, u, v ∈ [0, 1]; (x1, y1), (x2, y2), …, (x n , y n ) are the sample points taken from the sample population (X, Y); F n(x), G n (y) is the empirical cumulative distribution function of X and Y; I is the indicator function. If F n (x i ) ≤ u, then the value of I[F n (x i ) ≤ u] is 1, otherwise the value is 0.
[0031] Then calculate the Euclidean norm (L2 norm) between the Copula function to be tested and the sample empirical Copula function. The calculation method is shown in the following formula: where, u i = F n (x i ), v i = G n (y i ), i = 1, 2,..., n.
[0032] In a preferred embodiment, the specific calculation results of the Euclidean norms for the goodness-of-fit tests of various Copula function models to the above candidate influencing factors are shown in the following table. Among them, N, G, t, F, and C in the table respectively represent five Copula function models: Gaussian - Copula, Gumbel - Copula, t - Copula, Frank - Copula, and Clayton - Copula; E, C, H, H’, T, C’, P, and W respectively represent eight candidate influencing factors: electrical load, cooling load, heating load, humidity, temperature, calendar information, air pressure, and wind speed. Finally, calculate the correlation coefficient index: The correlation measure indexes derived based on the Copula theory mainly include two types: Kendall rank correlation coefficient and Spearman rank correlation coefficient. Compared with the Kendall rank correlation coefficient, the Spearman rank correlation coefficient has a smaller relative error when describing the correlation characteristics between multiple loads and between various loads and other influencing factors. Therefore, the present invention uses the Spearman rank correlation coefficient to describe the correlation characteristics between the candidate influencing factors. The calculation formula of the Spearman rank correlation coefficient is as follows: ρ = 12∫0∫0C(u, v)dudov - 3 Among them, ρ is the Spearman rank correlation coefficient, which is used to describe the order correlation between two random variables. The value range is [-1, 1]. A value close to 1 indicates a strong positive correlation, a value close to -1 indicates a strong negative correlation, and a value close to 0 indicates no correlation. C(u, v) is the Copula function, which is used to describe the structural characteristics of the joint distribution of two variables. Here, it takes (u, v) as the input and measures the dependence relationship between the two variables without being affected by the marginal distribution.
[0033] In a preferred embodiment, the Spearman rank correlation coefficient indexes of each influencing factor derived from the Euclidean norm calculation results are shown in the following table. Among them, E, C, H, H', T, C’, P, and W in the table respectively represent eight candidate influencing factors: electric load, cooling load, heating load, humidity, temperature, calendar information, air pressure, and wind speed. The results in the table record the absolute values of the actual calculated values. The Spearman rank correlation coefficient between each candidate influencing factor is the correlation coefficient between the candidate influencing factors. It should be noted here that the first three rows of this table record the Spearman rank correlation coefficients between each candidate influencing factor under various load forecasting subtasks, and the last row records the Spearman rank correlation coefficients of each influencing factor under the multi-load forecasting task of electricity, cooling, and heating. The calculation method is to take the average value of the Spearman rank correlation coefficients of each influencing factor under various load forecasting subtasks. Observing the calculation results of this table, it can be found that due to the difficulty of reasonably digitalizing calendar information, the correlation between the two influencing factors of air pressure and wind speed in meteorological information and multi-load forecasting is much greater than that of calendar information, which does not conform to real-world experience. From this, it can be seen that if only relying on the rank correlation coefficients of each influencing factor derived from the Copula function model, sometimes due to data quality and other data-related problems, the calculation results cannot accurately provide a basis for determining input features. In view of this, this article uses AHP to weight each influencing factor from the index dimension to handle the above situation.
[0034] Step S3: Calculate and generate the weight coefficients of each candidate influencing factor according to the candidate influencing factors.
[0035] In a preferred embodiment, the calculating and generating the weight coefficients of each candidate influencing factor according to the candidate influencing factors includes: Construct a comparison matrix between the candidate influencing factors based on a preset relative importance scale according to the candidate influencing factors; calculate and generate the weight coefficients of each candidate influencing factor based on the geometric mean algorithm according to the comparison matrix.
[0036] It should be noted here that in a preferred embodiment, the present invention uses AHP to calculate the weight coefficients of each influencing factor index dimension, and weights each influencing factor from the index dimension. AHP is a systematic and hierarchical multi-criteria decision analysis method that combines qualitative analysis and quantitative analysis. This method can provide a comprehensive framework for dealing with the problem of selecting input features of the multivariate load short-term prediction model in this article, and weight each influencing factor that affects the change of various load data, so as to determine the index dimension weight coefficient ω of each influencing factor for the short-term prediction of the multivariate load of the regional energy system.
[0037] First is to construct the hierarchical structure of the index system. For the problem of calculating the weight coefficients of each influencing factor index dimension in the study of short-term prediction of multivariate loads in the regional energy system, the constructed hierarchical structure of the index system is as follows: The target layer is the short-term prediction of the multivariate load of the regional energy system; the criterion layers are the electric load prediction, the cooling load prediction, and the heating load prediction respectively; the scheme layer consists of historical data of various loads, meteorological information, and calendar information, where the historical data of various loads includes the historical data of electric load, cooling load, and heating load, and the meteorological information includes temperature, humidity, wind speed, and air pressure.
[0038] After that is to calculate the weight coefficients of each layer of indexes. In this article, the geometric mean method is used to calculate the weight coefficients, and the maximum eigenvalue λ of each pairwise comparison judgment matrix is calculated max and its corresponding eigenvector ω. Specifically, first multiply the index elements of the pairwise comparison judgment matrix M column by column through the following formula: where M j is the product result of the j-th column in the pairwise comparison judgment matrix; M ij is the element value of the i-th row and the j-th column in the pairwise comparison judgment matrix.
[0039] Then find its geometric mean through the following formula: where is the geometric mean of the j-th index in the pairwise comparison judgment matrix.
[0040] Finally, normalize the obtained geometric mean through the following formula to obtain the weight coefficient of this index: where is the normalized weight coefficient representing the j-th index; The maximum eigenvalue λ of matrix M is calculated through the following formula max : where
[0041] Take the index parameters of each layer in the above index system hierarchy as the criteria for comparing each index parameter in the next layer respectively. According to the AHP relative importance scale shown in the following table, construct a judgment matrix: Then conduct a consistency test on the judgment matrix. Specifically, the consistency test compares the calculated pairwise comparison judgment matrix test coefficient CR with 0.1. If CR is less than 0.1, it is considered that the consistency test passes; otherwise, it is considered that the consistency test fails and the constructed pairwise comparison judgment matrix needs to be corrected. Among them, CR is calculated through the following formula: And thus obtain the weight coefficients of each influencing factor index dimension. In a preferred embodiment, the calculation results are shown in the following table: Step S4: Generate the association degrees between the candidate influencing factors according to the correlation coefficients between the candidate influencing factors and the weight coefficients of the candidate influencing factors.
[0042] It should be explained here that the correlation analysis method based on the Copula theory focuses on mining the correlation information contained in the multivariate load and the time series data of each influencing factor itself, while AHP focuses on describing the weight coefficients of the electricity, cooling, and heating loads and each influencing factor index dimension. In order to more effectively quantify and analyze the coupling relationship between the multivariate loads and improve the scientificity and reliability of the correlation index, in this paper, the correlation coefficients derived from the Copula theory are weighted from the index dimension, and the calculation formula is as follows: γ = βω Where ω is the weight value of each influencing factor index dimension; β is the correlation weight coefficient.
[0043] Furthermore, determine the final weighted association degree index. In a specific embodiment, the specific results are shown in the following table: The main influencing factors obtained through the analysis of the load law changes, including historical load information, meteorological information, and calendar information, all have non-linear associations with the multi-source load forecasting. There are only differences in the degree of association. The proposed correlation measurement method based on the AHP-Copula theory in the present invention can precisely quantify this difference, thereby quantitatively analyzing the different degrees of influence of this difference on the short-term forecasting of the multi-source load in the regional energy system. According to the specific results in the above table, in an embodiment of the present invention, the temperature and humidity with a relatively large degree of association in the historical load data, calendar information, and meteorological information are determined as the input features of the model. The preprocessed various types of load data are plotted into load curves on different time scales to analyze the load change law, so as to obtain the main factors affecting the changes of various types of loads and use the correlation measurement method based on the AHP-Copula theory to measure their correlations; based on the weighted correlation measurement index, the input features of the multi-source load forecasting model are determined.
[0044] Among them, the correlation measurement method based on the AHP-Copula theory mainly includes three parts: calculating the correlation coefficient between the load and each influencing factor, calculating the weight coefficient of each influencing factor index dimension, and determining the weighted correlation measurement index. Compared with the traditional correlation analysis method, this method can accurately capture the complex coupling relationships among various types and the non-linear correlation relationships between various types of loads and numerous influencing factors. At the same time, according to the importance of the influencing factor indicators themselves, the weight coefficients of each influencing factor index dimension are given, fully integrating expert suggestions and the actual operation experience of relevant staff, and finally obtaining a correlation measurement result that comprehensively considers both subjective and objective dimensions.
[0045] Step S5: Take the candidate influencing factors with the degree of association greater than the preset threshold as the target influencing factors; Step S6: Input the target influencing factors into a preset multi-source load forecasting model, so that the multi-source load forecasting model outputs the electrical load, heat load, and cooling load of the target day in the area to be forecast according to the target influencing factors.
[0046] In a preferred embodiment, the preset multi-source load forecasting model is trained in the following manner: Obtain the target influencing factors of historical typical days and the corresponding historical multi-source load forecasting labels; the historical multi-source load forecasting labels are the electrical load, heat load, and cooling load of historical typical days; Construct a training set according to the target influencing factors of historical typical days and the corresponding historical multi-source load forecasting labels; Randomly divide the training set into several batches of training samples according to the preset batch size; Input the training samples of each batch into the multi - load prediction model in sequence to train the multi - load prediction model until the preset number of training times is reached; wherein, when the multi - load prediction model receives each batch of training samples, it outputs the predicted electrical load, predicted heat load, and predicted cooling load corresponding to the training samples; calculate the loss function value through the loss function according to the predicted electrical load, predicted heat load, predicted cooling load, and the corresponding historical multi - load prediction labels; update the multi - load prediction model according to the loss function value.
[0047] It should be noted here that according to the load characteristic analysis results, the multi - load has obvious time characteristics, and the load consumption shows a certain regularity with the change of time; due to different electricity demands and diverse influencing factors, the load consumption shows non - linearity and randomness; according to the correlation analysis results, there is obvious coupling between the cooling load, heat load, and electrical load. Considering these characteristics comprehensively, the present invention selects the LSTM neural network as the prediction model. The LSTM neural network is a variant of the RNN. By constructing a memory cell on the basis of the RNN structure, it transfers the memory information from the initial position of the sequence to the end of the sequence, and controls the modification of the memory information value at any time step through 3 interactive thresholds, so that the weight of the self - loop depends on the information of the previous and subsequent moments rather than a fixed value, thus effectively solving problems such as gradient disappearance and gradient explosion, improving the ability to process samples with long intervals or delays in time series, effectively mining the time correlation of samples, and at the same time having the ability of neural networks to process non - linear data, which is very suitable for predicting the multi - load of regional energy systems.
[0048] Compared with the RNN, the LSTM neural network redesigned the memory unit on the basis of maintaining the basic structure, and set 3 control gates, namely the input gate i t 、output gate o t 、forget gate f t , whose function is to selectively remember the correction parameters of the error function feedback with gradient descent, optimize the weight of the self - loop, and maintain the dynamic change of the weight. Its architecture is as Figure 2 shown. The input data of the LSTM at time t is x t , the output value is h t , and the memory state is c t .
[0049] The forward propagation of the information flow in the LSTM model is similar to that of BP and RNN, except that when passing through the hidden layer, the information flow is calculated through 3 gated units. The working principles of the three gated units are as follows.
[0050] (1) Forget gate unit The function of the forget gate: By selectively calculating and forgetting the useless information in the early stage, output f tThe value h passed from the hidden layer at the previous moment to the current moment t-1 and the input value x at the current moment t The combined vector is obtained through an activation function to achieve the purpose of forgetting useless information. The mathematical expression is shown in the formula: Among them, W f and b f are the weight matrix and bias vector of the forget gate respectively; σ is the activation function.
[0051] (2) Input gate unit The function of the input gate is to receive the current input and the state at the previous moment, and control the memory selection of information. The input gate consists of two layers of neurons. The input information first passes through a sigmoid function to obtain the data that needs to be input into the memory unit, and the information is selected in sequence according to the importance. At the same time, the input information passes through the tanh layer to create a new candidate state As the state quantity for the next cycle, the mathematical formula is as follows: Among them, W i and b i 、W c and b c are the weight matrix and bias vector of the two layers of neurons in the input layer respectively.
[0052] (3) Input gate unit The output gate unit consists of two parts, one is the output of the cell unit state, and the other is the output of the output gate. The cell unit state is one of the outputs of the memory structure unit. Multiply the result i t calculated by the input gate with the calculated new unit state value , multiply the result f t obtained by the forget gate with the unit state value C t-1 at the previous moment, and finally add the two to get the unit state value at the current moment, and the update is shown in the following formula: The function of the output gate: Select the state information as the output of the memory unit and pass it to the next layer and the input of the hidden layer at the next moment. The data o t after classifying the input state information through the sigmoid function is passed to the hidden layer data value h t at the next moment, which is obtained from the new unit state processed by the tanh function and the data o t . The mathematical expression is as follows: h t =o t·tanh(C t ) where W o and b o are the weight matrix and bias vector of the output gate.
[0053] According to the results of the forward pass calculation, the error δ of each neuron is calculated in reverse. Due to the time loop structure of the LSTM, the backpropagation includes two directions: one is the backpropagation along time, that is, starting from the current moment, the error of each moment is calculated; the other is the backpropagation along the neuron layer. The weights and gradients need to be calculated according to the corresponding errors in both directions.
[0054] For a training sample (X, y), the mathematical expression of the error function is: where Y W,b (X) is the predicted value; y is the true value.
[0055] The input weight matrix is partitioned according to the value h t-1 transmitted from the hidden layer at the previous moment to the current moment and the input value at the current moment, and defined as follows: where W fh , W oh , W ih , W ch are the weights of the forget gate, output gate, input gate, and cell state at the previous moment respectively; W fx , W ox , W ix , W cx are the weights of the forget gate, output gate, input gate, and cell state at the current moment respectively; Define the derivative of the loss function with respect to the output value at time t as the error term: To calculate the backpropagation of the error along time, the derivative of the error at time t - 1 needs to be derived, and the formula is as follows: To calculate the backpropagation of the error between layers, the derivative of the error at layer l - 1 needs to be derived, and the formula is as follows: From the above formula, the model that backpropagates to the previous layer can be calculated, and then the gradient is calculated. First, calculate the gradient of the weight matrix of the first part of the partition. Taking W fh as an example, the gradient at time t is: Add the gradients at each moment to calculate the final gradient: Similarly, the other weight gradients are obtained as follows: Among them, are the weight partial derivatives of the output gate, input gate, and cell state, that is, the gradient values; δ o,j , δ i,j , are the error values of the output gate, input gate, and cell state respectively; is the state output of the hidden layer at the previous moment.
[0056] Next, the gradients of the weight matrix of the second part of the partition are solved in turn: Among them, the left - hand expressions are the gradient values of the forget gate, output gate, input gate, and cell state at the current moment; W fx , W ox , W ix , are the weights of the forget gate, output gate, input gate, and cell state at the current moment respectively; By performing the above data transfer process multiple times, continuously correcting the weights and thresholds of the network, making the difference between the network output value and the expected value meet the requirements, the final LSTM model is determined.
[0057] From the above reasoning, it can be seen that the LSTM neural network retains the advantages of the RNN in processing time series, and through three gating units, it better retains the effective information of the previous moment, deeply mines the time - correlation information in the sample data, and can effectively prevent phenomena such as gradient disappearance and explosion.
[0058] As Figure 3 shown, in a preferred embodiment, to improve the prediction accuracy of the neural network, an attention mechanism is added to the LSTM neural network.
[0059] The attention mechanism is a model that simulates the human brain's ability to focus on specific things. When observing things, the human brain will focus its attention on the areas that need to be focused on, reducing or even ignoring the attention to other areas, in order to obtain more detailed information that needs to be focused on and suppress other useless information.
[0060] Applying the Attention mechanism in the neural network model can allocate more attention to the key parts of the input sequence that affect the output result, better learn the information in the input sequence, and improve the prediction accuracy.
[0061] The structure of the Attention mechanism is as Figure 3 shown. Among them, Softmax is the activation function, x t(t ∈ [1, n]) represents the input of the LSTM network, h t (t ∈ [1, n]) corresponds to the output of the hidden layer of the LSTM model, α t (t ∈ [1, n]) is the attention probability distribution value of the Attention mechanism for the output of the LSTM hidden layer. y is the output value of the LSTM with the Attention mechanism introduced. The calculation formulas for the attention weight matrix α and the feature vector matrix v in the Attention mechanism are as follows: e t = u s tanh(w s h t + b s ) Among them, e t refers to the unnormalized weight matrix; w s , b s and u s are the weight, bias, and time series matrix of the attention mechanism respectively.
[0062] Based on the above method embodiment, the present invention correspondingly provides an apparatus embodiment.
[0063] As Figure 4 shown, an embodiment of the present invention provides a multi-load prediction apparatus for a regional energy system based on big data, including: a data acquisition module, a correlation coefficient calculation module, a weight coefficient calculation module, a target influencing factor determination module, and a multi-load prediction module; The data acquisition module is used to acquire candidate influencing factors of the area to be predicted; among them, the candidate influencing factor data includes historical electrical load, historical heat load, historical cooling load, humidity, temperature, calendar information, air pressure, and wind speed; The correlation coefficient calculation module is used to calculate and generate the correlation coefficients between the candidate influencing factors according to the candidate influencing factors; The weight coefficient calculation module is used to calculate and generate the weight coefficients of the candidate influencing factors according to the candidate influencing factors; the target influencing factor determination module is used to generate the association degree between the candidate influencing factors according to the correlation coefficients between the candidate influencing factors and the weight coefficients of the candidate influencing factors; and use the candidate influencing factors with the association degree greater than the preset threshold as the target influencing factors; The multi-load prediction module is used to use the candidate influencing factors with the association degree greater than the preset threshold as the target influencing factors; input the target influencing factors into a preset multi-load prediction model, so that the multi-load prediction model outputs the electrical load, heat load, and cooling load of the target day in the area to be predicted according to the target influencing factors.
[0064] In an optional embodiment, for the multi - load prediction device of the regional energy system based on big data, the correlation coefficient calculation module includes: an edge cumulative distribution calculation unit, a correlation function parameter estimated value calculation unit, an optimal correlation function model determination unit, and a correlation coefficient determination unit; The edge cumulative distribution calculation unit is configured to calculate and generate the edge cumulative distribution of each candidate influencing factor based on the kernel density estimation algorithm according to the candidate influencing factors; The correlation function parameter estimated value calculation unit is configured to calculate the estimated values of the correlation function parameters of each candidate influencing factor based on the distribution maximum likelihood estimation algorithm according to the edge cumulative distribution of each candidate influencing factor; The optimal correlation function model is configured to determine the optimal correlation function model of each candidate influencing factor based on the Euclidean norm minimization algorithm according to the estimated values of the correlation function parameters of each candidate influencing factor; The correlation coefficient determination unit is configured to calculate and generate the correlation coefficients between each candidate influencing factor according to the optimal correlation function model of each candidate influencing factor.
[0065] In an optional embodiment, for the multi - load prediction device of the regional energy system based on big data, the weight coefficient calculation module includes: a comparison matrix determination unit and a weight coefficient determination unit; The comparison matrix determination unit is configured to construct a comparison matrix between candidate influencing factors based on a preset relative importance scale according to the candidate influencing factors; The weight coefficient determination unit is configured to calculate and generate the weight coefficients of each candidate influencing factor based on the geometric mean algorithm according to the comparison matrix.
[0066] In an optional embodiment, for the multi - load prediction device of the regional energy system based on big data, the edge cumulative distribution of each candidate influencing factor is calculated and generated by the following formula: where f(x) is the probability density function of the candidate influencing factor x; n is the total number of sample points of the candidate influencing factor x; h is the window frame constant coefficient; K(·) is the kernel function; x is the candidate influencing factor; x i is the value of the i - th sample point among the candidate influencing factors.
[0067] In an optional embodiment, for the multi - load prediction device of the regional energy system based on big data, the preset multi - load prediction model is trained in the following manner: Obtain the target influencing factors of historical typical days and the corresponding historical multi - load prediction labels; the historical multi - load prediction labels are the electrical load, heat load, and cooling load of historical typical days; Construct a training set according to the target influencing factors of historical typical days and the corresponding historical multi - load prediction labels; Randomly divide the training set into several batches of training samples according to a preset batch size; Input each batch of training samples into the multi - load prediction model in turn to train the multi - load prediction model until a preset number of training times is reached; among them, when the multi - load prediction model receives each batch of training samples, it outputs the predicted electrical load, predicted heat load, and predicted cooling load corresponding to the training samples; according to the predicted electrical load, predicted heat load, predicted cooling load, and the corresponding historical multi - load prediction labels, calculate the loss function value through the loss function; update the multi - load prediction model according to the loss function value.
[0068] It should be noted that the embodiments of the device described above correspond to the above - mentioned embodiments of the present invention and can implement any one of the above - mentioned methods for multi - load prediction of regional energy systems based on big data. In addition, the embodiments of the above - mentioned device are only illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided by the present invention, the connection relationship between modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement without creative work.
[0069] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0070] The above are the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.
Claims
1. A method for predicting multiple loads in a regional energy system based on big data, characterized in that Including: Obtain candidate influencing factors for the area to be predicted; wherein, the candidate influencing factors include historical electric load, historical heat load, historical cooling load, humidity, temperature, calendar information, air pressure, and wind speed; Calculate the correlation coefficients between the candidate influencing factors according to the candidate influencing factors; Calculate the weight coefficients of the candidate influencing factors according to the candidate influencing factors; Generate the association degrees between the candidate influencing factors according to the correlation coefficients between the candidate influencing factors and the weight coefficients of the candidate influencing factors; Use the candidate influencing factors with the association degrees greater than the preset threshold as the target influencing factors; Input the target influencing factors into a preset multivariate load prediction model, so that the multivariate load prediction model outputs the electric load, heat load, and cooling load of the target day in the area to be predicted according to the target influencing factors.
2. The multi-load prediction method for a regional energy system based on big data according to claim 1, characterized in that The calculating the correlation coefficients between the candidate influencing factors according to the candidate influencing factors includes: Calculate the marginal cumulative distributions of the candidate influencing factors based on the kernel density estimation algorithm according to the candidate influencing factors; calculate the estimated values of the correlation function parameters of the candidate influencing factors based on the distribution maximum likelihood estimation algorithm according to the marginal cumulative distributions of the candidate influencing factors; Determine the optimal correlation function model of the candidate influencing factors based on the Euclidean norm minimization algorithm according to the estimated values of the correlation function parameters of the candidate influencing factors; Calculate the correlation coefficients between the candidate influencing factors according to the optimal correlation function model of the candidate influencing factors.
3. The multi-load prediction method for a regional energy system based on big data according to claim 2, characterized in that The calculating the weight coefficients of the candidate influencing factors according to the candidate influencing factors includes: Construct a comparison matrix between the candidate influencing factors based on a preset relative importance scale according to the candidate influencing factors; calculate the weight coefficients of the candidate influencing factors based on the geometric mean algorithm according to the comparison matrix.
4. The multi-load prediction method for a regional energy system based on big data according to claim 3, wherein Calculate the marginal cumulative distributions of the candidate influencing factors through the following formula: Among them, f(x) is the probability density function of the candidate influencing factor x; n is the total number of sample points of the candidate influencing factor x; h is the window frame constant coefficient; K(·) is the kernel function; x is the candidate influencing factor; x i is the value of the i-th sample point among the candidate influencing factors.
5. The multi-load prediction method for a regional energy system based on big data according to claim 4, wherein Train a preset multivariate load prediction model in the following manner: Obtain the target influencing factors of historical typical days and the corresponding historical multivariate load prediction labels; the historical multivariate load prediction labels are the electric load, heat load, and cooling load of historical typical days; Construct a training set according to the target influencing factors of historical typical days and the corresponding historical multivariate load prediction labels; Randomly divide the training set into several batches of training samples according to a preset batch size; Input each batch of training samples into the multivariate load prediction model in turn to train the multivariate load prediction model until the preset number of training times is reached; wherein, when the multivariate load prediction model receives each batch of training samples, it outputs the predicted electric load, predicted heat load, and predicted cooling load corresponding to the training samples; calculate the loss function value through the loss function according to the predicted electric load, predicted heat load, predicted cooling load, and the corresponding historical multivariate load prediction labels; update the multivariate load prediction model according to the loss function value.
6. A multi-load prediction device for a regional energy system based on big data, characterized in that, Including: A data acquisition module, a correlation coefficient calculation module, a weight coefficient calculation module, a target influencing factor determination module, and a multi-load prediction module; the data acquisition module is used to acquire candidate influencing factors for the area to be predicted; wherein, the candidate influencing factors include historical electricity load, historical heat load, historical cold load, humidity, temperature, calendar information, air pressure, and wind speed; The correlation coefficient calculation module is used to calculate and generate the correlation coefficients between the candidate influencing factors according to the candidate influencing factors; The weight coefficient calculation module is used to calculate and generate the weight coefficients of the candidate influencing factors according to the candidate influencing factors; the target influencing factor determination module is used to generate the association degrees between the candidate influencing factors according to the correlation coefficients between the candidate influencing factors and the weight coefficients of the candidate influencing factors; and take the candidate influencing factors with the association degrees greater than the preset threshold as the target influencing factors; The multi-load prediction module is used to input the target influencing factors into a preset multi-load prediction model, so that the multi-load prediction model outputs the electricity load, heat load, and cold load of the target day in the area to be predicted according to the target influencing factors.
7. The multi-load prediction device for a regional energy system based on big data according to claim 6, characterized in that, The correlation coefficient calculation module includes: a marginal cumulative distribution calculation unit, a correlation function parameter estimate calculation unit, an optimal correlation function model determination unit, and a correlation coefficient determination unit; The marginal cumulative distribution calculation unit is used to calculate and generate the marginal cumulative distributions of the candidate influencing factors based on the kernel density estimation algorithm according to the candidate influencing factors; The correlation function parameter estimate calculation unit is used to calculate the correlation function parameter estimates of the candidate influencing factors based on the distribution maximum likelihood estimation algorithm according to the marginal cumulative distributions of the candidate influencing factors; The optimal correlation function model is used to determine the optimal correlation function models of the candidate influencing factors based on the Euclidean norm minimization algorithm according to the correlation function parameter estimates of the candidate influencing factors; The correlation coefficient determination unit is used to calculate and generate the correlation coefficients between the candidate influencing factors according to the optimal correlation function models of the candidate influencing factors.
8. The multi-load prediction device for a regional energy system based on big data according to claim 7, characterized in that The weight coefficient calculation module includes: a comparison matrix determination unit and a weight coefficient determination unit; The comparison matrix determination unit is used to construct a comparison matrix between the candidate influencing factors based on a preset relative importance scale according to the candidate influencing factors; The weight coefficient determination unit is used to calculate and generate the weight coefficients of the candidate influencing factors based on the geometric mean algorithm according to the comparison matrix.
9. The multi-load prediction device for a regional energy system based on big data according to claim 8, wherein The marginal cumulative distributions of the candidate influencing factors are calculated and generated through the following formula: Among them, f(x) is the probability density function of the candidate influencing factor x; n is the total number of sample points of the candidate influencing factor x; h is the window frame constant coefficient; K(·) is the kernel function; x is the candidate influencing factor; x i is the value of the i-th sample point among the candidate influencing factors.
10. The multi-load prediction device for a regional energy system based on big data according to claim 9, characterized in that The preset multi-load prediction model is trained in the following manner: Obtain the target influencing factors of historical typical days and the corresponding historical multi-load prediction labels; the historical multi-load prediction labels are the electricity load, heat load, and cold load of historical typical days; Construct a training set according to the target influencing factors of historical typical days and the corresponding historical multi-load prediction labels; Randomly divide the training set into several batches of training samples according to a preset batch size; Input the training samples of each batch into the multi-load prediction model in sequence to train the multi-load prediction model until a preset number of training times is reached; wherein, when the multi-load prediction model receives a batch of training samples each time, it outputs the predicted electric load, predicted heat load, and predicted cooling load corresponding to the training samples; calculate the loss function value through the loss function according to the predicted electric load, predicted heat load, predicted cooling load, and the corresponding historical multi-load prediction labels; update the multi-load prediction model according to the loss function value.