A short-term prediction method for multi-energy load

Through dynamic Copula correlation analysis and fuzzy clustering method combined with machine learning algorithms, the problem of accuracy and coupling relationship analysis in the short-term prediction model of multi-energy load is solved, and efficient multi-energy load prediction is achieved to adapt to the dynamic needs of the integrated energy system.

CN115271161BActive Publication Date: 2025-07-11SOUTH CHINA UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210681305.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-15
Publication Date
2025-07-11
Estimated Expiration
2042-06-15

AI Technical Summary

Technical Problem

The existing multi-energy load short-term prediction model has poor prediction accuracy and overall prediction accuracy at the characteristic points of the load curve, and it is difficult to effectively analyze the complex nonlinear coupling relationship between multi-energy loads.

Method used

The combination model of dynamic Copula correlation analysis and two-stage fuzzy optimization load feature recognition is used to quantify the nonlinear coupling relationship between multi-energy loads through dynamic Copula correlation analysis, and feature extraction and prediction are performed by combining fuzzy clustering analysis and machine learning algorithms. The probability transfer chain model is used for classification and judgment, and the predicted features are reconstructed to the original dimension through fitting reconstruction technology.

Benefits of technology

It improves the accuracy and calculation efficiency of short-term prediction of multi-energy loads, can adapt to the dynamic changes of multi-energy load prediction, and has excellent generalization performance and practical application value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115271161B_ABST
    Figure CN115271161B_ABST
Patent Text Reader

Abstract

The present invention discloses a short-term prediction method for multi-energy load, including: acquiring data; performing data preprocessing, and executing dynamic Copula correlation analysis and load fuzzy clustering analysis; based on the fuzzy optimization curve morphology analysis algorithm, respectively performing feature recognition and extraction on each class of load curve clusters; based on the statistical frequency distribution, calculating and determining the overall feature distribution set of each class of load curve clusters; based on the machine learning algorithm classification, establishing a combined load prediction model, and training the combined model according to the historical feature data of each class of load curve clusters; using the probability transition chain model to judge the classification to which the load in the prediction period belongs, predicting the load characteristics based on the load prediction model trained for this class, and analyzing the future electricity consumption characteristics. The two-stage fuzzy optimization load feature recognition combined model has an efficient improvement in both calculation efficiency and specific performance compared with the traditional model, and can adapt to the development dynamics of actual multi-energy load prediction work.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electric power, and particularly relates to a short-term prediction method for multi-energy loads. Background Art

[0002] Power load forecasting plays an indispensable role in the stable and economic operation of the power system. Electric load forecasting is to explore the influence of the historical data change law of the electric load on the future load according to the historical data of power load, economy, society, meteorology, etc., and seek the internal connection between the historical change of the electric load and its own and various related factors, so as to accurately predict the future electric load. The main task of the power system is to provide economic, stable and high-quality electric energy for various types of power users, and should be able to respond in a timely manner to the requirements of different power load demands and power load characteristics of power users. Therefore, in the whole process of power system planning and design, operation management and power market trading, it is necessary to accurately predict the change in the amount and quality of the electric load. After the power enters the market-oriented operation, the electric load forecasting is actually the forecasting of the power market demand. The characteristic of instantaneous balance between power supply and demand determines that the forecasting demand of the power industry is more urgent than that of other industries. To better carry out the electric load forecasting work to ensure the safe and economic operation of the power system and at the same time steadily promote the marketization process of the power industry has become the industry consensus in the power system.

[0003] With the continuous improvement of the grid intelligence level and the gradual deepening of the power marketing mechanism, the information interaction frequency between the user side and the grid side is increasing and the scope is wider. At the grid operation level, it is necessary to accurately grasp the major adjustment characteristics of the load side as much as possible, so as to ensure the reliable operation of the system through corresponding dispatching measures; while the market operation system is more concerned about the power consumption behavior characteristics of users to reasonably adjust the power generation plan and enhance the competitiveness in the power market environment. Based on the above analysis, the power system under the new development background puts forward more refined requirements for load forecasting, that is, not only to meet the needs of grid production dispatching, but also to more accurately depict the load characteristics.

[0004] Integrated energy systems integrate various energy forms such as electricity, heat, cold, and gas, involving energy conversion, distribution, and organic coordination. They are important physical carriers for realizing the multi-energy complementarity characteristics of the energy Internet. As time goes by, the fields and industries of the energy Internet and integrated energy systems that meet the requirements of the times and respond to the guidance of policy documents will be further broadened and strengthened. The multi-energy load forecasting service, which shoulders the core and fundamental responsibility of the operation of integrated energy systems, urgently needs extensive and in-depth research to promote the major plan of China's energy transformation. However, the difficulty in conducting short-term multi-energy load forecasting research on emerging integrated energy systems lies in that, on the one hand, the non-linear and complex nature of single-load forecasting has not changed. Therefore, it is necessary to flexibly analyze the intertwined connotations of the load and various external correlation information, as well as the internal historical development logic of the load itself and the influence mechanism of the load's future changes under the guidance of scientific and correct theories. On the other hand, in integrated energy systems, various forms of energy can be efficiently converted into each other with the help of energy coupling and conversion equipment on the premise of dynamically responding to the multi-energy demands of users. Therefore, based on traditional load forecasting, multi-energy load forecasting also requires reasonable analysis and research on multi-energy coupling characteristics, so as to fully understand the changing laws of the development nature of multi-energy loads.

[0005] With the popularization of the concept of artificial intelligence and the development of technology, machine learning and deep learning have been widely supported and applied in load forecasting research. Among them, machine learning methods such as neural networks and support vector machine regression have developed rapidly due to their powerful data mining capabilities and superiority in solving complex non-linear problems.

[0006] (1) Feature processing

[0007] Feature processing work such as feature selection, dimensionality reduction, and extraction of artificial intelligence-based prediction models is an important factor affecting prediction accuracy. In typical engineering applications of machine learning, feature selection methods can be roughly classified into filter methods, wrapper methods, and embedded methods; feature dimensionality reduction methods include principal component analysis and linear discriminant analysis, etc.; feature extraction methods include principal component analysis, intelligent learning methods, etc.

[0008] (2) Correlation analysis

[0009] In integrated energy systems, the complex coupling relationships among multi-energy loads directly affect the accuracy performance of short-term multi-energy load forecasting. Therefore, considering load feature analysis is the key entry point for improving the accuracy of short-term multi-energy load forecasting. Existing load correlation analysis techniques mainly include linear correlation analysis, quantitative information correlation analysis, static correlation analysis, etc. Summary of the invention

[0010] To solve at least one of the technical problems existing in the above-mentioned background technology, the present invention provides a method for short-term multi-energy load forecasting.

[0011] To achieve the above object, the technical solution of the present invention is as follows:

[0012] A multi-energy load short-term prediction method, comprising:

[0013] Obtaining data, where the data includes multi-energy load historical curves and historical information of load influencing factors;

[0014] Performing data preprocessing, and performing dynamic Copula correlation analysis and load fuzzy clustering analysis;

[0015] Based on the fuzzy optimization curve shape analysis algorithm, respectively performing feature recognition and extraction on each type of load curve cluster to achieve primary load feature extraction;

[0016] Based on the statistical frequency distribution, calculating and determining the overall feature distribution set of each type of load curve cluster to achieve secondary load feature extraction;

[0017] Based on the machine learning algorithm classification, establishing a combined load prediction model, and training the combined model according to the historical feature data of each type of load curve cluster;

[0018] Using the probability transition chain model to judge the classification to which the load in the prediction period belongs, predicting the load characteristics based on the load prediction model trained for this type, and analyzing the future electricity consumption characteristics;

[0019] Based on the fitting reconstruction technology, reconstructing the prediction features to the original dimension of the load curve to complete the load prediction.

[0020] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0021] Aiming at the problems of poor prediction accuracy at the characteristic points of the multi-energy load curve and overall prediction accuracy existing in the short-term prediction models of multi-energy loads in integrated energy systems, the present invention proposes a short-term prediction method for multi-energy loads based on a combined model of dynamic Copula correlation analysis and two-stage fuzzy optimization load feature recognition. First, through dynamic Copula correlation analysis, this method can dynamically quantify and analyze the complex non-linear coupling relationships between multi-energy loads, thus providing strong reference support for the establishment of the input of the prediction model. Secondly, this method uses fuzzy clustering analysis to cluster historical load curves, and then analyzes the characteristics of each type of load according to the clustering results, making the application of the prediction model targeted. Based on the fact that similar load curves of the same type have similar load characteristics, a fuzzy optimization curve shape analysis method based on an adaptive adjustment threshold is used for primary feature extraction, and a statistical frequency distribution model is applied for secondary feature extraction, so as to perform two-stage load feature recognition and extraction on load curves of the same type. This not only solves the problems of feature selection and dimensionality reduction in load prediction work, but also makes the features obtained from load extraction clearer and more interpretable. At the same time, the prediction model pointing to load characteristics can reasonably solve the problem of low prediction accuracy at load characteristics in the prediction results. In addition, this method applies a probability transfer chain model to solve the classification and discrimination problem of the day to be predicted, thus having more practical application value. Finally, the two-stage fuzzy optimization load feature recognition method can be combined with a variety of machine learning methods to form a combined prediction model, which has excellent generalization performance. At the same time, the predicted future load characteristics can also be combined with a variety of reverse reconstruction methods to reconstruct the load from the feature dimension to the original load dimension, having excellent application potential. The two-stage fuzzy optimization load feature recognition combined model has a highly efficient improvement in both computational efficiency and specific performance compared with traditional models, and can adapt to the development dynamics of actual multi-energy load prediction work. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 is a flowchart of the short-term prediction method for multi-energy loads provided by an embodiment of the present invention;

[0023] Figure 2 is a structure diagram of a NARX neural network;

[0024] Figure 3 is a key technology implementation framework diagram of the short-term prediction method for multi-energy loads of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] Embodiment:

[0026] The technical solutions of the present invention will be further described below with reference to the drawings and embodiments.

[0027] Refer to Figure 1As shown in the figure, the short-term multi-energy load prediction method provided in this embodiment includes:

[0028] 101. Obtain data, where the data includes the historical curves of multi-energy loads and the historical information of load influencing factors.

[0029] 102. Perform data preprocessing, and execute dynamic Copula correlation analysis and load fuzzy clustering analysis.

[0030] For short-term multi-energy load prediction work, accurately grasping the internal load characteristics represented by the respective historical development laws of multi-energy loads and the coupling conversion relationship between multi-energy loads, as well as the external load characteristics of multi-energy loads and related influencing factors such as meteorology, is the key to improving the final prediction accuracy. Compared with the respective historical development laws of multi-energy loads and the correlation characteristics between multi-energy loads and related influencing factors such as meteorology, there is less research on the complex and flexible coupling conversion related characteristics between multi-energy loads. Therefore, a more accurate characterization method is needed to assist in the analysis, so as to effectively optimize the overall performance of the short-term multi-energy load prediction model. The dynamic Copula correlation analysis method is introduced to model and analyze the non-linear coupling conversion correlation characteristics between multi-energy load time series, so as to more precisely quantify and study the characteristic analysis method of multi-energy loads, and provide a reliable decision-making reference for optimizing the input characteristics of the short-term multi-energy load prediction model. In this way, in this step, through dynamic Copula correlation analysis, the complex non-linear coupling relationship between multi-energy loads can be dynamically quantified and analyzed, so as to provide strong information reference support for the establishment of the prediction model input.

[0031] 103. Based on the fuzzy optimization curve shape analysis algorithm, perform feature recognition and extraction on each type of load curve cluster (primary load feature extraction);

[0032] 104. Based on the statistical frequency distribution, calculate and discriminate to obtain the overall feature distribution set of each type of load curve cluster (secondary load feature extraction);

[0033] In this way, through steps 103 and 104, on the basis that the same type of load curves have similar load characteristics, the primary feature extraction is carried out based on the fuzzy optimization curve shape analysis method with an adaptive adjustment threshold, and the secondary feature extraction is applied using the statistical frequency distribution model, so as to perform two-stage load feature recognition and extraction on the same type of load curves, which not only solves the complexity problems of feature selection and dimensionality reduction in load prediction work, but also makes the extracted load features clearer and more interpretable. At the same time, the prediction model pointing to the load features can reasonably solve the problem of low prediction accuracy at the load features in the load prediction results.

[0034] 105. A combined load forecasting model is established based on the classification of machine learning algorithms, and the combined model is trained according to the historical feature data of each type of load curve cluster.

[0035] After the load feature recognition and extraction in the form of two-stage (primary and secondary) fuzzy optimization, the output unified load feature set can be combined with a variety of machine learning models to form a combined prediction model, such as a neural network model, a support vector machine regression model, etc. Among them, the NARX neural network model has stronger competitiveness compared with other typical machine learning models because of its reasonable structural performance, excellent ability to capture the nonlinearity of time series, and at the same time its parallel distributed training mode improves the fault tolerance and stability of the model.

[0036] Compared with the traditional feedforward neural network composed of an input layer, a hidden layer, and an output layer, the NARX neural network adds a feedback connection from the output layer to the input layer, thus forming a recurrent neural network. The structure of the NARX neural network is as Figure 2 shown. Among them, [d -1 ,...,d -n represents the time delay operator; ω H , ω V are the connection weight coefficients from the input layer to the hidden layer and from the hidden layer to the output layer respectively; Ψ and Φ represent the activation functions of the neurons in the hidden layer and the output layer respectively. Therefore, the input-output relationship of the NARX neural network model can be described by the following formula:

[0037]

[0038] In formula (1), I(·) is the input of the NARX model, O(·) is the output of the NARX model, F(·) is a typical non-linear activation function, and n is the time delay order of the external input and its own output. Combining formula (1), it can be seen that the output O(t + 1) of the NARX neural network at the next moment is always determined by the previous external input of the network [I(t), I(t - 1),..., I(t - n)] and the network's own output [O(t), O(t - 1),..., O(t - n)]. Therefore, the NARX model can fully consider the historical information contained in the time series, so as to be used to finely depict the state of the time series at the prediction moment.

[0039] In this way, the two-stage fuzzy optimization load feature recognition method can be combined with a variety of machine learning methods to form a combined prediction model, which has excellent generalization performance; at the same time, the predicted future load features can also be combined with a variety of reverse reconstruction methods, so as to reconstruct the load from the feature dimension to the original load dimension, with broad application prospects.

[0040] 106. Use the probability transition chain model to determine the category to which the load in the prediction period belongs, predict the load characteristics based on the load prediction model trained for this category, and analyze the future electricity consumption characteristics.

[0041] Applying the probability transition chain model to solve the classification discrimination problem of the day to be predicted when some information of the relevant factors for load prediction is unknown has practical application value.

[0042] 107. Based on the fitting and reconstruction technology, reconstruct the prediction features to the original dimension of the load curve to complete the load prediction.

[0043] To complete the electricity load prediction work, it is necessary to reconstruct the prediction feature set to the original dimension of the daily load curve. In the case of having the characteristic points of the future daily load curve, the intuitive reverse reconstruction method of the electricity load curve is mainly the piecewise linear interpolation method, as shown in Equation (2):

[0044]

[0045] In addition, there is also the B-spline interpolation method with smoother connection and more accurate results for the reverse reconstruction method of the electricity load curve, as shown in Equation (3):

[0046]

[0047] Such as Figure 3 shown, is the key technology implementation framework diagram of the multi-energy load short-term prediction method of the present invention. In step 101, the historical curves of the multi-energy load and the historical information data of the load influencing factors include electric load sequence data, heat load sequence data, cold load sequence data, meteorological data, and day type data.

[0048] In step 102, the dynamic Copula correlation analysis includes:

[0049] Modeling and analyzing the non-linear coupling conversion correlation characteristics among the electric load sequence data, heat load sequence data, and cold load sequence data; the modeling and analysis include:

[0050] The joint distribution function of n-dimensional variables can be constructed by the combination of the marginal distribution functions of n variables and the Copula function representing the complex correlation characteristics between variables. Assume that there is a certain group of n-dimensional variables x = [x1, x2,..., x n , where the marginal distribution functions of each dimensional variable are F1(x1), F2(x2),..., F n (x n ), then the joint distribution function and joint density function of this n-dimensional variable can be analyzed and described by Equations (4) and (5):

[0051] F(x1, x2,..., xn ) = C[F1(x1), F2(x2), …, F n (x n )] (4)

[0052]

[0053] Wherein, F(·) is the joint distribution function of n-dimensional variables; C(·) is the Copula distribution function characterizing the complex correlation characteristics between n-dimensional variables; f(·) represents the joint density function of n-dimensional variables; c(·) represents the Copula density function; f(·) is the density function of each dimensional variable.

[0054] It can be seen from formulas (4) and (5) that by differentiating the Copula distribution function C(·), the Copula density function c(·) can be obtained, as shown in formula (6).

[0055]

[0056] By obtaining the Copula distribution function and the Copula density function, the complex correlation between multivariate random variables can be described in detail. Based on the dynamic Copula distribution function for correlation feature analysis of random variables, the correlation characteristics of random variables can be reflected from different perspectives. Assuming that the Copula function of the second-order random variables (X, Y) with marginal distribution functions u = F(x) and v = G(y) is C(u, v), then the dynamic Copula distribution function and its density function are as follows:

[0057]

[0058] Wherein, are the time-varying upper tail and time-varying lower tail correlation coefficients of the dynamic Copula function. The correlation coefficient derived from the dynamic Copula function has time-varying characteristics and can describe in detail the non-linear dynamic change characteristics of the variable sequence, so as to dynamically quantify the complex non-linear coupling correlation between multi-energy loads.

[0059] In the above step 102, the load fuzzy clustering analysis includes:

[0060] For each type of data among the electric load sequence data, heat load sequence data, cold load sequence data, meteorological data, and day type data, the following steps are performed:

[0061] Step 1: Initialization: Set the total data sample set as X = {x j |j = 1, 2, … N}, and the number of clusters is M (2 ≤ M ≤ N); that is, X is to be divided into M classes, denoted as X 1 , X2 , …, X M ; At this time, there are Meanwhile, set M initial clustering centers, denoted as C = (c1(k), c2(k), …, c M (k)); Step 2: Calculate the Euclidean distance d j from the sample x i to the clustering center c ji ;

[0062]

[0063] where t is the total number of clustering indicators of the data samples;

[0064] Step 3: Calculate the membership degree u j of the sample x ji to the i-th class according to formula (9);

[0065]

[0066] Cluster X according to the minimum distance principle; Assume that the relational expression (10) is satisfied

[0067]

[0068] Then there is where i0 = 1, 2, …, M;

[0069] Step 4: Update the clustering center according to formula (11);

[0070]

[0071] Step 5: Assume that at this time, there exists i = {1, 2, …, M}, satisfying c i (k + 1) ≠ c i (k), then go back to Step 2 to execute the process again; Otherwise, the fuzzy clustering ends.

[0072] In the above Step 103, the process of the load initial feature extraction method can be briefly described as:

[0073] Step 1: Virtually connect the head and tail points of the target curve with a straight line, and calculate the perpendicular distance from each remaining target point on the current target curve to this straight line;

[0074] Step 2: Set the threshold of the curve shape analysis algorithm, and select the maximum vertical distance calculated in Step 1 to compare with the threshold. If it is greater than the threshold, retain the data point corresponding to the maximum vertical distance from the straight line; Otherwise, discard all data points between the two endpoints of the straight line;

[0075] Step 3: Based on the retained data points, divide the target curve into two parts for processing. Each part is regarded as a new target curve, and repeat the operations in Step 1 and Step 2. Iterate repeatedly based on the idea of the dichotomy method, that is, still select the maximum vertical distance to compare with the threshold, and judge and select in turn until there are no points to discard. Finally, obtain the curve feature points that meet the predetermined accuracy threshold, and discard other points, then the feature extraction of the target curve can be completed.

[0076] Define the domain \(E\in[0,1]\) as the threshold value range for the initial feature extraction. To simplify the operation, this paper sets the threshold to take values within the interval \([0,1]\), and the value interval is \(0.1\). \(Sat(\varepsilon)\) represents the membership degree of the threshold value of the curve shape analysis algorithm for a cluster of similar curves.

[0077]

[0078] In Equation (12), the threshold membership degree \(Sat(\varepsilon)\) is composed of the sum of two parts, and can be regarded as the overall curve feature extraction satisfaction degree of the specified similar curve cluster under a certain threshold. Among them, the first part \(D(\varepsilon)\) is the average matching degree between the curve features identified and extracted in the specified similar curve cluster and the original curve, so as to judge whether the extracted curve features can fully reflect the shape features of the original curve. The second part \(Z(\varepsilon)\) is the percentage value of the number of curve feature points to the number of original curve points, that is, the average percentage ratio of the curve feature extraction to compress the original curve. \(a\) and \(b\) are the corresponding proportionality coefficients.

[0079] 1) Average similarity matching degree \(D(\varepsilon)\).

[0080] Since the time dimension of the curve features is reduced compared with the original curve and is no longer a one-to-one mapping relationship, the Dynamic Time Warping (DTW) distance is introduced to calculate the matching degree between the curve features and the original curve. By calculating the DTW distance, the similarity of two time series with length differences can be reasonably measured. Let the original sequence and the feature sequence of the curve be \(X\) and \(Y\), then their sequence lengths are \(L\) X 、\(L\) Y . Define the warping path \(w = [w_1, w_2, \ldots, w\) K , where \(w\) i = (p\) i , q\) i ) \(\in [1:L\) X \(\times [1:L\) Y , \(1\leq i\leq K\). The warping path needs to satisfy boundary conditions, continuity and monotonicity, as shown in Equation (13).

[0081]

[0082] The cumulative distance F of the alignment path between the original sequence X and the feature sequence Y w (X, Y) can be calculated by Equation (14).

[0083]

[0084] The optimal alignment path w * The cumulative distance of the alignment path reaches the minimum. At this time, the cumulative distance of the alignment path is the DTW distance.

[0085]

[0086] Based on the idea of dynamic programming, the cumulative distance of the optimal alignment path can be recursively calculated by Equation (16).

[0087] D dtw (x i , y j ) = d(x i , y j ) + min(D dtw (x i-1 , y j ), D dtw (x i , y j-1 ), D dtw (x i-1 , y j-1 )) (16)

[0088] Iteratively calculate until D dtw (x LX , y LY ) is the DTW distance between X and Y we want.

[0089] In the extreme case, the curve features only include the start and end points of the original curve. Denote the straight line connecting the start and end points of the original curve as Y 0 , then the average matching degree D(ε) between the curve features identified and extracted from the similar curve cluster and the original curve is calculated by formula (17):

[0090]

[0091] In Equation (13), n is the number of curves included in the current similar curve cluster.

[0092] 2) The average compression ratio Z(ε).

[0093] Calculate the average compression ratio Z(ε) between the curve features identified and extracted from the similar curve cluster and the original curve through formula (18):

[0094]

[0095] In formula (14), the Num(·) function represents obtaining the number of data in the sequence.

[0096] 3) The proportionality coefficients a and b.

[0097] The setting of the proportionality coefficients a and b is related to the degree of emphasis on the average matching degree D(ε) and the average compression ratio Z(ε).

[0098] In the above step 103, the secondary feature extraction of the load includes:

[0099] After the initial extraction of the load characteristics of each load curve in a certain type of load curve cluster is completed by the curve shape analysis algorithm based on the fuzzy optimization threshold, the overall characteristics of this type of load curve cluster are extracted secondly by applying the idea of statistical frequency distribution. Denote the set of all non-repeating m load characteristic numbers generated during the load characteristic recognition and extraction process of this type of load curve cluster as I = [I1, I2, …, I i , …, I m , and the corresponding occurrence frequencies are G = [g1, g2, …, g i , …, g m . Then, the statistical frequency of each load characteristic of this type of load curve cluster is calculated by formula (19). Then, it is judged whether this load characteristic is suitable as one of the overall characteristics of this type of load curve cluster by formula (20). If the formula (20) is satisfied, the corresponding I i is added to the updated set I′ of the characteristic numbers of this type of load curve cluster, that is, I i ∈I′.

[0100]

[0101]

[0102] Finally, the set I′ obtained through the initial feature extraction by the curve shape analysis algorithm based on the fuzzy optimization threshold and the secondary feature extraction calculation of the statistical frequency distribution is used as the overall characteristic flag of this type of load curve cluster.

[0103] In the above step 106, the classification and discrimination of the day to be predicted based on the probability transition chain model includes:

[0104] Applying the probability transition chain model, first, it is necessary to determine the size of the state space of the system under study. Assume that the category coding sequence of the historical load contains a total of r states, denoted as [s1, s2, …, s i , …, s j , …, s r . When the result of the load FCM clustering is M categories, the value range of any state s i is [1, M]. According to the definition of the probability transition chain, the state s iThe probability of transferring to state s after n steps is calculated by formula (21). j as follows

[0105] P ij (n) = T ij (n) / T i (21)

[0106] where T ij (n) represents the number of times that state s transfers to state s i after n steps in the historical state sequence; T j represents the total number of times that state s appears in the historical state sequence i i i .

[0107] Repeatedly according to formula (21), the n-step state transition probability matrix of the probability transition chain can be calculated, as shown in formula (22) and formula (23):

[0108]

[0109]

[0110] When predicting the state at the (r + 1)-th step from the probability transition matrix, first find the maximum transition probability of state s in the r-th row of the probability transition matrix, assumed to be P r (n), where u = 1, 2,..., r. Then according to the definition of the probability transition chain, it can be determined that state s ru is the state with the maximum probability at the (r + 1)-th step. Through the above method, it is possible to reasonably determine which known classification the day to be predicted belongs to, so as to carry out the corresponding prediction work. u In summary, compared with the prior art, the present invention has the following advantages:

[0111] (1) Using the fuzzy clustering analysis method to cluster the historical load curves, and then analyzing the characteristics of each type of load according to the clustering results respectively, so that the application of the prediction model has a clear pertinence.

[0112] (2) On the basis that the load curves of the same type have similar load characteristics, based on the fuzzy optimization curve shape analysis method with an adaptive adjustment threshold for primary feature extraction, and applying the statistical frequency distribution model for secondary feature extraction, so as to perform two-stage load feature recognition and extraction on the load curves of the same type. This not only solves the complexity problems of feature selection and dimensionality reduction in the load prediction work, but also makes the features obtained from the load extraction more clear and interpretable. At the same time, the prediction model pointing to the load characteristics can reasonably solve the problem of low prediction accuracy at the load characteristics in the load prediction results.

[0113] (3) By using the above method, the load prediction work has a clear pertinence, and the load characteristics obtained from the load extraction are more clear and interpretable, and the prediction model pointing to the load characteristics can reasonably solve the problem of low prediction accuracy at the load characteristics in the load prediction results.

[0114] (3) Applying the probability transition chain model to solve the classification and discrimination problem of the day to be predicted when the information of relevant factors for partial load prediction is unknown has practical application value.

[0115] (4) The two-stage fuzzy optimization load feature recognition method can be combined with various machine learning methods to form a combined prediction model, which has excellent generalization performance; at the same time, the predicted future load features can also be combined with various reverse reconstruction methods, so as to reconstruct the load from the feature dimension to the original load dimension, having broad application prospects.

[0116] The above embodiments are only for illustrating the technical concept and features of the present invention, and the purpose is to enable those of ordinary skill in the art to understand the content of the present invention and implement it accordingly, and it should not be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the essence of the content of the present invention should be covered within the protection scope of the present invention.

Claims

1. A short-term prediction method for multi-energy load, characterized in that, Including: Obtaining data, where the data includes the historical curve of multi-energy load and the historical information of load influencing factors; Performing data preprocessing, and executing dynamic Copula correlation analysis and load fuzzy clustering analysis; Based on the fuzzy optimization curve morphology analysis algorithm, respectively performing feature recognition and extraction on each class of load curve clusters to achieve primary load feature extraction; Based on the statistical frequency distribution, calculating and discriminating to obtain the overall feature distribution set of each class of load curve clusters to achieve secondary load feature extraction; Based on the machine learning algorithm classification, establishing a combined load prediction model, and training the combined model according to the historical feature data of each class of load curve clusters; Using the probability transfer chain model to judge the classification to which the load in the prediction period belongs, predicting the load characteristics based on the load prediction model trained for this class, and analyzing the future power consumption characteristics; Based on the fitting reconstruction technology, reconstructing the prediction features to the original dimension of the load curve to complete the load prediction; The establishing of the combined load prediction model based on the machine learning algorithm classification and training the combined model according to the historical feature data of each class of load curve clusters includes: After the fuzzy optimization load feature recognition and extraction in a two-stage form, the output unified load feature set can be combined with multiple machine learning models to form a combined prediction model; The combined prediction model adopts the NARX neural network model, and the input-output relationship of the NARX neural network model is described by the following formula: In formula (1), I(·) is the input of the NARX model, O(·) is the output of the NARX model, F(·) is a typical non-linear activation function, and n is the delay order of the external input and its own output; The calculating and discriminating based on the statistical frequency distribution to obtain the overall feature distribution set of each class of load curve clusters includes: After the curve shape analysis algorithm based on the fuzzy optimization threshold completes the initial extraction of the load characteristics of each load curve in a certain type of load curve cluster, the overall characteristics of this type of load curve cluster are extracted secondly by applying the idea of statistical frequency distribution; Denote the set of m non-repeating load characteristic numbers generated during the load characteristic identification and extraction process of this type of load curve cluster as I = [I1, I2, …, I i , …, I m , and the corresponding occurrence frequencies are G = [g1, g2, …, g i , …, g m , then the statistical frequencies of the load characteristics of this type of load curve cluster are calculated by formula (19); Then, it is judged whether the load characteristic is suitable as one of the overall characteristics of this type of load curve cluster by formula (20); if the formula (20) is satisfied, then the corresponding I i is added to the set I' of the characteristic numbers of the updated load curve cluster of this type, that is, I i ∈I'; Finally, the set I′ obtained by the curve morphology analysis algorithm primary feature extraction with a fuzzy optimization threshold and the statistical frequency distribution secondary feature extraction calculation is used as the overall feature mark of this class of load curve clusters.

2. The multi-energy load short-term prediction method according to claim 1, characterized in that The reconstructing the prediction features to the original dimension of the load curve based on the fitting reconstruction technology to complete the load prediction includes: To complete the power load prediction work, it is necessary to reconstruct the prediction feature set to the original dimension of the daily load curve; in the case of the characteristic points of the future daily load curve already existing, the intuitive reverse reconstruction method of the power load curve is mainly the piecewise linear interpolation method, as shown in formula (2): Or, as shown in formula (3):

3. The multi-energy load short-term prediction method according to claim 1, characterized in that The historical curve of multi-energy load and the historical information data of load influencing factors include electric load sequence data, heat load sequence data, cooling load sequence data, meteorological data, and day type data.

4. The multi-energy load short-term prediction method according to claim 3, wherein, The dynamic Copula correlation analysis includes: Modeling and analyzing the non-linear coupling conversion correlation characteristics among the electric load sequence data, heat load sequence data, and cooling load sequence data; the modeling and analysis include: The joint distribution function of n-dimensional variables is constructed by the combination of the marginal distribution functions of n variables and the Copula function that characterizes the complex correlation characteristics between variables; assume that there is a certain set of n-dimensional variables x = [x1, x2, …, x n , where the marginal distribution functions of each dimension of variables are F1(x1), F2(x2), …, F n (x n ), then the joint distribution function and joint density function of the n-dimensional variables are analyzed and described by equations (4) and (5): F(x1,x2,…,x n ) = C[F1(x1), F2(x2),…, F n (x n )] (4) In the formula, F(·) is the joint distribution function of n-dimensional variables; C(·) is the Copula distribution function representing the complex correlation characteristics among n-dimensional variables; f(·) represents the joint density function of n-dimensional variables; c(·) represents the Copula density function; f(·) is the density function of each dimension variable; As can be seen from Equations (4) and (5), by differentiating the Copula distribution function C(·), the Copula density function c(·) is obtained, as shown in Equation (6). By obtaining the Copula distribution function and the Copula density function, the complex correlations between multivariate random variables can be described; based on the dynamic Copula distribution function for correlation feature analysis of random variables, the correlation characteristics of random variables can be reflected from different perspectives; assuming that the Copula function of the second-order random variables (X,Y) with marginal distribution functions u = F(x) and v = G(y) is C(u,v), the forms of the dynamic Copula distribution function and its density function are as follows: In the formula, is the time-varying upper tail and time-varying lower tail correlation coefficients of the dynamic Copula function; the correlation coefficients derived from the dynamic Copula function have time-varying characteristics and can describe the non-linear dynamic change characteristics of the variable sequence, so as to dynamically quantify the complex non-linear coupling correlation between multi-energy loads.

5. The multi-energy load short-term prediction method according to claim 3 or 4, characterized in that The load fuzzy clustering analysis includes: Perform the following steps on each type of data among the electric load sequence data, heat load sequence data, cold load sequence data, meteorological data, and day type data: Step 1: Initialization: Set the total data sample set as X = {x j | j = 1, 2, … N}, the number of clusters is M (2 ≤ M ≤ N); that is, X is to be divided into M classes, denoted as X 1 , X 2 , …, X M ; at this time, there is Meanwhile, set M initial cluster centers, denoted as C = (c1(k), c2(k), …, c M (k)); Step 2: Calculate the sample x by formula (8) j to the cluster center c i of the Euclidean distance d ji ; where t is the total number of clustering indicators for data samples; Step 3: Calculate the sample x by formula (9) j The membership degree u of the i-th class ji ; Cluster X according to the minimum distance principle; assume that the relational expression (10) is satisfied Then there is where \(i_0 = 1, 2, \ldots, M\); Step 4: Update the clustering center according to formula (11); Step 5: Assume that at this time there exists \(i = \{1, 2, \ldots, M\}\) such that \(c i (k + 1) \neq c i (k)\), then go back to Step 2 to execute the process again; otherwise, the load fuzzy clustering ends.

6. The multi-energy load short-term prediction method according to claim 5, characterized in that The feature recognition and extraction of each type of load curve cluster based on the fuzzy optimization curve shape analysis algorithm includes: Step 1: Virtually connect the head and tail points of the target curve with a straight line, and calculate the perpendicular distance from each remaining target point on the current target curve to this straight line; Step 2: Set the threshold of the curve shape analysis algorithm, and compare the maximum vertical distance calculated in Step 1 with the threshold. If it is greater than the threshold, retain the data point corresponding to the maximum vertical distance from the straight line; otherwise, discard all data points between the two endpoints of the straight line; Step 3: Based on the retained data points, divide the target curve into two parts for processing. Each part is regarded as a new target curve, and repeat the operations in Step 1 and Step 2. Iterate repeatedly based on the idea of the dichotomy method, that is, still select the maximum vertical distance to compare with the threshold, and judge whether to retain or discard in turn until there are no points to discard; finally, obtain the curve feature points that meet the predetermined accuracy threshold, and discard other points to complete the feature extraction of the target curve.

7. The multi-energy load short-term prediction method according to claim 6, wherein The feature recognition and extraction of each type of load curve cluster based on the fuzzy optimization curve shape analysis algorithm also includes: Define the domain E ∈ [0,1] as the threshold value range for the initial feature extraction; set the threshold to take values in the interval [0,1], and the value interval is 0.1; Sat(ε) represents the membership degree of the threshold value of the curve shape analysis algorithm for a cluster of similar curves; In Equation (12), the threshold membership degree Sat(ε) is composed of the sum of two parts, regarded as the overall curve feature extraction satisfaction degree of a specified similar curve cluster under a certain threshold; among them, the first part D(ε) is the average matching degree between the curve features identified and extracted in the specified similar curve cluster and the original curve, to judge whether the extracted curve features can fully reflect the shape features of the original curve; the second part Z(ε) is the percentage value of the number of curve feature points to the number of original curve points, that is, the average percentage ratio of the curve feature extraction to compress the original curve; a and b are the corresponding proportionality coefficients; 1) Average similarity matching degree D(ε) Introduce the dynamic time warping distance to calculate the matching degree between the curve feature and the original curve; assume the original sequence and the feature sequence of the curve are X and Y, and their sequence lengths are L X , L Y; Define the warping path w = [w1, w2, …, w K , where w i = (p i , q i ) ∈ [1: L X × [1: L Y , 1 ≤ i ≤ K; the warping path needs to satisfy boundary conditions, continuity, and monotonicity, as shown in Equation (13): Then the cumulative distance F of the alignment path between the original sequence X and the feature sequence Y w (X, Y) is calculated by Equation (14); Optimal alignment path w * The cumulative distance of the alignment path under this is minimized, and the cumulative distance of the alignment path at this time is the DTW distance; Based on the idea of dynamic programming, the cumulative distance of the optimal alignment path is recursively calculated by Equation (16); D dtw (x i ,y j ) = d(x i ,y j ) + min(D dtw (x i-1 ,y j ), D dtw (x i ,y j-1 ), D dtw (x i-1 ,y j-1 )) (16) Iterative calculation to D dtw (x LX ,y LY ) is the DTW distance between the required X and Y; The curve features in the extreme case only include the starting and ending points of the original curve. Denote the straight line connecting the starting and ending points of the original curve as Y 0 , then the average matching degree D(ε) between the curve features identified and extracted from the similar curve cluster and the original curve is calculated by formula (17): In Equation (13), n is the number of curves included in the current similar curve cluster; 2) Average compression ratio Z(ε) The average compression ratio Z(ε) of the curve features identified and extracted in the similar curve cluster and the original curve is calculated by Equation (18): In Equation (14), the Num(·) function represents obtaining the number of data in the sequence; 3) Proportional coefficients a, b The setting of the proportional coefficients a and b is related to the degree of emphasis on the average matching degree D(ε) and the average compression ratio Z(ε).

8. The multi-energy load short-term prediction method according to claim 1, characterized in that The probability transition chain model is as follows: Applying the probability transition chain model, it is first necessary to determine the size of the state space of the system under study; assuming that the category coding sequence of historical loads contains a total of r states, denoted as [s1, s2, …, s i , …, s j , …, s r , when the result of FCM clustering of the load is M categories, the value range of any state s i is [1, M]; according to the definition of the probability transition chain, the probability that state s i transfers to state s j after n steps is calculated by formula (21). P ij P(n) = T ij P(n) / T i (21) Where, T ij (n) represents the number of times the state s i in the historical state sequence transfers to the state s j after n steps; T i represents the total number of times the s i state appears in the historical state sequence; Repeatedly, the n-step state transition probability matrix of the probability transition chain can be calculated according to Equation (21), as shown in Equation (22) and Equation (23): When predicting the state at the (r + 1)-th step using the probability transition matrix, first find the state s of the r-th row of the probability transition matrix r with the maximum transition probability, assumed to be P ru (n), where u = 1, 2, …, r; then, according to the definition of the probability transition chain, determine that the state s u is the state with the maximum probability at the (r + 1)-th step.

Citation Information

Patent Citations

  • Comprehensive energy system regulation and control method and device based on multivariate load prediction

    CN112232575A