Method and device for establishing short-term probability prediction model of comprehensive energy load polymer
By constructing a model that integrates Copula adaptive correlation analysis, cross feature-time graph neural network and hybrid density network, the complex coupling relationship and uncertainty problems in the prediction of comprehensive energy load aggregates are solved, and a higher precision short-term probability prediction is achieved.
Patent Information
- Application Number
- CN202510520450.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-04-24
AI Technical Summary
The prior art fails to effectively consider the complex coupling relationship and uncertainty between multiple energy loads in the prediction of integrated energy load aggregates, resulting in insufficient prediction accuracy.
The data are grouped using affinity propagation clustering algorithm based on comprehensive similarity, and a comprehensive energy load aggregate short-term probability prediction model is constructed that integrates Copula adaptive correlation analysis, cross-feature-time graph neural network and hybrid density network. By mining the complex dynamic correlation and uncertainty between loads, prediction performance is improved.
It significantly improves the accuracy and robustness of the short-term probability prediction of comprehensive energy load polymers, can better cope with complex coupling relationships and uncertainties, and improves the engineering adaptability and promotion value of predictions.
Smart Images

Figure CN120448810A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of energy big data analysis, and in particular to a method and device for establishing a short-term probability prediction model for an integrated energy load aggregate. Background Art
[0002] With the rapid development of integrated energy systems, traditional demand response is gradually shifting towards integrated demand response, and the role of integrated energy load aggregates in supply and demand coordination is becoming increasingly important. Accurate forecasting of integrated energy load aggregates is not only related to the stable operation of the system but also plays a vital role in market planning and scheduling optimization. Especially in the context of multi-energy complementarity and regional collaborative optimization, load forecasting research has expanded from single energy sources or individual users to the load aggregate level. Improving forecast accuracy under complex coupling relationships and high uncertainty conditions has become a critical issue that needs to be addressed.
[0003] However, during research and practice on existing technologies, the inventors discovered that most current research focuses on forecasting for a single energy load aggregate. While this considers the aggregate effect of load aggregates, it fails to address the complex coupling relationships and uncertainties between multiple energy loads. Therefore, there is an urgent need to propose a short-term probabilistic forecasting method and apparatus for integrated energy load aggregates to more effectively address these complex coupling relationships and uncertainties. Summary of the Invention
[0004] The purpose of the present invention is to provide a short-term probabilistic forecasting model for an integrated energy load aggregate, so as to solve the problem in the prior art of insufficient ability to characterize the complex coupling relationship between integrated energy loads, the dynamic correlation with external influencing factors, and the uncertainty of the correlation relationship between each group load, thereby improving the overall forecasting performance.
[0005] The present invention provides a method for establishing a short-term probability forecasting model for a comprehensive energy load aggregate, the method comprising:
[0006] Preprocess and reduce the dimension of the collected historical data of comprehensive energy load aggregates;
[0007] The affinity propagation clustering algorithm based on comprehensive similarity is used to group the data after dimensionality reduction, and the centroid features of each group are extracted as representative load features to input into the prediction model;
[0008] A short-term probability forecasting model for integrated energy load aggregates is constructed by integrating Copula adaptive correlation analysis, cross-feature-time graph neural network, and hybrid density network. The model is used to perform short-term probability forecasting for integrated energy load aggregates.
[0009] Furthermore, the preprocessing and dimensionality reduction of the collected comprehensive energy load aggregate historical data includes:
[0010] Preprocessing the collected comprehensive energy load historical data, including data cleaning and normalization;
[0011] The preprocessed data is subjected to dimensionality reduction using the three-segment maximum triangle algorithm.
[0012] Furthermore, the affinity propagation clustering algorithm based on comprehensive similarity is used to group the data after dimensionality reduction, and the centroid features of each group are extracted as representative load features to input into the prediction model, including:
[0013] To address the problem that traditional clustering methods only measure similarity based on a single Euclidean distance, which makes it difficult to fully reflect the overall similarity of integrated energy users in multiple energy load curves such as electricity, cooling, heating, and gas, a comprehensive similarity measurement method that integrates Euclidean distance and cosine distance is proposed. The calculation formula of the comprehensive similarity distance is as follows:
[0014]
[0015] Where, d e represents the Euclidean distance, d cos Represents the cosine distance, and n represents the number of energy load types included in the user, taking into account the four dimensions of electric load, cooling load, heating load, and gas load.
[0016] Based on the similarity measurement method, a comprehensive similarity matrix between users is constructed. The similarity matrix is used as input to execute the affinity propagation clustering algorithm, and the centroid of each cluster is extracted as a representative load feature to be input into the prediction model.
[0017] Furthermore, the construction of a comprehensive energy load aggregate short-term probability forecasting model integrating Copula adaptive correlation analysis, cross-feature-time graph neural network and hybrid density network includes:
[0018] Through Copula adaptive correlation analysis, the complex dynamic correlations between comprehensive energy loads, between loads and external influencing factors, and between different groups are explored;
[0019] The cross-feature-time graph neural network is used to jointly model the complex coupling relationship between feature dimension, time dimension and group dimension, and combined with Bi-SRU for multivariate time series modeling;
[0020] The input features are nonlinearly mapped through the mixture density network, the mean, standard deviation and corresponding weight of each mixture component are estimated, and the combined mixture probability density distribution is output.
[0021] Furthermore, the Copula adaptive correlation analysis is used to explore the complex dynamic correlations between comprehensive energy loads, between loads and external influencing factors, and between different groups, including:
[0022] A directed graph structure based on load characteristics is constructed, adjacency relationships are defined to characterize the correlation between different load variables, the copula method is used to model the joint distribution between node pairs, and significant edges are retained based on the correlation threshold.
[0023] Construct a cross-feature graph, use the Copula function to model the joint distribution between load characteristics and external factors, build an initial feature correlation matrix, and obtain the weight matrix of graph construction through regularization processing. The specific formula is as follows:
[0024]
[0025] Where, E F [i,j] represents the feature node and The weight of the correlation between .
[0026] Construct a cross-time graph, obtain the dominant period information through frequency domain transformation and perform multi-scale sampling, use the Copula function to construct the initial time correlation matrix, and obtain the weight matrix for graph construction through regularization processing. The specific formula is as follows:
[0027]
[0028] Where, E T [i,j] is the time node and The weight of the correlation between .
[0029] Construct an inter-group association graph, using the centroids of each clustered group as graph nodes. Model the correlation between different groups based on the Copula function, and obtain the weight matrix for graph construction through regularization. The specific formula is as follows:
[0030]
[0031] Where, E C [i,j] is the centroid node and The weight of the correlation between .
[0032] Furthermore, the cross-feature-time graph neural network is used to jointly model the complex coupling relationship between the feature dimension, the time dimension, and the group dimension, and is combined with Bi-SRU to perform multivariate time series modeling, including:
[0033] By designing a message passing mechanism in the feature, time, and group dimensions, we model the dynamic interaction characteristics of load aggregates in terms of time sequence, features, and groups, and achieve adaptive learning of multi-dimensional dynamic dependencies.
[0034] In the feature dimension, the nonlinear interaction relationship between feature nodes is captured through the message passing mechanism in the graph neural network. For a feature node v, its message passing process can be expressed as:
[0035]
[0036] Where σ is the activation function, is the learnable parameter matrix, Represents the feature of feature node v at the kth layer, and They represent the neighbor sets that are homogeneous and heterogeneous with the feature node v, respectively, and Norm represents the normalization operation.
[0037] In the time dimension, the dynamic dependency relationship between time nodes is depicted by stacking K layers of graph neural network structures. The specific process is as follows:
[0038]
[0039] Where σ is the activation function, W i (k) is the learnable parameter matrix, represents the feature of time node i at the kth layer, N i is the set of time points adjacent to time node i, and Norm is the normalization operation.
[0040] In the group dimension, considering the complexity of load types within the load aggregate and the heterogeneity between groups, a separate message transmission mechanism is defined for each group. Assuming that the group node is represented as , its message transmission process can be expressed as:
[0041]
[0042] Where σ is the activation function, is the learnable parameter matrix, Represents the characteristics of group node c at the kth layer, N c They represent the neighbor sets of group node c respectively, and Norm is the normalization operation.
[0043] After completing the message transmission in the three dimensions of features, time, and groups, the feature-time graph neural network generates multi-step prediction results through the prediction model. The specific formula is as follows:
[0044] Y=f P (H final )
[0045] Where, f P (·) is the prediction model, where the Bi-SRU model is used, H final Representation features - comprehensive features output by temporal graph neural networks.
[0046] Furthermore, the nonlinear transformation of the input features is performed through the mixture density network, the mean, standard deviation and weight parameters of each basis distribution are estimated, and the corresponding mixture probability density distribution is output, including:
[0047] Assuming that each mixture component follows a Gaussian distribution, the mixture weight is modeled as a categorical distribution to satisfy the weight normalization condition. The mixture density network outputs the mean, standard deviation, and weight of each component, and the weighted summation forms the mixture probability distribution function.
[0048] The present invention provides a device for establishing a short-term probability forecasting model for a comprehensive energy load aggregate, comprising:
[0049] The preprocessing module is used to preprocess and reduce the dimension of the collected historical data of comprehensive energy load aggregates;
[0050] The cluster analysis module is used to group the data after dimensionality reduction using the affinity propagation clustering algorithm based on comprehensive similarity, and extract the centroid features of each group as representative load features to input into the prediction model;
[0051] A prediction model module is established to construct a short-term probability prediction model for integrated energy load aggregates that integrates Copula adaptive correlation analysis, a cross-feature-time graph neural network, and a hybrid density network. The short-term probability prediction model for integrated energy load aggregates is used to perform short-term probability prediction of integrated energy load aggregates.
[0052] The present invention provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the steps of the method for establishing a short-term probability prediction model of an integrated energy load aggregate are implemented.
[0053] The present invention also provides a computer-readable storage medium, on which an information transmission implementation program is stored. When the program is executed by a processor, the steps of the method for establishing the above-mentioned comprehensive energy load aggregate short-term probability prediction model are implemented.
[0054] Compared with the existing technology, the present invention fully considers the complex coupling relationship, uncertainty and dynamic correlation between multiple energy loads and external factors in the short-term probabilistic prediction of integrated energy load aggregates. By introducing the Copula adaptive correlation analysis method, it effectively mines the nonlinear dependency structure between different load types, between loads and external influencing factors, and between different groups; combined with the cross-feature-time graph neural network, it realizes the multidimensional coupling relationship modeling between feature dimensions, time dimensions and group dimensions, and uses the Bi-SRU structure to improve the efficiency and performance of multivariate time series modeling; at the same time, by characterizing the probability distribution of the prediction results through the mixed density network, the expressive ability of the model and the prediction uncertainty representation ability are enhanced, thereby significantly improving the prediction accuracy and robustness in complex load environments, and having stronger engineering adaptability and promotion value.
[0055] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are specifically listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0057] Figure 1 It is a flow chart of a method for establishing a short-term probability forecasting model for an integrated energy load aggregate according to an embodiment of the present invention;
[0058] Figure 2 Schematic diagram of a device for establishing a short-term probability forecasting model for an integrated energy load aggregate according to a first embodiment of the present invention;
[0059] Figure 3 It is a schematic diagram of a device for establishing a short-term probability prediction model of an integrated energy load aggregate according to the second embodiment of the device of the present invention. DETAILED DESCRIPTION
[0060] In order to make up for the shortcomings of the existing technology, the embodiments of the present invention propose a method and device for short-term probability prediction of integrated energy load aggregates, which are used to perform dimensionality reduction and cluster analysis on integrated energy load aggregate data and establish a short-term probability prediction model for integrated energy load aggregates.
[0061] The technical scheme of the present invention is clearly and completely described below in conjunction with the embodiments. It should be understood that the described embodiments are only a part of the present invention, rather than all embodiments. For those of ordinary skill in the art, various other embodiments made without departing from the concept of the present invention should be included in the scope of protection of the present invention.
[0062] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise" and the like to indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as limiting the present invention.
[0063] In addition, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the said features. In the description of the present invention, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined. In addition, the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be a communication between the two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0064] The technical solutions of the embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0065] The embodiment of the method of the present invention provides a method for short-term probability forecasting of integrated energy load aggregates based on an adaptive Copula cross-feature-time graph mixed density network model. The flowchart of the method is as follows: Figure 1 As shown, the specific processing includes the following:
[0066] Step 1 (S101) pre-processes and reduces the dimension of the collected historical data of the integrated energy load aggregate. Specifically, the following processing can be adopted:
[0067] Step 1.1: Preprocess the collected load data. The specific process is as follows:
[0068] (1) Clean the user load data to remove outliers and missing data;
[0069] (2) Normalize the user load data to ensure that the dimensions of each feature are consistent.
[0070] Step 1.2: Use the three-segment maximum triangle algorithm to reduce the dimension of the preprocessed data. The specific process is as follows:
[0071] The original time series data is divided into several equal segments, and in each segment, the most important point (denoted as point F) is selected based on the maximum effective area to represent all the points in the current segment to achieve dimensionality compression.
[0072] Step 2 (S102) uses affinity propagation clustering algorithm based on comprehensive similarity to group the data after dimensionality reduction, extracts the centroid features of each group as representative load features and inputs them into the prediction model. Specifically, the following processing can be adopted:
[0073] Step 2.1: Calculate the similarity between any two user energy consumption data sequences using a comprehensive similarity metric. The comprehensive similarity includes the following two aspects:
[0074] a. Distance metric: Euclidean distance is used to calculate the distance between energy consumption data sequences of different users;
[0075] b. Shape fluctuation measurement: cosine distance is used to measure the shape similarity of energy consumption data sequences between different users.
[0076] Combining the above Euclidean distance and cosine distance, the comprehensive similarity distance formula is constructed as follows:
[0077]
[0078] In the formula, n represents the number of energy load types contained in a single user, taking into account the four dimensions of electric load, cooling load, heating load, and gas load. e is the Euclidean distance, d cos is the cosine distance. The N*N comprehensive similarity matrix S is obtained by calculation (where N is the total number of particles), and s(j,j) is set as the median of the similarity matrix.
[0079] Step 2.2: Clustering is achieved through message passing. In this process, two types of messages are exchanged: attraction information and affiliation information. t (i, j) reflects the suitability of data point j as the cluster center of data point i at time t, indicating the message from i to j; the attribution information a t (i, j) reflects the suitability of data point i to select data point j as the cluster center at time t, and represents the message from j to i. Based on the above two types of information, the steps of message transmission are as follows:
[0080] Attract information t+1 (i, j) Iterate according to the following formula to calculate the attraction of data point i as the cluster center of data point j, taking into account the influence of all other cluster centers.
[0081] r t+1 (i,j)=s(i,j)-max{a t (i,j′)+s(i,j′)},stj′∈{1,2,…,N},j′≠j
[0082] Where s(i,j) represents the similarity between data point i and data point j. t (i, j′) represents the information that data point i selects data point j′ as the cluster center at time t. s(i, j′) represents the similarity between data point i and data point j′.
[0083] Attribution informationa t+1 (i, j) is iterated according to the following formula to determine whether data point i selects data point j as the cluster center.
[0084]
[0085] a t+1 (j,j)=∑ i′≠j max{0,r t (i′,j)}
[0086] Where r t (j,j) represents the attraction information between data point j and itself at time t, r t (i′, j) represents the attraction information between data point i′ and data point j at time t. During the algorithm's iterations, if the changes in the attraction information and the attribution information tend to be stable, that is, no longer change significantly, or the number of iterations exceeds the preset maximum number, the algorithm will terminate.
[0087] In order to avoid oscillation, the affinity propagation clustering algorithm introduces an attenuation coefficient λ (its value range is between 0 and 1) when updating information. The updated value of each piece of information is composed of the weighted average of the previous iteration value and the current iteration value, where the weight of the previous iteration value is the attenuation coefficient λ and the weight of the current iteration value is (1-λ). Then, in the t+1th iteration, the information r is attracted t+1 (i,j) and the attribution information a t+1 The updated values of (i,j) are:
[0088] r′ t+1 (i,j)=(1-λ)r t+1 (i,j)+λr t (i,j)
[0089] a′ t+1 (i,j)=(1-λ)a t+1 (i,j)+λa t (i,j)
[0090] Step 3 (S103) is to construct a comprehensive energy load aggregate short-term probability forecasting model that integrates Copula adaptive correlation analysis, cross-feature-time graph neural network and hybrid density network. Specifically, the following methods can be adopted:
[0091] Step 3.1: Use Copula adaptive correlation analysis to explore the complex dynamic correlations between comprehensive energy loads, between loads and external influencing factors, and between different groups. The specific process is as follows:
[0092] Assume that the input matrix of multivariate time series is χ=[X 1 ,X 2 ,...,X L ],in, D is the number of variables. Construct a directed graph structure, whose topological relationship is represented by the adjacency matrix A, where each element a ij The definition is as follows:
[0093]
[0094] Where w i and w j are the features of the i-th and j-th rows of χ, and g is the threshold for judging the significance of the correlation.
[0095] Constructing a directed graph G can be expressed as:
[0096] G=(V,E)
[0097] Where V={v1,v2,…,v n} is a node set containing n nodes. Represents an ordered edge set, consisting of ordered pairs (v i ,v j ) is composed of (v i ,v j ) represents the slave node v i Points to node v j A directed edge of .
[0098] The cross feature map is represented as G F =(V F ,E F ),in is the feature node set, D i is the number of features, Represents the jth feature node. Use the Copula function to model the correlation between any two nodes and construct the initial feature correlation matrix R F :
[0099] R F (i,j)=Copula(Z :,i ,Z :,j ),i,j∈{1,2,…,D i}
[0100] Perform Softmax normalization on the original correlation matrix to obtain the initialized feature correlation matrix
[0101]
[0102] set up Representation and Node The associated positive neighbor set, represents the set of negative neighbors, and the corresponding homogeneous and heterogeneous correlation weights are renormalized as follows:
[0103]
[0104] Where, E F [i,j] is and The correlation weights between them are calculated. Through this process, homogeneous edges and heterogeneous edges are retained, while other types of edges are removed. Homogeneous edge weights are positively correlated with the correlation score, while heterogeneous edge weights are negatively correlated with the correlation score. Finally, after normalization, a decomposed homogeneous and heterogeneous correlation cross-feature map is constructed.
[0105] The cross-time graph is represented as G T =(V T ,E T ),in is the set of time nodes, L h is the time range, represents the i-th time node, Represents the correlation weight matrix between time nodes. To effectively extract the potential features of multi-feature time series at different time scales and reduce the interference of noise on the correlation modeling between time nodes, we combine frequency domain analysis, Copula correlation modeling, and graph structure optimization strategies to dynamically adjust the adjacency relationship of time nodes, thereby constructing the adjacency matrix of the time graph.
[0106] Fast Fourier transform is used to extract the frequency domain features of the input time series, calculate the frequency domain amplitude M of each feature, and take the mean in the feature dimension:
[0107]
[0108] Where Amp(·) represents the amplitude calculation, represents the average amplitude of each frequency, and Q is the number of frequency components. Select the s frequency components with the largest amplitude, and their corresponding period length is Downsample the original sequence:
[0109] X (s) =AvgPool(X,kernel=q s ,stride=q s )
[0110] The downsampled sequence is in is the length of the downsampled sequence. The downsampled sequences at all time scales are concatenated to generate a multi-scale time feature sequence:
[0111] X′=Concat(X (1) ,X (2) ,...,X (s) )
[0112] Where, The total length after splicing is calculated by the Copula method to calculate the correlation between the time nodes of X′ and construct the initial time correlation matrix R T :
[0113] R T (i,j)=Copula(Z′ i ,Z′ j ),i,j∈{1,2,...,L′}
[0114] Get E through the original Copula correlation matrix T Initialization value of
[0115]
[0116] Where R T is the original Copula time correlation matrix, ReLU(·) is the regularization function, and Softmax(·) ensures that the sum of the weights of all nodes related to a specific time node is 1. In order to capture the temporal trend characteristics, the edge connection relationship between the time node and its adjacent previous and next nodes is retained. It is represented as a trend neighbor set of a time node and is defined as follows:
[0117]
[0118] Where, The trend neighbor set consists of adjacent time nodes that share the same scale (i.e., |ij| ≤ 1). The final correlation weights are renormalized as follows:
[0119]
[0120] Where, E T [i,j] is and This step removes connections with insignificant correlation and retains a restricted set of adjacent nodes for each node. Subsequently, the retained correlation weights are renormalized to complete the construction of the cross-temporal correlation map.
[0121] The intergroup correlation graph is represented by G C =(V C ,E C ),in is the cluster centroid node set, N represents the number of clusters, is the jth centroid node. Represents the correlation weight matrix between centroid nodes. Use the Copula method to calculate the correlation between centroid nodes and construct the initial inter-group correlation matrix R C :
[0122] R C (i,j)=Copula(C i ,C j ),i,j∈{1,2,…,N}
[0123] E is obtained through the original Copula correlation matrix C Initialization value of
[0124]
[0125] use Representation and Node The associated positive neighbor set, represents the set of negative neighbors, and the corresponding homogeneous and heterogeneous correlation weights are renormalized as follows:
[0126]
[0127] Where, E C [i,j] is and During the decoupling process, only the weights of homogeneous and heterogeneous edges are retained, and normalization is performed on them separately. Ultimately, a correlation graph between homogeneous and heterogeneous groups is constructed, effectively depicting the complex correlations between groups.
[0128] Step 3.2: Use the cross-feature-time graph neural network to jointly model the complex coupling relationship between the feature dimension, time dimension, and group dimension, and combine it with Bi-SRU to perform multivariate time series modeling. Specifically, it includes:
[0129] By designing a message passing mechanism in the feature, time and group dimensions, a comprehensive expression of the dynamic interaction characteristics of load aggregates in terms of timing, features and groups is achieved, thereby adaptively learning the multi-level complex dependency structure between various dimensions to enhance prediction performance.
[0130] In the feature dimension, the nonlinear interaction between feature nodes is captured through the message passing mechanism in the graph neural network. For a feature node v, its message passing process can be expressed as:
[0131]
[0132] Where σ is the activation function, is the learnable parameter matrix, Represents the feature of feature node v at the kth layer, and Denote the sets of neighbors that are homogeneous and heterogeneous with respect to feature node v, respectively, and Norm represents the normalization operation. This mechanism effectively models the complex relationships between features by aggregating the features of positive and negative neighbor nodes.
[0133] In the time dimension, the dynamic dependency relationship between time nodes is depicted by stacking K layers of graph neural network structures. The specific process is as follows:
[0134]
[0135] Where σ is the activation function, W i (k) is the learnable parameter matrix, represents the feature of time node i at the kth layer, N i is the set of time points adjacent to time node i, and Norm is the normalization operation. By aggregating the feature information of temporal neighbor nodes, the model updates the time node features layer by layer and finally outputs the results of K layers.
[0136] In the group dimension, considering the complex load types within the load aggregate and the significant heterogeneity between groups, the feature-time graph neural network defines a separate message passing mechanism for each group. Assuming that the group node is represented as , its message passing process can be expressed as:
[0137]
[0138] Where σ is the activation function, is the learnable parameter matrix, Represents the characteristics of group node c at the kth layer, N c where represents the neighbor set of group node c, and Norm is the normalization operation. Through the interaction between groups, the feature-time graph neural network can effectively capture the correlation between loads of different groups, thereby improving the overall prediction performance of the load aggregate.
[0139] After completing the message transmission in the three dimensions of features, time, and groups, the feature-time graph neural network generates multi-step prediction results through the prediction model. The specific formula is as follows:
[0140] Y=f P (H final )
[0141] Where, f P (·) is the prediction model, where the Bi-SRU model is used, H final Representation features - comprehensive features of the temporal graph neural network output. By modeling the temporal dependencies of hidden layer features, intermediate prediction results of the graph neural network are generated, and probabilistic load forecasting modeling is then carried out.
[0142] The forward pass calculation formula of the Bi-SRU model is as follows:
[0143] f t =σ(W f x t +V f ⊙c t-1 +b f )
[0144] r t =σ(W r x t +V r ⊙c t-1 +b r )
[0145] c t =f t ⊙c t-1 +(1-f t )⊙(W c x t )
[0146] h t =r t ⊙c t +(1-r t )⊙x t
[0147] Where, f t Represents the output of the forget gate at time t, which determines the information c from the previous moment t-1Whether to continue to pass, partially pass or remove; σ represents the Sigmoid activation function in the model; x t represents the input at time t; r t Represents the output of the reset gate at time t, which determines the information c from the previous moment t-1 How many are written to the current candidate set; c t represents the new information value generated at time t; h t represents the state of the hidden layer at time t; W f 、W r 、V f 、V r 、W c is the corresponding weight coefficient, b f 、b r is the corresponding bias term; the symbol ⊙ represents element-by-element multiplication, also known as the Hadamard product.
[0148] The Bi-SRU network model processes information in two directions: the forward SRU unit generates the previous information, that is, the hidden state in, The backward SRU unit generates subsequent information, that is, generates a hidden state in, By connecting these two pieces of information (forward information and backward information), the resulting joint information is expressed as:
[0149]
[0150] Where, and They represent the hidden states of the forward and backward SRU units at time t, respectively.
[0151] Step 3.3, nonlinearly map the input features through the mixture density network, estimate the mean, standard deviation and corresponding weight of each mixture component, and output the combined mixture probability density distribution, specifically including:
[0152] Assume that the conditional distribution of the mixture components comes from the Gaussian distribution family and set the mixture weight δ k ∈[0,1] M Modeled as a category distribution with M possible states. k satisfy The posterior distribution P(δ k |k) can be embedded in e through deterministic routing k To calculate:
[0153]
[0154] Where, is the linear model f δThe parameters of the mixture density network output include the mean, standard deviation, and weight of each basic distribution. By weighted summing these distributions, a mixed probability distribution function is finally formed, as shown below:
[0155]
[0156] Where μ m (k) and σ m (k) represents the mean and variance parameters of the mth mixture component respectively. The parameters of the output conditional density function can be further expressed as:
[0157]
[0158] Where, It is a linear model. The mixture density network can effectively handle the uncertainty in load forecasting and provide decision makers with more reliable load change expectations by giving the probability density function of each predicted value.
[0159] In other words, the specific process of short-term probabilistic forecasting of comprehensive energy load aggregates is as follows:
[0160] (1) Prediction process
[0161] The data set selected in this embodiment is the electricity load, cooling load, heating load and gas load data of 12,004 groups of residents in Arizona from the U.S. building data set, and the data range is from 0:00 on January 1, 2018 to 0:00 on January 1, 2019. The data set also includes hourly weather data at the corresponding time of the local time, covering external factors such as temperature, relative humidity, wind speed, wind direction, global horizontal radiation, direct normal radiation, diffuse horizontal radiation, etc., which may affect the load. In view of the limited amount of data, in order to make full use of the data and optimize the model performance, this study divides the data of each season into training set, validation set and test set in a ratio of 7:2:1. In terms of the parameters of the prediction model, the dropout is 0.01, the initial learning rate is 0.0001, the number of iterations is 100, and the batch size is 80. The model consists of 2 encoder layers and 1 decoder layer. The number of heads of the multi-head attention of each layer is 8, and the hidden layer dimension is 64.
[0162] (2) Prediction performance evaluation
[0163] In order to evaluate the prediction performance, Pinball loss function and Winkler score were selected to evaluate the effect of the probability prediction model.
[0164] The Pinball loss calculation formula is as follows:
[0165]
[0166] Where, represents the qth quantile load forecast value of the i-th sample, y i is the actual observed load value. The smaller the pinball loss score, the better the prediction effect.
[0167] The Winkler score calculation formula is as follows:
[0168]
[0169] Where, δ i =U i -L i Indicates the prediction interval width of the i-th sample, U i and L i are the upper and lower bounds, respectively. When the actual value falls within the prediction interval, the Winkler score depends only on the response width; otherwise, a penalty term is added to reflect the deviation of the uncovered actual value. The smaller the Winkler score, the better the prediction effect.
[0170] In this embodiment, the method in this paper (denoted as ACCFTG-MDN) is used to compare and analyze the prediction results of six prediction models, including the adaptive Copula cross-feature-time graph quantile model (denoted as ACCFTG-MQ), the adaptive Copula cross-feature-time graph Gaussian model (denoted as ACCFTG-Gaussian), the adaptive Copula cross-feature-time graph Bayesian model (denoted as ACCFTG-Bayes), the adaptive Copula cross-feature-time graph convolutional mixture density network model (denoted as ACCFTGCN-MDN), the adaptive Copula cross-feature-time graph recurrent mixture density network model (denoted as ACCFTGRN-MDN), and the adaptive Copula cross-feature-time convolution attention mechanism bidirectional gated recurrent mixture density network model (denoted as ACC-CNN-BiGRU-ATT-MDN). The analysis results are shown in Table 1.
[0171] Table 1
[0172]
[0173]
[0174] In summary, the embodiment of the present invention proposes a method for short-term probability prediction of integrated energy load aggregates, which performs dimensionality reduction through a three-segment maximum triangle algorithm, and utilizes an affinity propagation clustering algorithm based on comprehensive similarity to group load aggregates and extract centroid features. By constructing a fusion Copula adaptive correlation analysis, feature-time graph neural network and mixed density network model, the accuracy of short-term probability prediction of integrated energy load aggregates is effectively improved.
[0175] Device Example 1
[0176] According to an embodiment of the present invention, a device for establishing a short-term probability forecasting model for an integrated energy load aggregate is provided. Figure 2 Schematic diagram of a device for establishing a short-term probability forecasting model for an integrated energy load aggregate according to an embodiment of the present invention. Figure 2 As shown, the apparatus for establishing a short-term probability forecast model for an integrated energy load aggregate according to an embodiment of the present invention specifically includes: a pre-processing module, a cluster analysis module, and a forecast model establishment module, thereby obtaining the result of a short-term probability forecast for an integrated energy load aggregate. Specifically:
[0177] The pre-processing module 60 is used to pre-process and reduce the dimension of the collected comprehensive energy load aggregate historical data. The pre-processing module 60 is specifically used to:
[0178] Preprocessing the collected comprehensive energy load historical data, including data cleaning and normalization;
[0179] The preprocessed data is subjected to dimensionality reduction using the three-segment maximum triangle algorithm.
[0180] The cluster analysis module 62 is used to group the data after dimensionality reduction using an affinity propagation clustering algorithm based on comprehensive similarity, and extract the centroid features of each group as representative load features to input into the prediction model. The cluster analysis module 62 is specifically used to:
[0181] Calculate the comprehensive similarity measure between any two user energy consumption data sequences,
[0182] Clustering is achieved using affinity propagation clustering algorithm.
[0183] Furthermore, the calculation of the comprehensive similarity measure between any two user energy consumption data sequences specifically includes:
[0184] The comprehensive similarity includes the following two aspects:
[0185] a. Distance metric: Euclidean distance is used to calculate the distance between energy consumption data sequences of different users;
[0186] b. Shape fluctuation measurement: cosine distance is used to measure the shape similarity of energy consumption data sequences between different users.
[0187] Combining the above Euclidean distance and cosine distance, the comprehensive similarity distance formula is constructed as follows:
[0188]
[0189] In the formula, n represents the number of energy load types contained in a single user, taking into account the four dimensions of electric load, cooling load, heating load, and gas load. e is the Euclidean distance, d cos is the cosine distance. The N*N comprehensive similarity matrix S is obtained by calculation (where N is the total number of particles), and s(j,j) is set as the median of the similarity matrix.
[0190] Furthermore, the clustering is implemented by using the affinity propagation clustering algorithm, specifically including:
[0191] Clustering is achieved through message passing. In this process, two types of messages are exchanged: attraction information and affiliation information. t (i, j) reflects the suitability of data point j as the cluster center of data point i at time t, indicating the message from i to j; the attribution information a t (i, j) reflects the suitability of data point i to select data point j as the cluster center at time t, and represents the message from j to i. Based on the above two types of information, the steps of message transmission are as follows:
[0192] Attract information t+1 (i, j) Iterate according to the following formula to calculate the attraction of data point i as the cluster center of data point j, taking into account the influence of all other cluster centers.
[0193] r t+1 (i,j)=s(i,j)-max{a t (i,j′)+s(i,j′)},stj′∈{1,2,…,N},j′≠j
[0194] Where s(i,j) represents the similarity between data point i and data point j. t (i, j′) represents the information that data point i selects data point j′ as the cluster center at time t. s(i, j′) represents the similarity between data point i and data point j′.
[0195] Attribution informationa t+1 (i, j) is iterated according to the following formula to determine whether data point i selects data point j as the cluster center.
[0196]
[0197] a t+1 (j,j)=Σ i′≠j max{0,r t (i′,j)}
[0198] Where r t (j,j) represents the attraction information between data point j and itself at time t, r t (i′, j) represents the attraction information between data point i′ and data point j at time t. During the algorithm's iterations, if the changes in the attraction information and the attribution information tend to be stable, that is, no longer change significantly, or the number of iterations exceeds the preset maximum number, the algorithm will terminate.
[0199] In order to avoid oscillation, the affinity propagation clustering algorithm introduces an attenuation coefficient λ (its value range is between 0 and 1) when updating information. The updated value of each piece of information is composed of the weighted average of the previous iteration value and the current iteration value, where the weight of the previous iteration value is the attenuation coefficient λ and the weight of the current iteration value is (1-λ). Then, in the t+1th iteration, the information r is attracted t+1 (i,j) and the attribution information a t+1 The updated values of (i,j) are:
[0200] r′ t+1 (i,j)=(1-λ)r t+1 (i,j)+λr t (i,j)
[0201] a′ t+1 (i,j)=(1-λ)a t+1 (i,j)+λa t (i,j)
[0202] The prediction model building module 64 is used to build a short-term probability prediction model for an integrated energy load aggregate that integrates Copula adaptive correlation analysis, a cross-feature-time graph neural network, and a hybrid density network. The short-term probability prediction model for an integrated energy load aggregate is used to perform short-term probability prediction of an integrated energy load aggregate. The prediction model building module 64 is specifically used to:
[0203] Through Copula adaptive correlation analysis, the complex dynamic correlations between comprehensive energy loads, between loads and external influencing factors, and between different groups are explored;
[0204] The cross-feature-time graph neural network is used to jointly model the complex coupling relationship between feature dimension, time dimension and group dimension, and combined with Bi-SRU for multivariate time series modeling;
[0205] The input features are nonlinearly mapped through the mixture density network, the mean, standard deviation and corresponding weight of each mixture component are estimated, and the combined mixture probability density distribution is output.
[0206] Furthermore, the Copula adaptive correlation analysis is used to explore the complex dynamic correlations between comprehensive energy loads, between loads and external influencing factors, and between different groups. The specific process is as follows:
[0207] Assume that the input matrix of multivariate time series is χ=[X 1 ,X 2 ,...,X L ],in, D is the number of variables. Construct a directed graph structure, whose topological relationship is represented by the adjacency matrix A, where each element a ij The definition is as follows:
[0208]
[0209] Where w i and w j are the features of the i-th and j-th rows of χ, and g is the threshold for judging the significance of the correlation.
[0210] Constructing a directed graph G can be expressed as:
[0211] G=(V,E)
[0212] Where V={v1,v2,...,v n} is a node set containing n nodes. Represents an ordered edge set, consisting of ordered pairs (v i ,v j ) is composed of (v i ,v j ) represents the slave node v i Points to node v j A directed edge of .
[0213] The cross feature map is represented as G F =(V F ,E F ),in is the feature node set, D i is the number of features, Represents the jth feature node. Use the Copula function to model the correlation between any two nodes and construct the initial feature correlation matrix R F :
[0214] R F (i,j)=Copula(Z :,i ,Z :,j),i,j∈{1,2,...,D i}
[0215] Perform Softmax normalization on the original correlation matrix to obtain the initialized feature correlation matrix
[0216]
[0217] set up Representation and Node The associated positive neighbor set, represents the set of negative neighbors, and the corresponding homogeneous and heterogeneous correlation weights are renormalized as follows:
[0218]
[0219] Where, E F [i,j] is and The correlation weights between them are calculated. Through this process, homogeneous edges and heterogeneous edges are retained, while other types of edges are removed. Homogeneous edge weights are positively correlated with the correlation score, while heterogeneous edge weights are negatively correlated with the correlation score. Finally, after normalization, a decomposed homogeneous and heterogeneous correlation cross-feature map is constructed.
[0220] The cross-time graph is represented as G T =(V T ,E T ),in is the set of time nodes, L h is the time range, represents the i-th time node, Represents the correlation weight matrix between time nodes. To effectively extract the potential features of multi-feature time series at different time scales and reduce the interference of noise on the correlation modeling between time nodes, we combine frequency domain analysis, Copula correlation modeling, and graph structure optimization strategies to dynamically adjust the adjacency relationship of time nodes, thereby constructing the adjacency matrix of the time graph.
[0221] Fast Fourier transform is used to extract the frequency domain features of the input time series, calculate the frequency domain amplitude M of each feature, and take the mean in the feature dimension:
[0222]
[0223] Where Amp(·) represents the amplitude calculation, represents the average amplitude of each frequency, and Q is the number of frequency components. Select the s frequency components with the largest amplitude, and their corresponding period length is Downsample the original sequence:
[0224] X(s) =AvgPool(X,kernel=q s ,stride=q s )
[0225] The downsampled sequence is in is the length of the downsampled sequence. The downsampled sequences at all time scales are concatenated to generate a multi-scale time feature sequence:
[0226] X′=Concat(X (1) ,X (2) ,...,X (s) )
[0227] Where, The total length after splicing is calculated by the Copula method to calculate the correlation between the time nodes of X′ and construct the initial time correlation matrix R T :
[0228] R T (i,j)=Copula(Z′ i ,Z′ j ),i,j∈{1,2,...,L′}
[0229] Get E through the original Copula correlation matrix T Initialization value of
[0230]
[0231] Where R T is the original Copula time correlation matrix, ReLU(·) is the regularization function, and Softmax(·) ensures that the sum of the weights of all nodes related to a specific time node is 1. In order to capture the temporal trend characteristics, the edge connection relationship between the time node and its adjacent previous and next nodes is retained. It is represented as a trend neighbor set of a time node and is defined as follows:
[0232]
[0233] Where, The trend neighbor set consists of adjacent time nodes that share the same scale (i.e., |ij| ≤ 1). The final correlation weights are renormalized as follows:
[0234]
[0235] Where, E T [i,j] is and This step removes connections with insignificant correlation and retains a restricted set of adjacent nodes for each node. Subsequently, the retained correlation weights are renormalized to complete the construction of the cross-temporal correlation map.
[0236] The intergroup correlation graph is represented by G C =(V C ,E C ),in is the cluster centroid node set, N represents the number of clusters, is the jth centroid node. Represents the correlation weight matrix between centroid nodes. Use the Copula method to calculate the correlation between centroid nodes and construct the initial inter-group correlation matrix R C :
[0237] R C (i,j)=Copula(C i ,C j ),i,j∈{1,2,...,N}
[0238] E is obtained through the original Copula correlation matrix C Initialization value of
[0239]
[0240] use Representation and Node The associated positive neighbor set, represents the set of negative neighbors, and the corresponding homogeneous and heterogeneous correlation weights are renormalized as follows:
[0241]
[0242] Where, E C [i,j] is and During the decoupling process, only the weights of homogeneous and heterogeneous edges are retained, and normalization is performed on them separately. Ultimately, a correlation graph between homogeneous and heterogeneous groups is constructed, effectively depicting the complex correlations between groups.
[0243] Furthermore, the cross-feature-time graph neural network is used to jointly model the complex coupling relationship between the feature dimension, the time dimension, and the group dimension, and is combined with Bi-SRU to perform multivariate time series modeling, specifically including:
[0244] By designing a message passing mechanism in the feature, time and group dimensions, a comprehensive expression of the dynamic interaction characteristics of load aggregates in terms of timing, features and groups is achieved, thereby adaptively learning the multi-level complex dependency structure between various dimensions to enhance prediction performance.
[0245] In the feature dimension, the nonlinear interaction between feature nodes is captured through the message passing mechanism in the graph neural network. For a feature node v, its message passing process can be expressed as:
[0246]
[0247] Where σ is the activation function, is the learnable parameter matrix, Represents the feature of feature node v at the kth layer, and Denote the sets of neighbors that are homogeneous and heterogeneous with respect to feature node v, respectively, and Norm represents the normalization operation. This mechanism effectively models the complex relationships between features by aggregating the features of positive and negative neighbor nodes.
[0248] In the time dimension, the dynamic dependency relationship between time nodes is depicted by stacking K layers of graph neural network structures. The specific process is as follows:
[0249]
[0250] Where σ is the activation function, W i (k) is the learnable parameter matrix, represents the feature of time node i at the kth layer, N i is the set of time points adjacent to time node i, and Norm is the normalization operation. By aggregating the feature information of temporal neighbor nodes, the model updates the time node features layer by layer and finally outputs the results of K layers.
[0251] In the group dimension, considering the complex load types within the load aggregate and the significant heterogeneity between groups, the feature-time graph neural network defines a separate message passing mechanism for each group. Assuming that the group node is represented as , its message passing process can be expressed as:
[0252]
[0253] Where σ is the activation function, is the learnable parameter matrix, Represents the characteristics of group node c at the kth layer, N cwhere represents the neighbor set of group node c, and Norm is the normalization operation. Through the interaction between groups, the feature-time graph neural network can effectively capture the correlation between loads of different groups, thereby improving the overall prediction performance of the load aggregate.
[0254] After completing the message transmission in the three dimensions of features, time, and groups, the feature-time graph neural network generates multi-step prediction results through the prediction model. The specific formula is as follows:
[0255] Y=f P (H final )
[0256] Where, f P (·) is the prediction model, where the Bi-SRU model is used, H final Representation features - comprehensive features of the temporal graph neural network output. By modeling the temporal dependencies of hidden layer features, intermediate prediction results of the graph neural network are generated, and probabilistic load forecasting modeling is then carried out.
[0257] The forward pass calculation formula of the Bi-SRU model is as follows:
[0258] f t =σ(W f x t +V f ⊙c t-1 +b f )
[0259] r t =σ(W r x t +V r ⊙c t-1 +b r )
[0260] c t =f t ⊙c t-1 +(1-f t )⊙(W c x t )
[0261] h t =r t ⊙c t +(1-r t )⊙x t
[0262] Where, f t Represents the output of the forget gate at time t, which determines the information c from the previous moment t-1 Whether to continue to pass, partially pass or remove; σ represents the Sigmoid activation function in the model; x trepresents the input at time t; r t Represents the output of the reset gate at time t, which determines the information c from the previous moment t-1 How many are written to the current candidate set; c t represents the new information value generated at time t; h t represents the state of the hidden layer at time t; W f 、W r 、V f 、V r 、W c is the corresponding weight coefficient, b f 、b r is the corresponding bias term; the symbol ⊙ represents element-by-element multiplication, also known as the Hadamard product.
[0263] The Bi-SRU network model processes information in two directions: the forward SRU unit generates the previous information, that is, the hidden state in, The backward SRU unit generates subsequent information, that is, generates a hidden state in, By connecting these two pieces of information (forward information and backward information), the resulting joint information is expressed as:
[0264]
[0265] Where, and They represent the hidden states of the forward and backward SRU units at time t, respectively.
[0266] Furthermore, the nonlinear mapping of the input features through the mixture density network is performed, the mean, standard deviation and corresponding weight of each mixture component are estimated, and the combined mixture probability density distribution is output, specifically including:
[0267] Assume that the conditional distribution of the mixture components comes from the Gaussian distribution family and set the mixture weight δ k ∈[0,1] M Modeled as a category distribution with M possible states. k satisfy The posterior distribution P(δ k |k) can be embedded in e through deterministic routing k To calculate:
[0268]
[0269] Where, is the linear model f δThe parameters of the mixture density network output include the mean, standard deviation, and weight of each basic distribution. By weighted summing these distributions, a mixed probability distribution function is finally formed, as shown below:
[0270]
[0271] Where μ m (k) and σ m (k) represents the mean and variance parameters of the mth mixture component respectively. The parameters of the output conditional density function can be further expressed as:
[0272]
[0273] Where, It is a linear model. The mixture density network can effectively handle the uncertainty in load forecasting and provide decision makers with more reliable load change expectations by giving the probability density function of each predicted value.
[0274] The embodiment of the present invention is an apparatus embodiment corresponding to the above-mentioned method embodiment. The specific operations of each module can be understood by referring to the description of the method embodiment, which will not be repeated here.
[0275] Device Example 2
[0276] An embodiment of the present invention provides an electronic device, such as Figure 3 As shown, it includes: a memory 70, a processor 72 and a computer program stored in the memory 70 and executable on the processor 72. When the computer program is executed by the processor 72, the steps described in the method embodiment are implemented.
[0277] Device Example 3
[0278] An embodiment of the present invention provides a computer-readable storage medium, on which a program for implementing information transmission is stored. When the program is executed by the processor 72, the steps described in the method embodiment are implemented.
[0279] The computer-readable storage medium in this embodiment includes but is not limited to: ROM, RAM, magnetic disk or optical disk, etc.
[0280] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0281] In the 1930s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly done using "logic compiler" software. This is similar to the software compiler used when developing programs. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.
[0282] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules that implement the method and structures within the hardware component.
[0283] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0284] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing the embodiments of this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0285] Those skilled in the art will appreciate that one or more embodiments of this specification may be provided as a method, system, or computer program product. Thus, one or more embodiments of this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0286] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0287] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0288] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0289] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0290] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0291] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0292] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0293] One or more embodiments of this specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0294] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.
[0295] The foregoing description is merely an example of the present invention and is not intended to limit the present invention. Persons skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims herein.
Claims
1. A method for establishing a short-term probability forecasting model for a comprehensive energy load aggregate, characterized in that: include: S101 preprocesses and reduces the dimension of the collected historical data of comprehensive energy load aggregates; S102 uses affinity propagation clustering algorithm based on comprehensive similarity to group the data after dimensionality reduction, and extracts the centroid features of each group as representative load features to input into the prediction model; S103 constructs a comprehensive energy load aggregate short-term probability forecasting model that integrates Copula adaptive correlation analysis, cross-feature-time graph neural network and hybrid density network; wherein the comprehensive energy load aggregate short-term probability forecasting model is used to perform comprehensive energy load aggregate short-term probability forecasting.
2. The method for establishing a short-term probability forecasting model for a comprehensive energy load aggregate according to claim 1 is characterized in that: The affinity propagation clustering algorithm based on comprehensive similarity is used to group the data after dimensionality reduction, and the centroid features of each group are extracted as representative load features to input into the prediction model, including: A comprehensive similarity measurement method that combines Euclidean distance and cosine distance is proposed. The calculation formula of the comprehensive similarity distance is as follows: Where, d e represents the Euclidean distance, d cos represents the cosine distance, n represents the number of energy load types included in the user, taking into account the four dimensions of electric load, cooling load, heating load, and gas load; Based on the similarity measurement method, a comprehensive similarity matrix between users is constructed; the similarity matrix is used as input, an affinity propagation clustering algorithm is executed, and the centroid of each cluster is extracted as a representative load feature input into a prediction model.
3. The method for establishing a short-term probability forecasting model for a comprehensive energy load aggregate according to claim 1, characterized in that: A comprehensive energy load aggregate short-term probabilistic forecasting model is constructed by integrating Copula adaptive correlation analysis, cross-feature-time graph neural network, and hybrid density network. Specifically, it includes: Through Copula adaptive correlation analysis, the complex dynamic correlations between comprehensive energy loads, between loads and external influencing factors, and between different groups are explored; The cross-feature-time graph neural network is used to jointly model the complex coupling relationship between feature dimension, time dimension and group dimension, and combined with Bi-SRU for multivariate time series modeling; The input features are nonlinearly mapped through the mixture density network, the mean, standard deviation and corresponding weight of each mixture component are estimated, and the combined mixture probability density distribution is output.
4. The method for establishing a short-term probability forecasting model for a comprehensive energy load aggregate according to claim 3 is characterized in that: Through Copula adaptive correlation analysis, we can explore the complex dynamic correlations between comprehensive energy loads, between loads and external influencing factors, and between different groups, including: A directed graph structure based on load characteristics is constructed, adjacency relationships are defined to characterize the correlation between different load variables, the copula method is used to model the joint distribution between node pairs, and significant edges are retained based on the correlation threshold. Construct a cross-feature graph, use the Copula function to model the joint distribution between load characteristics and external factors, build an initial feature correlation matrix, and obtain the weight matrix of graph construction through regularization processing. The specific formula is as follows: Where, E F [i,j] represents the feature node and The correlation weight between them; Construct a cross-time graph, obtain the dominant period information through frequency domain transformation and perform multi-scale sampling, use the Copula function to construct the initial time correlation matrix, and obtain the weight matrix for graph construction through regularization processing. The specific formula is as follows: Where, E T [i,j] is the time node and The correlation weight between them; Construct an inter-group association graph, using the centroids of each clustered group as graph nodes. Model the correlation between different groups based on the Copula function, and obtain the weight matrix for graph construction through regularization. The specific formula is as follows: Where, E C [i,j] is the centroid node and The weight of the correlation between .
5. The method for establishing a short-term probability forecasting model for a comprehensive energy load aggregate according to claim 3 is characterized in that: The cross-feature-time graph neural network is used to jointly model the complex coupling relationship between feature dimensions, time dimensions, and group dimensions, and combined with Bi-SRU for multivariate time series modeling, specifically including: By designing a message passing mechanism in the feature, time, and group dimensions, we model the dynamic interaction characteristics of load aggregates in terms of time sequence, features, and groups, and achieve adaptive learning of multi-dimensional dynamic dependencies. In the feature dimension, the nonlinear interaction relationship between feature nodes is captured through the message passing mechanism in the graph neural network. For a feature node v, its message passing process is expressed as: Where σ is the activation function, is the learnable parameter matrix, Represents the feature of feature node v at the kth layer, and They represent the neighbor sets that are homogeneous and heterogeneous with the feature node v, respectively. Norm represents the normalization operation. In the time dimension, the dynamic dependency relationship between time nodes is characterized by stacking K layers of graph neural network structure. The specific process is as follows: Where σ is the activation function, W i (k) is the learnable parameter matrix, represents the feature of time node i at the kth layer, N i is the set of time points adjacent to time node i, and Norm is the normalization operation; In the group dimension, considering the complexity of load types within the load aggregate and the heterogeneity between groups, a separate message transmission mechanism is defined for each group. Assuming that the group node is represented as , its message transmission process is expressed as: Where σ is the activation function, W c (k) is the learnable parameter matrix, Represents the characteristics of group node c at the kth layer, N c They represent the neighbor sets of group node c respectively, and Norm is the normalization operation; After completing the message transmission in the three dimensions of features, time, and groups, the feature-time graph neural network generates multi-step prediction results through the prediction model. The specific formula is as follows: Y=f P (H final ) Where, f P (·) is the prediction model, where the Bi-SRU model is used, H final Representation features - comprehensive features output by temporal graph neural networks.
6. The method for establishing a short-term probability forecasting model for a comprehensive energy load aggregate according to claim 3, characterized in that: The mixture density network performs nonlinear mapping on the input features, estimates the mean, standard deviation and corresponding weight of each mixture component, and outputs the combined mixture probability density distribution, including: Assuming that each mixture component obeys a Gaussian distribution, the mixture weight is modeled as a category distribution to meet the weight normalization condition; the mixture density network outputs the mean, standard deviation and weight of each component, and the weighted sum is used to form a mixture probability distribution function.
7. A device for establishing a short-term probability forecasting model for an integrated energy load aggregate, characterized in that: include: The preprocessing module is used to preprocess and reduce the dimension of the collected historical data of comprehensive energy load aggregates; The cluster analysis module is used to group the data after dimensionality reduction using the affinity propagation clustering algorithm based on comprehensive similarity, and extract the centroid features of each group as representative load features to input into the prediction model; A prediction model module is established to construct a comprehensive energy load aggregate short-term probability prediction model that integrates Copula adaptive correlation analysis, cross-feature-time graph neural network and hybrid density network.
8. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and capable of running on the processor, wherein when the computer program is executed by the processor, the steps of the method for establishing a short-term probabilistic forecasting model for an integrated energy load aggregate are implemented as described in any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores an implementation program for information transmission, and when the program is executed by the processor, the steps of the method for establishing a short-term probability forecasting model for an integrated energy load aggregate as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Method and device for establishing load prediction model of integrated energy system
CN113822482A
Comprehensive load short-term prediction method based on coupling characteristic matrix time sequence fragment analysis
CN115907118A
Comprehensive energy system multi-element load short-term prediction method and system
CN117674117A
Short-term power load prediction method and system considering multi-dimensional feature extraction
CN119448225A
Probability estimation method for photovoltaic power based on optimized copula function and photovoltaic power system
US20240014651A1
Cited By
Method for predicting electrical load of user
CN121813337A
A user power consumption load prediction method
CN121813337B