Power consumer price-quantity curve aggregation method based on Gaussian mixture model and Gaussian process regression

Through a two-stage method of mixing Gaussian model and Gaussian process regression, the simplification and aggregation of the massive volume-value curves in the power market are solved. The generated segmented step curve can better reflect the characteristics of user price response, and improve data analysis efficiency and decision-making accuracy.

CN120336694APending Publication Date: 2025-07-18CHONGQING UNIV OF TECH +2
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510461484.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the power market, as the number of users increases, the shape characteristics of the massive price-volume curve vary greatly, which leads to difficulty in information processing and optimization calculation. The existing methods are not effective when the data volume is large, making it difficult to effectively simplify and aggregate.

Method used

The two-stage method of mixed Gaussian model and Gaussian process regression is adopted. First, the price-quantity curve is classified through mixed Gaussian model (GMM), and then the regression aggregation is used using Gaussian process regression (GPR), and the segmented step linearization is performed to generate a segmented step curve that conforms to the power market.

Benefits of technology

It reduces the dimension and complexity of user data, improves the efficiency of data analysis and clearing decision-making, and the generated alternative curve is highly similar to the original curve, which truly reflects the characteristics of user market price response and has better adaptability and flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336694A_ABST
    Figure CN120336694A_ABST
Patent Text Reader

Abstract

The invention relates to a power consumer price-quantity curve aggregation method based on a Gaussian mixture model (GMM) and Gaussian process regression, which belongs to the technical field of power and comprises the following steps: S1, dividing a price-quantity curve of a consumer into a plurality of key categories according to key features of the price-quantity curve by using a GMM algorithm; s2, performing regression aggregation on the curve of each category by using a Gaussian process regression GPR algorithm; and S3, performing segmented step linearization processing on the obtained regression polymerization curve to obtain a segmented step curve conforming to the power market expression form.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of electric power technology and relates to an electric power user price-volume curve aggregation method based on a mixed Gaussian model and Gaussian process regression. Background Art

[0002] It is an important step to realize the transaction for users to declare the price-volume curve to the power trading center. However, the large-scale direct participation of users in the inter-provincial spot market has brought great challenges to the processing of the declared price-volume curve information. With the large increase in the number of users participating in the market, the number of price-volume curves they declare has increased dramatically, and the shape characteristics of each curve are very different, which makes the information processing and optimization calculation of the clearing model more difficult. Therefore, how to aggregate and simplify the original massive price-volume curve declarations with different characteristics under the premise of restoring the true intention characteristics of each trading entity to participate in the market as much as possible is an important problem that needs to be solved urgently. The meaning and function of the aggregation and simplification of price-volume curves is to classify curves with similar characteristics in the massive price-volume curves, and aggregate the same type of curves into an alternative curve that can retain the main characteristics of the original data. The power trading center only needs to process the aggregated alternative curve information, which significantly reduces the data complexity.

[0003] Some existing studies have made some progress in the processing methods of price-volume curves of power users. Some literatures have proposed a model for analyzing the bidding behavior of the power spot market based on the mean clustering algorithm, aiming to effectively identify market power manipulation behavior through new indicators and improved algorithms, and provide new perspectives and methods in theoretical models and indicator systems. Other studies have revealed the diverse bidding behaviors of users through clustering analysis, and then proposed methods to optimize demand response strategies. However, the above studies mainly focus on the identification and optimization of bidding strategies, rather than the simplified processing of price-volume curve declarations. Another part of the research uses methods such as spectral clustering to cluster time series data such as power load, electricity price, power generation, and grid frequency in the power market, so as to better capture the time-varying characteristics of user behavior. These methods have achieved some results in clustering analysis and capturing time-varying characteristics, but they are often not effective when processing large amounts of data and large differences in the shape characteristics of data curves. Summary of the invention

[0004] In view of this, the purpose of this invention is to propose a two-stage power user price-volume curve aggregation method to address the problem that the number of users participating in the current power market quotation is increasing, and the power trading center is insufficient in processing massive price-volume curve information. The proposed method can reduce the dimension and complexity of user data, so that the trading center can only process the simplified alternative curve after aggregation, thereby improving the efficiency of data analysis and clearing decision-making.

[0005] In order to achieve the above object, the present invention provides the following technical solutions:

[0006] A method for aggregating price-volume curves of power users based on a mixed Gaussian model and Gaussian process regression comprises the following steps:

[0007] S1: Use the Gaussian Mixture Model (GMM) algorithm to classify the user's price-volume curve into several key categories according to their key features;

[0008] S2: Use Gaussian process regression (GPR) algorithm to regress and aggregate the curves of each category;

[0009] S3: Perform piecewise step linearization on the obtained regression aggregation curve to obtain a piecewise step curve that conforms to the form of electricity market expression.

[0010] Furthermore, the user's price-volume curve is generated by simulation using the Monte Carlo sampling method.

[0011] Furthermore, the mathematical representation of the Gaussian mixture model GMM is as follows:

[0012]

[0013] Where p(x) is the probability density of data point x; K is the number of components in the mixed distribution, that is, the number of Gaussian distributions; π k is the weight of the kth Gaussian distribution, satisfying Represents the contribution of each distribution to the overall model; N(x丨μ k ,∑ k ) is the probability density function of the kth Gaussian distribution with mean μ k and the covariance matrix ∑ k .

[0014] Furthermore, the key features include slope, intercept, quantity-price correlation coefficient, quantity-price slope change rate and regression residual variance;

[0015] The slope Slope is expressed as:

[0016]

[0017] Among them, x i represents the electricity price, y i represents the power demand, represents the mean value of electricity price, represents the mean value of electricity demand, n is the number of data points;

[0018] The intercept is expressed as:

[0019]

[0020] The quantity-price correlation coefficient C P,Q is expressed as:

[0021]

[0022] C P,Q has a value range of [-1, 1];

[0023] The rate of change slope of the quantity-price, Rate of Change Slope, is expressed as:

[0024]

[0025] where Slope i is the local slope of each point, and is the average slope of all points;

[0026] The regression residual variance, Residual Variance, is expressed as:

[0027]

[0028] where y i represents the actual electricity demand, represents the predicted value of the Gaussian process regression model.

[0029] Furthermore, in the regression residual variance, the predicted value of the Gaussian process regression model needs to be iteratively optimized in combination with Gaussian process regression. The regression residual variance of the electricity quantity optimized by Gaussian process regression is used as the feature input of a mixture Gaussian model. The specific iterative process is as follows:

[0030] First, select a kernel function using the existing data and train the Gaussian process regression model. Then, subtract the predicted value obtained from the actual value to get the residual. Next, accumulate the squared residuals and take the average to get the residual variance. If the obtained residual variance value exceeds the preset threshold, enter the model optimization stage. Check the change trend of the residual variance by optimizing the Gaussian process regression model until it is less than the preset threshold. Use the value obtained at this time as the feature input of the mixture Gaussian model.

[0031] Furthermore, in step S1, using the GMM algorithm of the mixture Gaussian model to divide the price-quantity curve of the user into several key categories according to its key features, specifically including:

[0032] After inputting the key features into the mixture Gaussian model GMM, use the Expectation-Maximization (EM) algorithm for iterative optimization, continuously improving the parameter estimation of the model until convergence. The specific steps are as follows:

[0033] (1) E-step (Expectation Step): Given the current model parameters θ = {π k , μ k , Σ k}, using the feature matrix W = [w1, …, w d , …, w D extracted from the price - volume curves of power users, where w d represents the vector composed of the d - th eigenvalue of all data points, as the model input, calculate the responsibility value of each data point w d belonging to the k - th Gaussian distribution:

[0034]

[0035] where γ dk represents the probability that the data point w d belongs to the k - th Gaussian distribution;

[0036] (2) M-step (Maximization Step): According to the calculated responsibility γ ik , update the GMM model parameters π k , μ k , Σ k :

[0037] Mean update:

[0038]

[0039] Covariance update:

[0040]

[0041] Weight update:

[0042]

[0043] where N is the total number of data points; γ dk is the probability that the d - th data point belongs to the k - th distribution;

[0044] During the process of the EM algorithm, the Bayesian Information Criterion BIC is used to select the optimal value of the latent clusters. The calculation formula of the Bayesian Information Criterion is as follows:

[0045] BIC = -2·log(L) + q·log(N) (11)

[0046] where L represents the maximum likelihood estimate value obtained by model fitting, q is the number of model parameters, and N represents the number of data points;

[0047] Repeat steps E and M until the model parameters converge or meet the preset stopping criteria. Finally, all curves are divided into several clusters according to the Bayesian information criterion, and each cluster is denoted as n1, n2, …; if a certain curve R i is assigned to the first cluster, it is denoted as If it is assigned to the second cluster, it is denoted as And so on, all price - volume curves are assigned to the most suitable category according to their characteristics.

[0048] Furthermore, in step S2, for the price - volume curve of the first electricity user, Gaussian process regression is used for regression modeling to obtain the aggregated regression curve of this user, specifically as follows:

[0049] The input point is the declared price x1 of the first electricity user, which is associated with an output value of the declared demand y1 of the first electricity user. The relationship between the two is as follows:

[0050] y1 = f(x1)+ε (12)

[0051] where f(x1) represents the function value in Gaussian process regression, and ε is Gaussian noise;

[0052] The kernel function of the Gaussian process regression model is expressed as:

[0053] f(x)~gp(μ(x),k(x,x')) (13)

[0054] where k(x,x') represents the kernel function, which is used to measure the similarity between any two data points; μ(x) is the mean function.

[0055] Furthermore, the kernel function is the Matérn kernel, and its general form is expressed as:

[0056]

[0057] where v is the smoothness parameter; l is the length - scale parameter; ||x - x'|| is the Euclidean distance between input points; Γ(v) is the gamma function; K v is the modified Bessel function of order v; during the operation of the Matérn kernel, the length - scale parameter l is used to control the scope of action of the kernel function, the gamma function Γ(v) is used to perform a normalization process on the kernel function, and the modified Bessel function K v is used to simplify the function itself; flexible control of smoothness is achieved by adjusting the v value.

[0058] Furthermore, when v = 3 / 2, the Matérn kernel has a simple closed - form expression:

[0059]

[0060] At this time, the mean square error of the Matern kernel is the lowest, and the fitting effect is the best.

[0061] Furthermore, step S3 specifically includes two steps: linear interpolation and piecewise linear curve generation;

[0062] (1) Linear interpolation: Consider the smooth curve generated by GPR as a continuous function f(x). The points x on the curve obtain smooth predicted values f(x i ); First, divide the regression curve by the linear interpolation method and perform interpolation on a new uniformly distributed point set E = {e1, e2,..., e n}; For every two adjacent points e i and e i+1 , the interpolation formula is as follows:

[0063]

[0064] where f(e) is the estimated value at the new interpolation point x;

[0065] (2) Generate a piecewise linear curve: First, determine b piecewise points according to the form of equally spaced segmentation, and then in each interval [e i-1 , e i , use the following method to determine the new y i , define the original regression curve as f 回归 , and obtain the following mathematical expressions:

[0066] 1) Take the regression curve value corresponding to the midpoint of the interval:

[0067] y i = f 回归 ((e i-1 + e i ) / 2) (17)

[0068] 2) Mean value of the interval:

[0069]

[0070] Then, substitute the previously determined piecewise points and constant values into the two expressions to obtain the piecewise step curve pattern.

[0071] The beneficial effects of the present invention are as follows: the present method can reduce the dimension and complexity of user data, so that the trading center can only process the simplified alternative curve after aggregation, thereby improving the efficiency of data analysis and clearing decision-making. Compared with the existing methods, the alternative price-volume curve obtained by the method proposed in the present invention has a higher similarity with the original curve declared by the power user, and thus can more truly reflect the price response characteristics of users participating in the market. Compared with traditional methods, the present method can effectively reduce the dimension and complexity of data, and the trading center only needs to process the alternative curve after aggregation and simplification, thereby improving the declaration efficiency. The alternative price-volume curve generated by this method has a high similarity with the original curve, and more truly reflects the price response characteristics of users in the market.

[0072] (1) Able to better capture the complexity and nonlinear relationships of data

[0073] The Gaussian mixture model combines the aggregation method of Gaussian process regression. Through the non-parametric modeling capability of Gaussian process regression and the probability distribution modeling of the Gaussian mixture model, it can better capture these complex nonlinear relationships and improve the aggregation effect.

[0074] (2) Better adaptability and flexibility

[0075] The aggregation method that combines the mixed Gaussian model and Gaussian process regression can automatically discover potential aggregation structures in different types of quotation curves without presetting the number of clusters or relying on a fixed distance metric, thereby improving the adaptability and flexibility of the aggregation method.

[0076] (3) Generate a curve form that is more in line with the market

[0077] The price-volume curve in the electricity market often presents a piecewise step-type structure. By processing the regression curve into piecewise steps, the curve presents a piecewise step form that is more in line with market conditions and user needs.

[0078] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:

[0080] Figure 1 Price-volume curve for large users;

[0081] Figure 2It is the GMM-GPR aggregation curve (squared exponential kernel);

[0082] Figure 3 It is the GMM-GPR aggregation curve (Matern kernel);

[0083] Figure 4 It is the aggregated curve graph after piecewise processing (squared exponential kernel);

[0084] Figure 5 It is the aggregated curve graph after piecewise processing (Matern kernel). Specific implementation manners

[0085] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0086] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in its actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0087] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.

[0088] The technical means and implementation process of the present invention include: classifying the original price-volume curve using a Gaussian mixture model, aggregating each original curve under each classification using a Gaussian process regression algorithm, performing piecewise stepwise linearization processing on the obtained aggregated regression curve, and proposing an index system to evaluate the aggregation effect of the proposed algorithm. The relevant technical means and implementation process are specifically described as follows. The present invention uses the Monte Carlo method to simulate and generate the price-volume curve reported by each user as the data basis to verify the effectiveness of the method proposed in the present invention.

[0089] Step 1: Classification by Gaussian mixture model (GMM)

[0090] The Gaussian Mixture Model (GMM) assumes that the data is generated by the weighted sum of multiple Gaussian distributions. Here, we consider using the Gaussian Mixture Model to perform preliminary classification on the declared data of the price-volume curve because it can be well applied to aggregation and classification tasks, especially when there are multiple potential clusters or groups in the data. GMM can capture these potential patterns more effectively. The mathematical representation of the Gaussian Mixture Model is as follows:

[0091]

[0092] Where: p(x) is the probability density of data point x; K is the number of components in the mixed distribution (i.e. the number of Gaussian distributions); π k is the weight of the kth Gaussian distribution, satisfying Represents the contribution of each distribution to the overall model; N(x丨μ k ,∑ k ) is the probability density function of the kth Gaussian distribution with mean μ k and the covariance matrix ∑ k .

[0093] Before using GMM to perform a preliminary classification of the original electricity user price-volume curve, it is necessary to first prepare the data and build a feature matrix for the input feature quantity, and then use the GMM model to calculate the probability density of each data point for the input data x, and classify each data point into its appropriate category. In the analysis of the relationship between electricity demand and electricity prices, five key features were selected as input parameters for GMM classification. The selection of each feature is intended to more comprehensively reflect the demand patterns of different users in the electricity market. By selecting these input features, it is ensured that the model has a strong ability to distinguish when the original data is input. Specifically, these five features include: slope, intercept, quantity-price correlation coefficient, quantity-price slope change rate, and regression residual variance. The meaning of each feature and the reasons for selection are as follows:

[0094] (1) Slope:

[0095]

[0096] Among them, x i represents the electricity price, y i represents the power demand, represents the mean value of electricity price, represents the mean of electricity demand, and n is the number of data points.

[0097] The slope reflects the degree of inclination of the price-volume curve and can intuitively measure the impact of electricity price changes on electricity demand. The solution process is mainly to calculate the sum of the products of the electricity price and the demand deviation: And the sum of the squares of the electricity price deviation Dividing the above two equations gives the slope characteristic of the price - quantity curve. A larger slope indicates that the change in electricity price has a greater impact on the demand quantity, while a smaller slope indicates a smaller impact of electricity price on the demand quantity. Therefore, the slope helps to distinguish users with higher or lower sensitivity to electricity price in the market.

[0098] (2) Intercept:

[0099]

[0100] The intercept represents the intersection point of the price - quantity curve on the price axis and reflects the basic electricity demand when the electricity price is zero. The intercept characteristic of the electricity user's price - quantity curve is mainly obtained by subtracting the product of the mean of the electricity demand quantity and the means of the slope and the electricity price after obtaining the slope characteristic. A high intercept value means a higher basic electricity demand of the user. When analyzing the market demand, the intercept helps to understand the starting demand level of users, that is, the basic consumption of electricity in the absence of electricity price fluctuations.

[0101] (3) Price - Quantity Correlation:

[0102]

[0103] Among them, C P,Q represents the price - quantity correlation coefficient, and its value range is [-1, 1].

[0104] The price - quantity correlation coefficient measures the correlation between the electricity price and the electricity demand, and the value range is from -1 to 1. A positive value indicates a positive correlation between the electricity price and the demand quantity, while a negative value indicates a negative correlation. This characteristic can reveal the dependence relationship between the electricity price and the demand quantity, and further helps to identify the impact pattern of price changes on the demand quantity in the market. Through the price - quantity correlation coefficient, the impact degree of price fluctuations on the demand can be effectively captured.

[0105] (4) Rate of Change in Price - Quantity Slope:

[0106]

[0107] Among them, Slope i is the local slope of each point, is the average slope of all points. The rate of change of the price - volume slope refers to the rate of change of the slope of the price - volume curve, that is, whether the response of demand to price changes is stable. Calculating the rate of change of the price - volume slope feature of the electricity user's price - volume curve means calculating the second - order difference between the electricity price and the demand. A higher rate of slope change indicates that the reaction of demand to price changes fluctuates greatly, while a lower rate of change indicates that the demand response is relatively stable. Therefore, this feature can reveal the dynamic impact of market price fluctuations on demand.

[0108] (5) Residual Variance:

[0109]

[0110] where y i represents the actual electricity demand, represents the predicted value of the Gaussian process regression model. Here, the calculation of the predicted value needs to be iteratively optimized in combination with the Gaussian process regression in the second step. The residual variance of the electricity quantity optimized by the Gaussian process regression is used as a feature input for a mixture Gaussian model. The specific iterative process is as follows: First, select a kernel function using the existing data and train the Gaussian process regression model. Then, subtract the predicted value obtained from the actual value to get the residual. Next, accumulate the squared residuals and take the average to get the residual variance. If the obtained residual variance value is large, enter the model optimization stage. Check the trend of the residual variance by optimizing the Gaussian process regression model (optimizing hyperparameters, changing the kernel function, etc.) until it is less than the preset threshold. The value obtained at this time is used as the feature input for the mixture Gaussian model.

[0111] The residual variance of the regression reflects the goodness of fit of the Gaussian process regression model. A smaller residual variance means that the model can fit the data well. The solution process is mainly to find the difference between the actual electricity demand and the predicted electricity demand: Get the residual value, then square the residual, accumulate and take the average to get the residual variance eigenvalue. By evaluating the residual variance of the regression, it can be judged whether the model can effectively describe the change characteristics of the price - volume curve. A larger residual variance of the regression may indicate that there are large data deviations in some areas, suggesting that a more complex model is needed to fit the data.

[0112] The main reasons for selecting these features are as follows: The slope and intercept directly reflect the relationship between electricity price and electricity demand; the quantity-price correlation coefficient helps capture the dependence between electricity price and demand; the rate of change of the quantity-price slope reveals the dynamic stability of demand response; and the regression residual variance can measure the fitting effect of the linear regression model. Considering these five features comprehensively takes into account the sensitivity, stability, and fitting quality of demand response, which helps to comprehensively evaluate and classify the electricity demand patterns in the market. Compared with other high-order features, these features are more interpretable and closer to the actual situation of the electricity market.

[0113] The following are the steps for curve classification. Since the GMM model is a mixture model composed of multiple Gaussian distributions, the estimation of its parameters usually uses the Expectation-Maximization (EM) algorithm. The EM algorithm continuously improves the parameter estimation of the model through an iterative optimization process until convergence. The basic steps are as follows:

[0114] (1) E-step (Expectation Step): Given the current model parameters θ = {π k , μ k , Σ k}, using the feature matrix W = [w1, …, w d , …, w D extracted from the price-quantity curves of electricity users, where w d represents the vector composed of the d-th feature values of all data points, as the model input, calculate the responsibility value of each data point w d belonging to the k-th Gaussian distribution:

[0115]

[0116] where γ dk represents the probability that the data point w d belongs to the k-th Gaussian distribution. Here, "responsibility" mainly refers to the probability that the input features of each price-quantity curve belong to a specific Gaussian distribution, which specifically reflects the degree of belonging of this data point under different clusters.

[0117] (2) M-step (Maximization Step): According to the calculated responsibility γ ik , update the GMM model parameters π k , μ k , Σ k :

[0118] 1) Mean update:

[0119]

[0120] 2) Covariance update:

[0121]

[0122] 3) Weight update:

[0123]

[0124] where N is the total number of data points; γ dk is the probability that the d-th data point belongs to the k-th distribution. During the EM algorithm process, to achieve accurate classification, a model selection criterion (Bayesian Information Criterion) is used to select the optimal value of the potential clusters. The calculation formula of the Bayesian Information Criterion (BIC) is as follows:

[0125] BIC = -2·log(L) + q·log(N) (11)

[0126] where L represents the maximum likelihood estimate value obtained from model fitting, q is the number of model parameters, and N represents the number of data points.

[0127] The BIC is used to select the optimal number of clusters to ensure that the classification result is both reasonable and not overfitted. The core idea of this strategy is to ensure that each user is assigned to the cluster that best matches their demand pattern, avoiding overfitting of the classification result of the electricity user price - quantity curve. Adopting this classification strategy can better capture the demand characteristics of different users in the market and achieve more refined market stratification.

[0128] Repeat the E-step and M-step until the model parameters converge or meet the preset stopping criterion.

[0129] In the classification result of this study, assume that there are R price - quantity curves to be classified, denoted as R = (R1, R2, R3, …, R n ), through the Bayesian Information Criterion, these curves are divided into several clusters, denoted as n1, n2, … respectively. If a certain curve R i is assigned to the first cluster, it is denoted as If it is assigned to the second cluster, it is denoted as And so on, all price - quantity curves are assigned to the most suitable category according to their characteristics.

[0130] Step 2: Gaussian Process Regression (GPR) Aggregation

[0131] Regarding the classification result obtained by the Gaussian mixture model in Step 1, it is obtained that the first price - quantity curve is classified into the first cluster, denoted as Taking the price - quantity curve of the first electricity user as an example, Gaussian process regression is used to perform regression modeling on the price - quantity curve of this user to obtain the aggregated regression curve of this user. The input point x1 (the declared price of the first electricity user) is associated with an output value y1 (the declared demand of the first electricity user). For an input data x1, its output value y1 is given by the following formula:

[0132] y1 = f(x1)+ε (12)

[0133] where f(x1) represents the function value in Gaussian process regression, which is usually regarded as a multi - dimensional Gaussian distribution, and ε is Gaussian noise.

[0134] In the Gaussian process regression (GPR) model, the choice of the kernel function is crucial because it directly affects the similarity measure between different data points, thus determining the regression effect. Its mathematical expression is as follows:

[0135] f(x)~gp(μ(x),k(x,x')) (13)

[0136] where k(x,x') represents the kernel function, which is mainly used to measure the similarity between any two data points; μ(x) is the mean function, which is usually taken as a zero - mean. To effectively model the relationship between electricity demand and electricity price, we conducted extensive experimental comparisons on different types of kernel functions, including the common squared - exponential kernel, linear kernel, and Matern kernel. Through comparative analysis, it is found that the Matern kernel, especially under the configuration of its smoothness parameter v = 3 / 2, can better perform regression aggregation on the price - quantity curve. The general form of the Matern kernel is expressed as:

[0137]

[0138] where v is the smoothness parameter; l is the length - scale parameter; ||x - x'|| is the Euclidean distance between input points; Γ(v) is the gamma function; K v is the modified Bessel function of the second kind of order v. During the operation of the Matern kernel, the length - scale parameter l is used to control the range of action of the kernel function, the gamma function Γ(v) is used to perform a normalization process on the kernel function, and the modified Bessel function K v is used to simplify the function itself. When v = 3 / 2, the Matern kernel has a simple closed - form expression:

[0139]

[0140] When we select the Matérn kernel, we can flexibly control the smoothness by adjusting its v value. At the same time, compared with general kernel functions such as the squared exponential kernel, the Matérn kernel is more robust against high-noise data. The fitting degree of the regression model is evaluated by the mean squared error (MSE) as an evaluation index. Through experimental comparative analysis, it is found that the mean squared error of the Matérn kernel (v = 3 / 2) is the lowest and the fitting effect is the best. The accuracy and stability of the regression results are ensured by the selection of the kernel function, and significant advantages are shown in data fitting and prediction accuracy.

[0141] Step 3: Piecewise linear processing

[0142] After obtaining the smooth regression curve through Gaussian process regression (GPR), the further goal is to convert these smooth curves into a piecewise step form that meets the actual application requirements of the electricity market. This process involves two key steps: linear interpolation and piecewise linear curve generation, aiming to convert the smooth regression curve into the price-volume curve representation form commonly seen in the current electricity market.

[0143] (1) Linear interpolation: Assume that the smooth curve generated by GPR is regarded as a continuous function f(x), and the points x on the curve obtain the smooth predicted value f(x i ). To meet the market demand and convert the smooth curve into a piecewise step form, we first divide the regression curve by the linear interpolation method, hoping to perform interpolation on a new uniformly distributed point set E = {e1, e2,..., e n}. For every two adjacent points e i and e i+1 , the interpolation formula is as follows:

[0144]

[0145] where f(e) is the estimated value at the new interpolation point x.

[0146] (2) Generate piecewise linear curve: Once the new data points are obtained through interpolation, we then connect each segment of data into a piecewise linear curve.

[0147] First, b piecewise points need to be determined according to the equally spaced segmentation form, and then in each interval [e i-1 , e i , the following method can be used to determine the new y i , define the original regression curve as f 回归 , and the following mathematical expression can be obtained:

[0148] 1) Take the regression curve value corresponding to the midpoint of the interval:

[0149] y i = f回归 ((e i-1 +e i ) / 2) (17)

[0150] 2) Interval mean:

[0151]

[0152] Then, by substituting the previously determined segmentation points and constant values into the above two expressions, the segmented step curve pattern can be obtained.

[0153] In this way, we can transform the smooth curve obtained by Gaussian process regression into a segmented step form, which is widely used in the electricity market because it intuitively presents the response of price changes to demand. Specifically, the linear fitting of each segment represents the change trend of electricity demand within a specific price range, ensuring that the clustered results are consistent with the trend characteristics of the original data. During the process of generating the segmented linear curve, special attention is paid to maintaining the consistency of the aggregated results with the original curve in terms of trend characteristics. Through the designed interpolation and fitting steps, the segmented linear curve can not only present a clear price-demand relationship but also effectively retain the volatility of demand response within each aggregation. Finally, the generated segmented step curve conforms to the common manifestation form in the electricity market, ensuring the feasibility and practicality of the model output in actual operation.

[0154] Step 4: Aggregation effect evaluation

[0155] To evaluate the aggregation effect of the above method compared with traditional algorithms, internal evaluation indicators (DTW + MSE) are used for analysis.

[0156] 1. Dynamic time warping:

[0157] Dynamic time warping (DTW) is used to measure the similarity between the original curve and the aggregated price-volume curve. The smaller the DTW distance, the higher the curve similarity. Its calculation formula is as follows:

[0158] (1) Construct an n×m cumulative distance matrix D, where D(i,j) represents the cumulative minimum distance from X(i) to Y(j).

[0159] (2) Initialize matrix D:

[0160] D(1,1) = d(x1,y1) (19)

[0161] (3) Fill the matrix by dynamic programming:

[0162] D(i,j) = d(x i ,y j) + min{D(i - 1, j), D(i, j - 1), D(i - 1, j - 1)}(20)

[0163] (4) The final DTW distance is:

[0164] DTW(X, Y) = D(n, m) (21)

[0165] 2. Mean squared error:

[0166] The mean squared error (MSE) is used to quantify the average error between the original curve and the aggregated curve. The smaller the MSE, the smaller the curve error and the higher the similarity. Its calculation formula is as follows:

[0167] Suppose there are two time series:

[0168] X = (x1, x2, …, x n ), Y = (y1, y2, …, y n )(22)

[0169] The calculation formula of MSE is:

[0170]

[0171] where x i and y i are the corresponding points in the time series X and Y.

[0172] Combining DTW and MSE as a comprehensive similarity index helps to more comprehensively evaluate the aggregation effect. By assigning weights to this comprehensive index, the importance of various factors can be weighed during the evaluation process. The smaller the comprehensive score, the better the aggregation effect. According to the index evaluation results in Example 4, the aggregation method combining GMM and GPR has a better simplification effect and higher similarity for the price - volume curves of power users compared to traditional aggregation methods.

[0173] Example 1:

[0174] As Figure 1 shown, in this example, several piece - wise step - shaped price - volume curves are generated based on the original data for subsequent aggregation analysis. The generated curves retain the typical price - volume relationship in the electricity market by simulating the bidding behavior of market participants. These curves reflect the dynamic changes in supply - demand balance in the market and can effectively simulate the price fluctuation characteristics under different market conditions, providing data support for subsequent market behavior analysis and optimization.

[0175] Example 2:

[0176] As Figures 2-3As shown, this example analyzes the declared data of the price - volume curves of 10 electricity users. First, a mixture Gaussian model is used to perform aggregation analysis on the price - volume data of users. According to the posterior probability of each user's price - volume curve under each Gaussian distribution, it is classified into the most likely category, and finally the data is divided into three categories. Then, Gaussian process regression (GPR) is used to aggregate the curves of each category to generate three smooth regression curves. This process not only optimizes the aggregation effect of the original data but also provides a deeper understanding of the behavior of electricity market participants, especially the potential changes in user behavior under different market conditions. Through this method, trends in the market can be captured more accurately, and strong data support can be provided for price prediction and optimal scheduling in the electricity market.

[0177] Example 3:

[0178] As Figures 4-5 shown, this example extracts the electricity market data of 10 users for case analysis. For the aggregated curves obtained through Gaussian process regression (GPR), a piecewise step - like processing is further applied. To make the aggregated curves more accurately reflect the changing trend of the original price - volume curves, this example uses the equal - interval quotation method to segment the three aggregated curves. This processing method effectively improves the practical applicability of the aggregation results, ensuring that the segmented curves can better reflect the price - quantity relationship in the market while maintaining the original trend. In addition, the piecewise step - like processing optimizes the smoothness and response characteristics of the curves, facilitating subsequent in - depth analysis of market behavior and decision - making support.

[0179] Example 4:

[0180] To comprehensively evaluate the adopted aggregation algorithms, this example uses dynamic time warping (DTW) and mean squared error (MSE) as evaluation indicators and conducts a comparative analysis of four aggregation algorithms. The selected aggregation algorithms include: mixture Gaussian model (GMM) aggregation, K - means aggregation, mixture Gaussian model combined with Gaussian process regression (GPR) aggregation (using the Matérn kernel function), and mixture Gaussian model combined with Gaussian process regression aggregation (using the squared exponential kernel function). Under the conditions of the same aggregation quantity and unified feature matrix, we compare the comprehensive evaluation indicators of each algorithm. The evaluation results show that the smaller the comprehensive index score, the better the aggregation result fits the original price - volume curve. Through the comparative analysis results, the aggregation method of the mixture Gaussian model combined with Gaussian process regression performs the best among various aggregation algorithms. Especially when choosing the kernel function of Gaussian process regression, the aggregation effect of using the Matérn kernel function is significantly better than other kernel functions. Table 1 shows the results of using the DTW + MSE index scores (weight ratio 7:3).

[0181] Table 1

[0182]

[0183] In the above embodiments, the reference in the specification to "this embodiment" means that the specific features, structures, or characteristics described in connection with the embodiment are included in at least some embodiments, but not necessarily all embodiments. Multiple occurrences of "this embodiment" do not necessarily all refer to the same embodiment.

[0184] In the above embodiments, although the present invention has been described in connection with specific embodiments of the present invention, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other storage structures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed. The embodiments of the present invention are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims.

[0185] This embodiment also provides a computer-readable storage medium having a computer program stored thereon, and when the program is executed by a processor, it implements any one of the methods in this embodiment.

[0186] This embodiment also provides an electronic terminal, including: a processor and a memory;

[0187] The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory so that the terminal executes any one of the methods in this embodiment.

[0188] For the computer-readable storage medium in this embodiment, those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to the computer program. The foregoing computer program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disk, or optical disk that can store program codes.

[0189] The electronic terminal provided in this embodiment includes a processor, a memory, a transceiver, and a communication interface. The memory and the communication interface are connected to the processor and the transceiver and complete communication therebetween. The memory is used to store a computer program, the communication interface is used for communication, and the processor and the transceiver are used to run the computer program so that the electronic terminal executes each step of the method as described above.

[0190] In this embodiment, the memory may include a random access memory (Random Access Memory, abbreviated as RAM), and may also include a non-volatile memory, such as at least one disk memory.

[0191] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it can also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0192] The present invention can be used in numerous general-purpose or special-purpose computing system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on.

[0193] The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0194] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A method for aggregating electricity user price - volume curves based on a mixture Gaussian model and Gaussian process regression, characterized in that: It includes the following steps: S1: Using the Gaussian mixture model GMM algorithm, divide the price - volume curve of the user into several key categories according to its key features; S2: Using the Gaussian process regression GPR algorithm to perform regression aggregation on the curves of each category; S3: Perform piece - wise step - linearization processing on the obtained regression - aggregated curve to obtain a piece - wise step curve that conforms to the form of the electricity market performance.

2. The method for aggregating electricity user price - volume curves based on the mixture Gaussian model and Gaussian process regression according to claim 1, wherein: The price - volume curve of the user is generated by Monte Carlo sampling method.

3. The method for aggregating price - volume curves of power users based on the Gaussian mixture model and Gaussian process regression according to claim 1, wherein: The mathematical representation of the Gaussian mixture model GMM is as follows: Where p(x) is the probability density of data point x; K is the number of components in the mixed distribution, that is, the number of Gaussian distributions; π k is the weight of the kth Gaussian distribution, satisfying Represents the contribution of each distribution to the overall model; N(x丨μ k ,∑ k ) is the probability density function of the kth Gaussian distribution with mean μ k and the covariance matrix ∑ k .

4. The method for aggregating electricity user price - volume curves based on the Gaussian mixture model and Gaussian process regression according to claim 1, wherein: The key features include slope, intercept, price - volume correlation coefficient, rate of change of price - volume slope, and regression residual variance; The slope Slope is expressed as: Among them, x i represents the electricity price, y i represents the electricity demand, represents the mean value of the electricity price, represents the mean value of the electricity demand, and n is the number of data points; The intercept Intercept is expressed as: The quantity-price correlation coefficient C P,Q is expressed as: C P,Q The value range is [-1, 1]; The rate of change of price - volume slope Rate of Change Slope is expressed as: Among them, Slope i is the local slope of each point, is the average slope of all points; The regression residual variance Residual Variance is expressed as: Among them, y i represents the actual power demand, represents the predicted value of the Gaussian process regression model.

5. The method for aggregating price - volume curves of power users based on the mixture Gaussian model and Gaussian process regression according to claim 1, wherein: In the regression residual variance, the predicted value of the Gaussian process regression model needs to be iteratively optimized in combination with Gaussian process regression. The regression residual variance of the optimized electricity quantity after Gaussian process regression is used as the feature input of a mixture Gaussian model. The specific iterative process is as follows: First, use the existing data to select a kernel function and train the Gaussian process regression model. Then, subtract the predicted value obtained from the actual value to get the residual. Then, accumulate the squared residuals and take the average to get the residual variance. If the obtained residual variance value exceeds the preset threshold, enter the model optimization stage. Check the change trend of the residual variance by optimizing the Gaussian process regression model until it is less than the preset threshold, and take the value obtained at this time as the feature input of the Gaussian mixture model.

6. The method for aggregating the price - volume curves of power users based on the Gaussian mixture model and Gaussian process regression according to claim 1, wherein: In step S1, using the Gaussian mixture model GMM algorithm to divide the price - volume curve of the user into several key categories according to its key features specifically includes: After inputting the key features into the Gaussian mixture model GMM, use the expectation - maximization EM algorithm to iteratively optimize and continuously improve the parameter estimation of the model until convergence. The specific steps are as follows: (1)Step E: Given the current model parameters θ = {π k , μ k , Σ k}, use the feature matrix W = [w1, …, w d , …, w D extracted from the price - volume curve of electricity users, where w d represents the vector composed of the d - th eigenvalue of all data points. As the model input, calculate the responsibility value of each data point w d belonging to the k - th Gaussian distribution: where γ dk represents the probability that the data point w d belongs to the k-th Gaussian distribution; (2) Step M: Update the GMM model parameters π ik , μ k , μ k , Σ k : Mean update: Covariance update: Weight update: where N is the total number of data points; γ dk is the probability that the d-th data point belongs to the k-th distribution; During the process of performing the EM algorithm, use the Bayesian information criterion BIC to select the optimal value of the latent cluster. The calculation formula of the Bayesian information criterion is as follows: BIC = - 2·log(L)+q·log(N) (11) Where, L represents the maximum likelihood estimate value obtained by model fitting, q is the number of model parameters, and N represents the number of data points; Repeat the execution of step E and step M until the model parameters converge or meet the preset stopping criterion. Finally, all curves are divided into several clusters through the Bayesian information criterion, and each cluster is denoted as n1, n2, …; If a certain curve R i is assigned to the first cluster, it is denoted as If it is assigned to the second cluster, it is denoted as r n2,i , and so on. All price-volume curves are assigned to the most suitable category according to their characteristics.

7. The method for aggregating price - volume curves of power users based on the mixture Gaussian model and Gaussian process regression according to claim 1, characterized in that: In step S2, for the price - volume curve of the first electricity user, use Gaussian process regression to perform regression modeling to obtain the aggregated regression curve of this user, specifically as follows: The input point is the declared price x1 of the first electricity user, and it is associated with an output value of the demanded quantity y1 declared by the first electricity user. The relationship between them is as follows: y1 = f(x1)+ε (12) Where, f(x1) represents the function value in Gaussian process regression, and ε is Gaussian noise; The kernel function of the Gaussian process regression model is expressed as: f(x)~gp(μ(x),k(x,x')) (13) Where, k(x,x') represents the kernel function, which is used to measure the similarity between any two data points; μ(x) is the mean function.

8. The method for aggregating electricity user price - volume curves based on the mixture Gaussian model and Gaussian process regression according to claim 1, characterized in that: The kernel function is the Matérn kernel, and its general form is expressed as: Among them, v is the smoothing parameter; l is the length scale parameter; ||x - x'|| is the Euclidean distance between input points; Γ(v) is the gamma function; K v is the modified Bessel function of order v; during the action process of the Matérn kernel, the length scale parameter l is used to control the action range of the kernel function, the gamma function Γ(v) is used to perform a normalization process on the kernel function, and the modified Bessel function K v is used to simplify the function itself; flexible control of smoothness is achieved by adjusting the value of v.

9. The method for aggregating electricity user price - volume curves based on the Gaussian mixture model and Gaussian process regression according to claim 8, wherein: When v = 3 / 2, the Matérn kernel has a simple closed - form expression: At this time, the mean squared error of the Matérn kernel is the lowest, and the fitting effect is the best.

10. The method for aggregating electricity user price - volume curves based on the mixture Gaussian model and Gaussian process regression according to claim 1, wherein: Step S3 specifically includes two steps: linear interpolation and piecewise linear curve generation; (1) Linear interpolation: Consider the smooth curve generated by GPR as a continuous function f(x), and the points x on the curve obtain smooth predicted values f(x i ); First, divide the regression curve by the linear interpolation method and perform interpolation on a new uniformly distributed point set E = {e1, e2, …, e n}; For every two adjacent points e i and e i+1 , the interpolation formula is as follows: where f(e) is the estimated value at the new interpolation point x; (2) Generate a piecewise linear curve: First, determine b segmentation points in the form of equally spaced segmentation, and then in each interval [e i-1 , e i , adopt the following method to determine the new y i . Define the original regression curve as f 回归 , and obtain the following mathematical expression: 1) Obtain the regression curve value corresponding to the midpoint of the interval: y i = f 回归 ((e i-1 + e i ) / 2) (17) 2) Interval mean: Then substitute the previously determined segmentation points and constant values into the two expressions to obtain the piecewise step curve pattern.

Citation Information

Cited By

  • Water and fertilizer integrated control strategy recommendation method and system based on multi-objective optimization

    CN122228818A