Model guiding-data driving price response characteristic identification method and device

By using analytical models to generate synthetic data and combining it with dynamic small-batch training methods, the overfitting problem of data-driven models under conditions of insufficient data is solved, and efficient identification of price response characteristics is achieved.

CN120705574APending Publication Date: 2025-09-26CHINA THREE GORGES CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510795826.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

When data is insufficient, existing data-driven models are prone to overfitting problems, resulting in a significant decrease in generalization performance and making it difficult to effectively identify price response characteristics.

Method used

By using analytical models that reflect the market operation mechanism to generate synthetic data, expanding the amount of sample data, and combining dynamic small-batch training methods, the training process of the data-driven model can be guided to avoid the risk of overfitting.

Benefits of technology

It effectively solves the problem of efficient training of data-driven models under conditions of insufficient data, ensures the orderly development of offline training of the stacked architecture, and improves the training efficiency and accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705574A_ABST
    Figure CN120705574A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses a model-guided-data-driven price response characteristic identification method and device, and the method comprises the steps: obtaining historical data of an electricity market; the prior analysis model is calibrated by using historical data; generating synthetic data by using the calibrated prior analysis model, and combining the historical data with the synthetic data to generate a mixed data set; and training the data-driven model by using the mixed data set. The invention provides a data enhancement method guided by a prior analytical model. According to the method, analytic model resources accumulated in an application scene are ingeniously utilized, and the concept of knowledge guidance and data driving is embodied. The analytic model reflecting the market operation mechanism is fully utilized to generate the synthetic data, the sample data size is expanded, then the training process of the data-driven model is efficiently guided, and the over-fitting risk is effectively avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a model-guided and data-driven price response characteristic identification method and device. Background Art

[0002] Simulating the response behavior of distributed resources under different price incentives is a key task in developing interactive systems between power sources, grids, loads, and storage. These response behavior characteristics are generally referred to as price response characteristics. In real-world applications, identifying price response characteristics may face the problem of insufficient data, meaning that the amount of training data is significantly smaller than the number of model parameters. This problem arises, on the one hand, because practical applications are often subject to external constraints such as low data sampling frequency and data privacy. On the other hand, price response characteristics are generally complex and can only be accurately captured by large-scale models. It is generally believed that insufficient data can easily lead to overfitting, resulting in a significant decrease in the generalization performance of data-driven models.

[0003] Unfortunately, most existing research avoids this issue, often abandoning high-precision data-driven models or directly using simulated data. Currently, there is an urgent need to develop efficient and feasible methods to deal with the problem of insufficient data. Summary of the Invention

[0004] In view of this, the present invention provides a model-guided and data-driven price response characteristic identification method and device to solve the problem of insufficient training data for data-driven models.

[0005] In a first aspect, the present invention provides a model-guided and data-driven price response characteristics identification method, the method comprising:

[0006] Obtain historical data on the electricity market;

[0007] calibrating the prior analytical model using the historical data;

[0008] generating synthetic data using the calibrated prior analytical model, and merging the historical data with the synthetic data to generate a hybrid data set;

[0009] The data-driven model is trained using the mixed dataset.

[0010] The model-guided and data-driven price response characteristic identification method provided in this application proposes a data enhancement method guided by a priori analytical model. This method cleverly utilizes the analytical model resources accumulated in the application scenario and embodies the concept of "knowledge-guided, data-driven". By making full use of the analytical model that reflects the market operation mechanism to generate synthetic data, the amount of sample data is expanded, and the training process of the data-driven model is efficiently guided, the risk of overfitting is effectively avoided. After the transformation, the overall process is transformed into a data-driven modeling method guided by the analytical model, which effectively solves the problem of efficient training of data-driven models under conditions of insufficient data, thereby ensuring the orderly development of the offline training link of the stacked architecture.

[0011] In an optional embodiment, the training of the data-driven model using the mixed data set includes:

[0012] Dynamically extracting small batches of data from the mixed dataset;

[0013] The data-driven model is trained using the mini-batch data.

[0014] The dynamic mini-batch training method can efficiently process mixed data sets consisting of real data and synthetic data, taking into account the effectiveness and computational efficiency of the training process.

[0015] In an optional embodiment, the training of the data-driven model using the mixed dataset further includes:

[0016] During the training process, the proportion of historical data in the small batch data is gradually increased.

[0017] In an optional embodiment, the calibrating the prior analytical model using the historical data includes:

[0018] Construct multiple prior analytical models by randomly sweeping the model parameters and randomly removing the constraints from the equations;

[0019] Multiple prior analytical models are calibrated one by one according to the historical data.

[0020] In an optional embodiment, generating synthetic data using the calibrated prior analytical model includes:

[0021] Perform performance analysis on multiple calibrated prior analytical models to obtain the fitting accuracy and estimation results of each prior analytical model;

[0022] Screening out the prior analytical model whose fitting accuracy meets the preset requirements;

[0023] For the selected priori analytical models, weighted averaging is performed on the estimation results of the priori analytical models according to the fitting accuracy relationship to generate synthetic data.

[0024] In an optional embodiment, the prior analytical model is calibrated using the following optimization model:

[0025]

[0026] Where, e k is the model output gk(λi; w k ) and the actual observed value p i Minimizing the mean square error will yield a set of optimal parameters And the corresponding minimum error value

[0027] In an optional embodiment, the prediction error, loss function and gradient noise of the data-driven model are analyzed with the help of machine learning.

[0028] In a second aspect, the present invention provides a model-guided and data-driven price response characteristic identification device, the device comprising:

[0029] Acquisition module, used to obtain historical data of the power market;

[0030] a calibration module, configured to calibrate the prior analytical model using the historical data;

[0031] a merging module, configured to generate synthetic data using the calibrated prior analytical model, and merge the historical data with the synthetic data to generate a hybrid data set;

[0032] A training module is used to train a data-driven model using the mixed data set.

[0033] The model-guided, data-driven price response characteristics identification device provided by this invention cleverly leverages the analytical model resources accumulated in application scenarios. It fully utilizes analytical models that reflect market operating mechanisms to generate synthetic data, expand the sample data volume, and effectively guide the training process of the data-driven model, effectively avoiding the risk of overfitting. Furthermore, after modification, the overall process is transformed into a data-driven modeling method guided by analytical models, effectively solving the problem of efficient data-driven model training under insufficient data conditions, thereby ensuring the orderly execution of the offline training phase of the cascade architecture.

[0034] In a third aspect, the present invention provides a computer device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to thereby execute the model-guided-data-driven price response characteristic identification method of the first aspect or any corresponding embodiment thereof.

[0035] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the model-guided-data-driven price response characteristic identification method of the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0037] Figure 1 It is a comparison of price response characteristics under different electricity price scanning scenarios;

[0038] Figure 2 A schematic flow chart of a model-guided and data-driven price response characteristics identification method according to an embodiment of the present invention;

[0039] Figure 3 is a schematic diagram comparing four data sets according to an embodiment of the present invention;

[0040] Figure 4 2 is a schematic diagram of the back propagation process of the error in training according to an embodiment of the present invention;

[0041] Figure 5 is a schematic diagram of prediction errors corresponding to different function manifolds according to an embodiment of the present invention;

[0042] Figure 6 is a schematic diagram of calibration accuracy of a priori analytical model according to an embodiment of the present invention;

[0043] Figure 7 2 is a schematic diagram of the change in training error before and after data enhancement according to an embodiment of the present invention;

[0044] Figure 8 is a structural block diagram of a model-guided-data-driven price response characteristic identification device according to an embodiment of the present invention;

[0045] Figure 9Schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0046] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present invention.

[0047] Currently, the academic community has not yet reached a unified definition and modeling standard for price response characteristics. Therefore, it is necessary to first clarify the definition method used in this application. Specifically, this application focuses on the price response characteristics of the day-ahead market, which can be expressed using the following mapping relationship:

[0048]

[0049] Where f is a high-dimensional mapping function that describes the price response characteristics. is the feasible domain. T is the number of time periods in a day, and the default value is T = 24. λ is the electricity price vector for a day, which can be expanded to be expressed as λ = [λ1,…,λ T ]. p is the user's daily electricity consumption vector, which can also be expanded into p=[p1,…,p T ].

[0050] In fact, the price response characteristics have many similar concepts: price-sensitive demand response, price-driven demand response, time-series demand response, and price-responsive load. At the same time, there are three typical variants: considering the relationship between price change Δλ and electricity consumption change Δp, focusing only on the response characteristics during daytime peak hours, and considering random factors in modeling. It is not difficult to verify that the mathematical essence of the above concepts and variants is exactly the same as formula (1-1). The price response characteristics have the outstanding characteristics of high nonlinearity, time-series coupling, limited rationality, and uncertainty. Here, taking a certain user system as an example, the two characteristics of nonlinearity and time-series coupling are visualized, as shown below. Figure 1 shown.

[0051] Figure 1 The price response characteristics and self-elasticity at 11:00 am are shown for two electricity price scanning scenarios. Figure 1 In (a), only the electricity price at 11 o'clock is scanned, while Figure 1Figure (b) simultaneously scans electricity prices from 7:00 to 15:00. Observing the individual sub-graphs, we find that the price response curves in both cases are not straight lines. In particular, a significant saturation effect is observed when the electricity price is too high or too low. This demonstrates the nonlinearity of the price response, supported by the continuously changing self-elasticity values. Furthermore, comparing the two sub-graphs reveals the impact of price changes in adjacent time periods, demonstrating the temporal coupling of the price response characteristics.

[0052] The problem of identifying price response characteristics specifically involves using existing data resources to model and analyze price response characteristics, while striving for the highest possible fitting accuracy. The following describes this problem in detail from both the data and identification model perspectives. Regarding data, the existing data resource refers to a historical dataset HD, expressed as follows:

[0053] HD={(λ1,p1),(λ2,p2),…,(λ n ,p n )} (1-2)

[0054] Here, n represents the amount of historical data, which may not be large in practical applications. Also, note that the bold symbol p1 refers to the first power consumption vector data, which contains the power consumption measurements for T time periods. This is different from the first time period power consumption p1 mentioned earlier.

[0055] In terms of the identification model, the data-driven and analytical identification function expressions are given below:

[0056]

[0057] Where p^ represents the estimated power consumption. f(·;θ) represents the data-driven function, which can be a neural network, and θ is the weight parameter to be trained. In addition, consider K different forms of analytical functions, where the kth one is g k (·;w k ), the function can adopt any of the optimization model class, elasticity class, bidding class, and special function class, w k is the parameter to be estimated in the model.

[0058] In most extrinsic feature estimation scenarios, Equation (1-3) performs significantly better than Equation (1-4). This is because, unlike analytical models, data-driven models do not rely on any simplifying assumptions. However, in special cases where data is insufficient, the performance of data-driven models will be affected.

[0059] In recent years, data-driven modeling methods have achieved tremendous success, thanks to the ever-increasing size of datasets and computing resources. However, many practical application scenarios often struggle to obtain large amounts of data. For example, user behavior analysis in the electricity market involves a large amount of user privacy data, which is often difficult to obtain or even prohibited from being collected. Even if data can be collected through bilateral agreements, the time and labor costs are extremely high. Therefore, insufficient data is a common problem. The following lists four reasons for insufficient data, taking into account the characteristics of the electricity market:

[0060] First, much of the data in the electricity market is sampled infrequently. Aside from the real-time market, the frequency of transactions and information releases in the day-ahead and medium- to long-term markets is also low. For example, the day-ahead market releases data once a day, accumulating only about 360 sets of data annually. This volume of data is incomparable to the scale of databases used in image and speech recognition.

[0061] Second, much market data involves user privacy, particularly transaction-related data. While this sensitive data can fully reflect user characteristics and decision-making preferences, it is often inaccessible due to privacy restrictions. Even in some electricity markets with relatively high data transparency, data desensitization measures such as delayed release, anonymization, and differential privacy are often employed. These additional processing steps also complicate subsequent data analysis.

[0062] Third, market conditions and user preferences are in a dynamic and constantly changing process. Outdated data is unrepresentative, and generally only recent data can provide meaningful information. For example, consider using a simple fully connected neural network with a "24-2424" structure to fit the day-ahead price response characteristics. The model contains 1,152 parameters. Based on a 70%-30% training and validation set split, a minimum of 1,646 training data points is required, equivalent to 4.5 years of data. Unfortunately, most user characteristics change within 2-3 years.

[0063] Fourth, the coexistence of complex market structures, market rules and user behaviors, and the interweaving of multiple complexities and uncertainties, greatly increase the difficulty of related analysis work. It is necessary to rely on a large amount of high-precision data to fully characterize these complex characteristics.

[0064] As can be seen from the above, there is a sharp contradiction between the limited amount of available data and the massive data requirements for high-precision modeling. Insufficient data directly affects sample representativeness, causing a certain deviation between the empirical distribution described by the sample and the true distribution. This deviation is primarily caused by random factors. It is generally considered undesirable to use a small dataset to train a large-scale data-driven model. This can lead to the risk of overfitting, and the reliability of model training cannot be guaranteed. The risk of overfitting refers to the situation where the data-driven model is misled by these random factors, overfitting the empirical distribution and performing poorly on the true distribution.

[0065] The problem of overfitting can be further understood from the perspective of error decomposition. Generally, error can be decomposed into empirical error (bias) and generalization error (variance). If these two errors are understood using training and test sets, the empirical error reflects the fit performance on the training set, while the generalization error reflects the fit performance on the test set. For example, in machine learning, the training goal is to minimize the empirical error, but the goal is to obtain a model that minimizes the generalization error. Overfitting occurs when the empirical error is small but the generalization error is large. This risk is significantly increased when data is insufficient.

[0066] When identifying user price response characteristics, retailers often possess some user data and prior knowledge. This "prior knowledge" refers to the industry experience and domain knowledge accumulated through extensive business processes, not complete ignorance of the users they serve. For example, retailers likely know the user's load type or approximate boundary parameters. Unfortunately, using this knowledge alone is not enough to build a high-precision user model. Similarly, using only existing data is insufficient to achieve high accuracy.

[0067] To this end, this application proposes a model-guided and data-driven price response characteristic identification method, which uses an analytical model that reflects the market operation mechanism to generate synthetic data, expand the amount of sample data, and then efficiently guide the training process of the data-driven model to effectively avoid the risk of overfitting.

[0068] According to an embodiment of the present invention, an embodiment of a model-guided and data-driven price response characteristic identification method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0069] In this embodiment, a model-guided and data-driven price response characteristic identification method is provided. Figure 2 is a flow chart of model-guided-data-driven price response characteristic identification according to an embodiment of the present invention. Figure 2As shown, the process includes the following steps:

[0070] Step S11: Acquire historical data of the electricity market.

[0071] Specifically, the transaction information and release information of the recent real-time market, the day-ahead market and the medium- to long-term market are collected.

[0072] Step S12: Calibrate the prior analytical model using historical data.

[0073] Specifically, the above step S12 includes the following process:

[0074] Step S121 : construct multiple prior analytical models by randomly scanning the data-driven model parameters and randomly removing the constrained equations.

[0075] In this embodiment of the present invention, the a priori analytical models fall into four categories: time-shifted load models, temperature-controlled load models, and electric vehicle load models. For each model category, 12 a priori analytical models were constructed by randomly sweeping model parameters and randomly removing constraints from the equations. Four of these models were constructed for each type.

[0076] Step S122: calibrate the multiple prior analytical models one by one according to the historical data.

[0077] In this embodiment of the present invention, the prior analytical models are calibrated one by one based on historical data. These calibrated models are subsequently used to generate synthetic data. For the kth prior analytical model, its calibration problem can be expressed as the following optimization model:

[0078]

[0079] Where, e k is the model output g k (λ i ;w k ) and the actual observed value p i Minimizing the mean square error will result in a set of optimal parameters And the corresponding minimum error value Please note that It can be used to measure the fitting effect of the prior analytical model, and models with poor fitting effect should be screened out.

[0080] The following focuses on the solution of the optimization problem (1-5). k (·;w k) is a function with an explicit expression, such as a multivariate nonlinear function, a series of nonlinear optimization methods can be used to solve the optimization problem (1-5). The BFGS algorithm (Broyden–Fletcher–Goldfarb–Shanno algorithm) is a commonly used algorithm for solving such nonlinear optimization problems.

[0081] When gk(·;w k ) is given in the form of an implicit function, the solution method of the optimization problem (1-5) is different. An implicit function generally refers to a relationship between input and output variables that is not expressed by an explicit function, but is hidden in a set of equations or an optimization problem. Taking the optimization problem as an example, if the boundary conditions are regarded as input and the optimal solution is regarded as output, then a mapping relationship has been established between the two, because whenever a set of boundary conditions is given, the corresponding optimal solution result can be obtained by solving the optimization problem (if there are multiple solutions, additional restriction rules can be introduced to ensure that the optimization algorithm selects one of the solutions according to a fixed rule). Except for very few special cases, the above mapping relationship cannot be directly expressed by explicit analysis, so it is called an implicit function relationship. Since gk(·;w k ) has no specific expression, so appropriate transformation methods must be used to embed this implicit function relationship into the optimization problem (1-5). Without loss of generality, we can first transform the different implicit function sources into a unified system of equations. For the optimization problem, this can be equivalently transformed into its optimality conditions, that is, a set of nonlinear equations. Then, all these equations are embedded as constraints into the optimization problem (1-5), resulting in the following optimization problem:

[0082]

[0083] Where I is the constraint number of the optimization problem, h i is the constraint set, It is an auxiliary modeling variable used to construct the optimization problem expression. i (·) is the embedded constraint set. For the sake of simplicity, we use the less than or equal to form. In fact, this expression can also express the greater than or equal to, equal to relationship. For example, the equality constraint can always be equivalently expressed by two less than or equal to inequalities. i It is an additional optimization variable. When the implicit function is given by a pure system of equations, this variable can be omitted. When it is given by the optimization problem, this variable represents the dual multiplier variable.

[0084] Generally speaking, optimization problems (1-6) are high-dimensional nonlinear optimization problems that may contain a series of highly nonlinear constraints, including complementary slack conditions. There are two general approaches to solving such complex optimization problems: the first is to introduce 0 / 1 variables for linearization, typically employing piecewise linear approximation and the Big M method. This conversion results in a 0 / 1 integer programming problem, which is then solved using a branch-and-bound (BnB) algorithm. Another approach is to employ intelligent algorithms for approximate solutions, typically requiring repeated testing of the initial feasible solution and hyperparameter settings. Commonly used methods include particle swarm optimization (PSO) and genetic algorithms (GA).

[0085] Step S13: Generate synthetic data using the calibrated prior analytical model, merge the historical data with the synthetic data to generate a hybrid data set.

[0086] Specifically, synthetic data is generated as follows:

[0087] Step S131 , performing performance analysis on the calibrated multiple priori analytical models to obtain the fitting accuracy and estimation result of each priori analytical model.

[0088] Step S132: Screen out a priori analytical models whose fitting accuracy meets preset requirements.

[0089] Step S133 : For the selected priori analytical models, weighted averaging is performed on the estimation results of the priori analytical models according to the fitting accuracy relationship to generate synthetic data.

[0090] In the embodiment of the present invention, after completing the calibration of all prior annotation models, the next step is to analyze the model performance and generate the final synthetic data based on the model estimation results. The basic method of generation is to select a portion of models with the best fitting effect and perform a weighted average of their estimation results.

[0091] Take a set of electricity price series λ i ,i≥n+1, the synthetic data to be generated (λ i ,p i ) is calculated by the following formula:

[0092]

[0093] Where, γ k Is a normalized weighted coefficient, ensuring that the sum of all weighted coefficients is 1, and each coefficient is consistent with the model fitting error Inversely proportional. I(·) represents an indicator function with only two possible values: 0 and 1. It takes on the value 1 when the internal conditions are met and 0 otherwise. In addition, α and β are two user-defined parameters that can be set based on the overall performance of the model and the application scenario requirements.

[0094] Equations (1-7) demonstrate two strategies for improving the effectiveness of synthetic data. The first is to directly discard model estimation results that do not meet the fitting accuracy requirement of β. The second is to perform a weighted average of the estimation results that meet the estimation requirements based on the fitting accuracy relationship, where the estimation results of high-precision models are given greater weight.

[0095] The expression of the synthetic dataset is as follows:

[0096] SD={(λ n+1 ,p n+1 ),…,(λ N ,p N )} (1-8)

[0097] Here, we must ensure that N > dimθ. Note that to maintain a simple symbol system and avoid introducing too many symbols, this application uses the same vector symbols to represent historical data and synthetic data. The only difference between the two is the range of the subscript value.

[0098] The historical dataset is then merged with the synthetic dataset to obtain a hybrid dataset D = HD∪SD.

[0099] Step S14: training the data-driven model using the mixed data set.

[0100] Specifically, the above step S14 includes the following process:

[0101] Step S141, dynamically extract small batches of data from the mixed data set.

[0102] Step S142: train the data-driven model using small batch data.

[0103] In the embodiments of the present invention, after preparing the synthetic data, the next task is to design an appropriate training strategy to ensure that these data resources can be fully utilized while effectively reducing the side effects caused by the inaccuracy of the synthetic data during the training process. It should be noted that since this application uses neural networks as a specific example of data-driven models, the training method of this application is primarily applicable to neural network models.

[0104] In common regression tasks, the training and calibration process of neural networks is essentially to minimize the loss function J N (θ), and generally the gradient descent method is used to gradually optimize. The expression of the loss function is as follows:

[0105]

[0106] Where N is the amount of data.

[0107] A lot of practice shows that using the entire dataset in gradient descent calculations often leads to slow calculation speed and poor training effect. Based on this experience, in order to speed up training, a small batch Dm={(λ (1) ,p (1) ),…,(λ (m) ,p (m) )}, and use this small batch to iteratively correct the model parameters θ. And this small batch data will be randomly updated before each round of calculation. It should be noted that here λ (1) Different from the previous λ1, the subscripts use parentheses to represent the order within the mini-batch data.

[0108] Formula (1-10) shows the gradient calculation results based on small batch data:

[0109]

[0110] In the formula, the gradient term This can be calculated using the classic backpropagation algorithm (BP). Since the mini-batch size is randomly updated in each round, the above gradient formula needs to be adjusted accordingly. This mechanism actually introduces a certain amount of random perturbation in the gradient calculation.

[0111] According to the gradient result given by formula (1-10), the model weight parameters can be updated using formula (1-11):

[0112]

[0113] Where s represents the learning rate, which is essentially an iterative step size. It can be set using a fixed step size or a dynamic step size mechanism. To ensure efficient training, the Adam algorithm is used to update weights. This algorithm uses a dynamic step size mechanism.

[0114] The key difference between the dynamic mini-batch training method developed in this application and traditional mini-batch training lies in the nature of the dataset. In other words, the coexistence of real historical data and synthetic data must be carefully handled. Otherwise, the data noise introduced by the imprecision of the synthetic data can significantly degrade the training quality, seriously affecting the accuracy and robustness of the model. Therefore, the concept of dynamic sampling is introduced here. Its core idea is to gradually adjust the ratio of historical data to synthetic data as the training process progresses.

[0115] Given a small batch of data, we can divide it into two parts, m1 and m2, with m = m1 + m2. Here, m1 = |Dm∩HD|, corresponding to the amount of data collected from the historical dataset, and m2 = |Dm∩SD|, corresponding to the amount of data collected from the synthetic dataset. We then establish a proportional coefficient to specifically calculate the proportion of historical data:

[0116]

[0117] Where η max and η min Represents the upper and lower bounds of the dynamic ratio, which can be set according to specific needs. Generally speaking, the lower bound can be set by default here is a floor rounding function. The default value is the expected value when considering uniform sampling. As for the upper bound value, it is generally recommended not to exceed 90%.

[0118] The core idea of ​​dynamic sampling is simple: to ensure that η gradually increases during training, allowing real historical data to play a stronger guiding role in the later stages of training. This can, to a certain extent, suppress the potential impact of data noise in the synthetic data. Specifically, η can be initialized to a lower bound and then gradually updated according to the following calculation formula:

[0119]

[0120] Where T is a large constant that determines the rate of change of the dynamic ratio. At the same time, the ratio η is limited to the upper and lower bounds.

[0121] It is easy to see that the proportional coefficient η can be used to establish a dynamic mini-batch. The specific method is as follows: First, randomly sample from the historical data set data, and then randomly sampled from the synthetic dataset Finally, merge the two sets of data.

[0122] By utilizing the dynamic mini-batch training method, it can efficiently process mixed data sets consisting of real data and synthetic data, taking into account the effectiveness and computational efficiency of the training process.

[0123] Figure 3 The four data sets mentioned above have been sorted out and summarized. The gray vertical lines in the figure roughly indicate the model parameter quantities of the data-driven model and the prior analytical model. It can be seen that the historical data set can effectively fit the prior analytical model, but it is still a long way from the data demand of the data-driven model. However, after data enhancement, the hybrid data set can effectively guide the training of the data-driven model.

[0124] The model-guided and data-driven price response characteristic identification method provided by the present invention includes: obtaining historical data of the electricity market; calibrating the prior analytical model using the historical data; generating synthetic data using the calibrated prior analytical model, merging the historical data with the synthetic data to generate a hybrid data set; and training the data-driven model using the hybrid data set. The present application proposes a data enhancement method guided by a priori analytical model. This method cleverly utilizes the analytical model resources accumulated in the application scenario, and embodies the concept of "knowledge-guided, data-driven". By making full use of the analytical model that reflects the market operation mechanism to generate synthetic data, the amount of sample data is expanded, and the training process of the data-driven model is efficiently guided, effectively avoiding the risk of overfitting. After the transformation, the overall process is transformed into a data-driven modeling method guided by the analytical model, which effectively solves the problem of efficient training of the data-driven model under the condition of insufficient data, thereby ensuring the orderly development of the offline training link of the stacked architecture.

[0125] In an optional embodiment, the method for setting the relevant hyperparameters is specifically described, including: data volume N, electricity price series λ n+1 ,…,λ N , parameter α and parameter β.

[0126] Generally speaking, the amount of data N needs to be large enough to ensure that it exceeds the number of parameters of the data-driven model, which can be expressed mathematically as N>dimθ. n+1 ,…,λ N The setting of is more flexible, but the basic principle to be followed is to minimize the overlap with existing historical data. This ensures that the generated synthetic data effectively provides additional information not provided by the historical data. Specifically, this series can be generated by randomly sampling from an interval or space, and the sampling density can be increased in spaces where the historical data is sparsely distributed.

[0127] The settings of parameters α and β are highly flexible and can be adjusted according to the needs of specific application scenarios. Common value ranges are: 75%≤β≤85%.

[0128] In an optional embodiment, the prediction error, loss function and gradient noise of the data-driven model are analyzed with the help of machine learning.

[0129] Specifically, some cutting-edge theories in the field of machine learning are used to explain the effectiveness of the data enhancement and dynamic mini-batch methods established in this application. Figure 4 Based on the backpropagation process of the training error, we have sorted out three improvements brought about by the proposed method. The following content will further explain these three aspects.

[0130] (1) More effective prediction error information feedback

[0131] The imprecision of synthetic data is a significant issue, requiring careful consideration of potential uncertainties arising from the inaccuracy of some data. This naturally raises concerns about the effectiveness of the proposed method. It should be noted that the introduction of synthetic data presupposes insufficient data, so the comparison should also be to training conditions with insufficient data. From this perspective, inaccurate training data can still be beneficial for improving model generalization capabilities, as it can provide inaccurate information about the function manifold, thus playing a guiding role.

[0132] Figure 5 An explanatory case is used to show the prediction error under different function manifolds. It should be noted that the function manifold is an abstract mathematical concept, which can be simply understood as the interface form of a high-dimensional function. Five data points are specifically marked in the figure, where A and B represent two historical data, that is, A, B∈HD; C2 is a synthetic data, that is, C2∈SD; C1 is an imaginary point taken from the real function manifold, and its horizontal coordinate is consistent with C2; and C3 is the estimated result obtained at the same horizontal coordinate position when the data is insufficient. Comparison Figure 4 From the three function manifold curves in , we can find that although there is still a certain difference between the synthetic data C2 and the real result C1, it is still better than the case where there is no information at all at this position.

[0133] The following introduces some specific symbols, assuming that the coordinates of C1, C2, and C3 are (λ i ,p i ) and (λ i ,f(x i ;θ))(i≥n+1), then the manifold difference can be analyzed using the following formula:

[0134]

[0135] Observe the right side of equation (1-14), the first term represents the difference between C1 and C3, while the second term represents the difference between C1 and C2. Figure 5 In the example, although the synthetic data is not perfect, as shown by the second term being greater than zero, this imprecise data is still valuable, at least better than having no information at all, as shown by the first term being greater than the second term in the formula.

[0136] It should be noted that the above situation is not an isolated one. In fact, as long as the prior analytical model is set up reasonably, synthetic data can provide a certain degree of performance improvement. In other words, if the characteristics described by Equation (1-15) can be achieved in a probabilistic manner, it can be considered that synthetic data can help improve the accuracy of manifold characterization.

[0137]

[0138] Where E(·) represents the mean function.

[0139] (2) Loss function with better generalization performance

[0140] Generalization performance is a crucial aspect of evaluating data-driven models, directly determining their usability in real-world applications. Neural network training naturally focuses on generalization performance. Although training is performed on a training set, the expectation is that the model will demonstrate excellent performance on a test set. This fully demonstrates the fundamental difference between neural network training and general optimization problems.

[0141] Generalization performance is closely related to the quality of the loss function. The loss function is actually an empirical risk criterion. For regression tasks, an ideal loss function can be expressed as follows:

[0142] J(θ)=∫(f(λ;θ)-p) 2 dp(λ,p) (1-16)

[0143] Where dp(λ,p) represents the true distribution of the data.

[0144] The true distribution is often not available in real applications. A feasible approximation is to estimate the ideal function J(θ) using a limited number of data points. For example, using the historical data set HD or the mixed data set D can produce two estimation results J n (θ) and J N (θ), its specific mathematical expression is:

[0145]

[0146] Theoretically, if the number of data points used is insufficient, the error of the above approximation will be larger. This also shows that when the data is insufficient, the model training efficiency is low and overfitting may occur. The reason behind this is that the loss function J n (θ) deviates greatly from the theoretical situation J(θ), and can no longer guarantee that the empirical risk is effectively reduced, and the generalization ability of the model is seriously affected.

[0147] Using synthetic data to expand the overall data volume is a feasible way to improve the above approximation effect. Although the inaccuracy of synthetic data will bring some negative effects, it is found in actual applications that the benefits it brings are often greater. In short, J N (θ) is generally higher than J n (θ) is more reliable and the training effect will be better.

[0148] (3) Richer gradient noise

[0149] In addition to improving the prediction error and loss function, we can also analyze it from the perspective of gradient noise. Gradient noise describes the degree to which the training gradient deviates from the average gradient due to small batch sampling. The specific definition of gradient noise is as follows:

[0150]

[0151] Where GN(θ) is the gradient noise for the parameter vector θ. It should be noted that the dimension of the gradient noise is the same as the parameter vector, so it is also a high-dimensional vector mathematically. m (θ) is the loss function value on the small batch, which is exactly the same as formula (1-18), and The expression of has been given in formula (1-10).

[0152] According to modern machine learning theory, gradient noise improves the local search capabilities of neural network training, effectively preventing training from becoming trapped in undesirable local areas. Mathematically, gradient noise introduces randomness into the gradient descent algorithm used in training, creating a random search mechanism. The greatest advantage of random search is its ability to effectively escape local depressions and search for optimal results in a wider range of spaces.

[0153] It is generally believed that the stronger the gradient noise, the better the generalization effect of model training. Some researchers also regard gradient noise as a special regularization method. It is generally believed that gradient noise is mainly caused by the use of small batch sampling technology in training.

[0154] However, in addition to the randomness introduced by small batches, the method proposed in this application also provides an additional layer of diversity, namely the diversity introduced by synthetic data and its dynamic proportions. These factors make the gradients during training more variable, which helps improve training efficiency.

[0155] Combining knowledge guidance with data-driven approaches maximizes the use of existing data and knowledge resources, thereby designing a more efficient price response characteristic identification model. In practical applications, insufficient data is a common problem, and the method proposed in this application has great potential for promotion and application:

[0156] First of all, this method fully embodies the knowledge-guided and data-driven design concept. By utilizing data enhancement and dynamic mini-batch technology, it can efficiently integrate the data-driven flexibility with the mechanism-driven characteristics of the prior analytical model, ultimately greatly improving the sample efficiency during the training process.

[0157] Secondly, real-world application scenarios often contain a wealth of auxiliary information resources, which are generally difficult to directly utilize with traditional data-driven models and are therefore often overlooked and not effectively utilized. In contrast, a priori analytical models can more comprehensively consider these potential resources and improve the training efficiency of data-driven modeling by synthesizing data.

[0158] In one specific implementation, the case study considers two test systems: the first test system is a single user with time-shifted loads, and the second test system is a factory with a series of time-shifted loads. This application will first carefully analyze the practical effects of the "model-guided, data-driven" approach using the first test system, while simultaneously conducting robustness testing using the second test system to verify the applicability of this "model-guided, data-driven" approach to systems with more diverse characteristics and larger scale.

[0159] There are two common ways to model time-series load shifting: one is based on linear programming and the other is based on mixed integer programming. The former generally takes utility maximization as the objective function, considering the upper and lower bounds of electricity consumption, the limit on the rate of change of electricity consumption, and the daily minimum electricity consumption constraint. The latter generally considers the characteristics of load components and the shift characteristics of discreteness in more detail, and usually further refines the modeling based on the former. This application will consider both modeling methods, and the model details and parameter values ​​are taken from the paper.

[0160] The example uses day-ahead electricity price data from the PJM market, covering the period from January 1, 2018, to December 31, 2019. This dataset is split into three parts for training, validation, and testing. Specifically, the training set covers the period from January 1, 2018, to June 30, 2019; the validation set covers the period from July 1, 2019, to September 30, 2019; and the test set covers the period from October 1, 2019, to the end of the year. The data volume ratio of these three datasets is approximately 6:1:1.

[0161] The price response dataset can be generated through model optimization. Here, a time-shifted load model using mixed integer programming is used, and the electricity price data is the PJM market electricity price. The resulting dataset spans two years and includes 546 training data points, 92 validation data points, and 92 test data points. Obviously, this is a very small dataset and is generally unsuitable for training medium- to large-scale data-driven models.

[0162] The a priori analytical models considered in this section fall into four categories: time-shifted load models, temperature-controlled load models, and electric vehicle load models. The time-shifted loads used as a priori analytical model are modeled using linear programming. For each model type, 12 a priori analytical models were constructed by randomly sweeping model parameters and randomly removing constraints from the equations, with four models of each type constructed.

[0163] In addition, the subsequent examples analyze fully connected neural networks as a specific example of data-driven models. In fact, neural networks are indeed more suitable for this type of high-dimensional regression learning task.

[0164] All subsequent examples are written in Python 3.6.0, the neural network is modeled using TensorFlow 1.12.0, and the optimization problem is solved using Gurobipy 8.0.0. The computer is equipped with an i7-8550U CPU and 16.0GB of RAM.

[0165] Consider the first test system. At this time, there are only 546 historical data available for training. Therefore, it is necessary to calibrate the prior analytical models first and pay attention to the fitting accuracy of these models. Figure 6 Specifically shows the overall situation of estimation accuracy.

[0166] Depend on Figure 6 As can be seen, the prior analytical model for time-shifted loads performs best. This result is not surprising, as different time-shifted loads generally share some similarities. Even if the optimization decision models are not identical, they are often reflected in some local characteristics. The fitting accuracy of both the temperature control load and the electric vehicle load models is poor, and they can be eliminated based on the fitting accuracy.

[0167] Then, according to the synthetic data aggregation method established in this application, the results of the time series shift model and the elastic matrix model are weighted averaged to generate generated data. Specifically, 1500 data points are generated here. These data points are merged with the historical training data set to obtain a hybrid data set, containing a total of 2046 data points. Once the data volume increases, it can be used to train medium-sized neural networks.

[0168] Figure 7 The convergence curves of neural network learning and training before and after the data enhancement method in this chapter are compared. Figure 7It can be seen that after data augmentation and dynamic mini-batch training, the training efficiency of the neural network has improved significantly, as evidenced by the faster descent of the data-augmented curve in the figure. Furthermore, the convergence value of the data-augmented curve is lower than that of the original curve, indicating that the training quality of the model has improved significantly after data augmentation. A closer look at the convergence curve also reveals that the curve with data augmentation has slightly stronger fluctuations than the original curve, a phenomenon that is partly related to the richer gradient noise during training.

[0169] Next, we compare the estimation accuracy of the neural network before and after using the method in this paper, and the results are summarized in Table 1.

[0170] Table 1 Effect differences before and after data augmentation

[0171]

[0172] Note that Table 1 shows the estimated accuracy on the test set. The numbers before the slash represent the original situation, while the numbers after the slash represent the results after applying the data augmentation method in this paper. Comparing the numbers before and after the slash shows that for different network structures (corresponding to different rows) and different activation function settings (corresponding to different columns), the data augmentation method proposed in this application can roughly improve the estimated accuracy by about 5%.

[0173] Nevertheless, after adopting this method, if it is combined with an appropriate network structure, it is estimated that the performance can be further improved. Table 1 shows a general trend, that is, choosing a larger neural network size within the allowable range to improve nonlinear fitting capabilities is a more appropriate selection strategy.

[0174] To verify the universality of this proposed method, we conducted further analysis on a second test system. To conduct large-scale, multi-scenario testing, we varied the number of time-shifted loads in the factory from 10 to 50, focusing on the effectiveness of data augmentation and dynamic mini-batch training at varying system scales. The comparative results are summarized in Table 2.

[0175] Table 2 Performance of the algorithm under large-scale testing

[0176]

[0177] The data in Table 2 represent the estimation accuracy of the neural network. A careful analysis of the data in the table shows that the performance of this application is basically stable in systems of different sizes. Although there are slight fluctuations in some scenarios, the overall performance has neither increased significantly nor decreased significantly. In real-world applications, this stable performance is important, which shows that the "knowledge-guided-data-driven" approach has guiding significance in both large and small systems, and the data enhancement and dynamic mini-batch methods designed based on this approach can also maintain relatively stable performance. The general conclusions of Table 2 are similar to those of Table 1, and it also shows that selecting larger neural networks within the allowable range is beneficial to improving estimation accuracy.

[0178] Numerical results, across multiple test systems, demonstrate the superiority of the proposed method. Testing has shown that the proposed method can reduce data requirements by 30%–50% while maintaining modeling accuracy. Furthermore, its performance is further enhanced in application scenarios with richer domain knowledge and more comprehensive auxiliary information. Calibration of the prior analytical model provides data-driven models with imprecise manifold information and a smoother loss function interface, significantly improving the efficiency of data-driven model training.

[0179] This embodiment also provides a model-guided, data-driven price response characteristics identification device, which is used to implement the above-mentioned embodiments and preferred implementations. Details already described are omitted for clarity. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. While the devices described in the following embodiments are preferably implemented using software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0180] This embodiment provides a model-guided and data-driven price response characteristic identification device. Figure 8 Shown, including:

[0181] The acquisition module 81 is used to acquire historical data of the electricity market.

[0182] The calibration module 82 is used to calibrate the prior analytical model using historical data.

[0183] The merging module 83 is used to generate synthetic data using the calibrated prior analytical model, and merge the historical data with the synthetic data to generate a hybrid data set.

[0184] The training module 84 is used to train the data-driven model using the mixed data set.

[0185] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0186] The model-guided and data-driven price response characteristic identification device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above-mentioned functions.

[0187] The model-guided, data-driven price response characteristics identification device provided by this invention cleverly leverages the analytical model resources accumulated in application scenarios. It fully utilizes analytical models that reflect market operating mechanisms to generate synthetic data, expand the sample data volume, and effectively guide the training process of the data-driven model, effectively avoiding the risk of overfitting. Furthermore, after modification, the overall process is transformed into a data-driven modeling method guided by analytical models, effectively solving the problem of efficient data-driven model training under insufficient data conditions, thereby ensuring the orderly execution of the offline training phase of the cascade architecture.

[0188] The embodiment of the present invention also provides a computer device having the above Figure 8 The model-guided-data-driven price response characteristics identification device is shown.

[0189] See also Figure 9 , Figure 9 is a structural diagram of a computer device provided by an optional embodiment of the present invention, such as Figure 9 As shown, the computer device includes: one or more processors 10, memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components utilize different buses to communicate with each other and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in the memory or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Equally, multiple computer devices can be connected, and each device provides part of the necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 9 A processor 10 is taken as an example.

[0190] The processor 10 may be a central processing unit, a network processor, or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.

[0191] The memory 20 stores instructions that can be executed by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.

[0192] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0193] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0194] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or a communication network.

[0195] The embodiment of the present invention also provides a computer-readable storage medium. The above-mentioned method according to the embodiment of the present invention can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.

[0196] Although the embodiments of the present invention have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention. Such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A model-guided and data-driven price response characteristics identification method, characterized in that: The method comprises: Obtain historical data on the electricity market; calibrating the prior analytical model using the historical data; generating synthetic data using the calibrated prior analytical model, and merging the historical data with the synthetic data to generate a hybrid data set; The data-driven model is trained using the mixed dataset.

2. The model-guided and data-driven price response characteristics identification method according to claim 1, characterized in that: The training of the data-driven model using the mixed data set includes: Dynamically extracting small batches of data from the mixed dataset; The data-driven model is trained using the mini-batch data.

3. The model-guided and data-driven price response characteristics identification method according to claim 2, characterized in that: The training of the data-driven model using the mixed data set further includes: During the training process, the proportion of historical data in the small batch data is gradually increased.

4. The model-guided and data-driven price response characteristics identification method according to claim 1, characterized in that: The calibrating the prior analytical model using the historical data includes: Construct multiple prior analytical models by randomly sweeping the model parameters and randomly removing the constraints from the equations; Multiple prior analytical models are calibrated one by one according to the historical data.

5. The model-guided and data-driven price response characteristics identification method according to claim 4, characterized in that: The generating of synthetic data using the calibrated prior analytical model comprises: Perform performance analysis on multiple calibrated prior analytical models to obtain the fitting accuracy and estimation results of each prior analytical model; Screening out the prior analytical model whose fitting accuracy meets the preset requirements; For the selected priori analytical models, weighted averaging is performed on the estimation results of the priori analytical models according to the fitting accuracy relationship to generate synthetic data.

6. The model-guided and data-driven price response characteristics identification method according to claim 4, characterized in that: The prior analytical model is calibrated using the following optimization model: Where, e k is the model output gk(λi; w k ) and the actual observed value p i Minimizing the mean square error will yield a set of optimal parameters And the corresponding minimum error value 7. The model-guided and data-driven price response characteristics identification method according to claim 1, characterized in that: Analyze the prediction error, loss function, and gradient noise of the data-driven model using machine learning.

8. A model-guided and data-driven price response characteristics identification device, characterized in that: The device comprises: Acquisition module, used to obtain historical data of the power market; a calibration module, configured to calibrate the prior analytical model using the historical data; a merging module, configured to generate synthetic data using the calibrated prior analytical model, and merge the historical data with the synthetic data to generate a hybrid data set; A training module is used to train a data-driven model using the mixed data set.

9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the model-guided-data-driven price response characteristic identification according to any one of claims 1 to 7 by executing the computer instructions.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions for causing a computer to execute the model-guided-data-driven price response characteristic identification according to any one of claims 1 to 7.