Method, system and equipment for predicting price difference of electricity price

By using the random forest algorithm and the minimum mean square deviation algorithm in the power spot market for iterative segmentation and multi-decision tree average calculation, the problem of low real-time price difference prediction accuracy in the existing technology has been solved, and higher prediction accuracy has been achieved.

CN120146882AInactive Publication Date: 2025-06-13BEIJING LANMUDA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510145233.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When the prior art predicts the recent real-time price difference in the spot power market, it is easy to overfit due to data distribution limitations, resulting in low prediction accuracy.

Method used

The random forest algorithm and the minimum mean square deviation algorithm are used to randomly select data samples and iteratively divide them to form multiple subsets, and calculate the average value of the subset output values ​​as the output value of the decision tree. Finally, multiple decision trees are used to determine the prediction result.

Benefits of technology

Improve the accuracy of the recent real-time spread predictions, and avoid the possible underfitting and overfitting of a single decision tree model when dealing with complex problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146882A_ABST
    Figure CN120146882A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data integration, and discloses an electricity price difference prediction method, system and equipment, and the method comprises the steps: obtaining a data set of a to-be-predicted day, the data set comprises multi-class boundary condition data, multi-class weather condition data and day-ahead real-time price difference data, the day-ahead real-time price difference data is a difference value between the day-ahead price and the real-time price; sending the data set to a price difference prediction model, wherein the price difference prediction model randomly extracts data of the data set to determine a plurality of sample sets; and traversing the plurality of sample sets, carrying out iterative segmentation on the current sample set according to data in the sample sets to form a plurality of subsets of the sample sets, obtaining a target value of each subset, and determining a prediction result according to the target value of each subset. According to the technical scheme provided by one or more embodiments, the accuracy of predicting the day-ahead real-time price difference in the electric power spot market can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data integration, and in particular, to a method, system, and device for predicting electricity price spreads. Background Art

[0002] The spot trading in the electricity spot market mainly includes the day-ahead, intraday, and real-time electricity energy trading markets. Among them, when a power selling company declares electricity quantity and electricity price in the day-ahead trading market, it declares the electricity demand and corresponding price for each period of the next 24 hours to the power dispatching agency. Therefore, how the power selling company makes a reasonable quotation has a great impact on its electricity consumption cost and revenue level.

[0003] In practical applications, the day-ahead real-time spread can be predicted to determine the quotation for the next day. The day-ahead real-time spread is the difference between the day-ahead price and the real-time price. In related technologies, the k-nearest neighbor model or elastic net is usually used for spread prediction, but limited by the data distribution, it is prone to overfitting, resulting in low prediction accuracy.

[0004] In view of this, how to improve the accuracy of predicting the day-ahead real-time spread is an urgent problem to be solved. Summary of the Invention

[0005] This application provides a method, system, and device for predicting electricity price spreads, which can improve the accuracy of predicting the day-ahead real-time spread in the electricity spot market.

[0006] In the first aspect of this application, a method for predicting electricity price spreads is provided. The method includes: obtaining a data set for a day to be measured, where the data set includes multi-category boundary condition data, multi-category weather condition data, and day-ahead real-time spread data. The multi-category weather condition data is the ratio of the weather condition difference to the reference value, and the day-ahead real-time spread data is the difference between the day-ahead price and the real-time price; sending the data set to a spread prediction model, and the spread prediction model randomly extracts data from the data set to determine multiple sample sets; traversing the multiple sample sets, and iteratively dividing the current sample set according to the condition data in the sample set to form multiple subsets of the sample set, obtaining the target values of the respective subsets, and determining a prediction result according to the target values of the respective subsets.

[0007] The second aspect of the present application provides an electricity price spread prediction system, which includes: a data processing unit for obtaining a data set of a to-be-measured day, where the data set includes multi-category boundary condition data, multi-category weather condition data, and day-ahead real-time spread data. The multi-category weather condition data is the ratio of the weather condition difference to a reference value, and the day-ahead real-time spread data is the difference between the day-ahead price and the real-time price; a model prediction unit for sending the data set to a spread prediction model. The spread prediction model randomly extracts condition data from the data set of the data set to determine multiple sample sets. Each sample set includes multiple condition data, and the type of the condition data is one of multi-category boundary condition data, multi-category weather condition data, and day-ahead real-time spread data. Traverse the multiple sample sets, and iteratively divide the current sample set according to the condition data in the sample set to form multiple subsets of the sample set, obtain the target values of each subset, and determine the prediction result according to the target values of each subset.

[0008] The technical solution provided by one or more embodiments of the present application can improve the accuracy of predicting the day-ahead real-time spread according to the relevant data of the to-be-measured day and the spread prediction model constructed based on the random forest algorithm. Specifically, according to the random forest algorithm and the least mean square algorithm, multiple subsets are divided and the average value of the subset output values is calculated as the output value of the current decision tree, and then the average value of multiple decision trees is used to determine the final output result.

[0009] It can be seen that the technical solution provided by the present application, using the random forest algorithm of multiple decision trees, can capture the linear relationship in the data, avoid the underfitting situation that may occur when a single decision tree model processes complex problems, and thus improve the accuracy of predicting the day-ahead real-time spread. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0011] Figure 1 It is a schematic diagram of the steps of a method for predicting an electricity price spread provided by an embodiment of the present application;

[0012] Figure 2 It is a schematic diagram of the steps for training a spread prediction model provided by an embodiment of the present application;

[0013] Figure 3 It is a schematic diagram of the structure of an electricity price spread prediction system provided by an embodiment of the present application;

[0014] Figure 4 Schematic structural diagram of a computer device provided for an embodiment of the present application. Specific embodiments

[0015] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0016] In addition, the descriptions involving "first", "second", etc. in the present application are only for descriptive purposes, and cannot be construed as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" is two or more. In addition, the use of "based on" or "according to" means open and inclusive, because a process, step, calculation, or other action "based on" or "according to" one or more of the stated conditions or values may in practice be based on additional conditions or values beyond the stated ones.

[0017] In the electricity spot market, market participants include power generation enterprises, electricity retailers, electricity users, ancillary service providers, etc. The above-mentioned market participants conduct transactions according to market rules. As an important part of the electricity spot market, the day-ahead trading market buys and sells the electricity demand for the next day according to the declared price. Among them, the declared price can be adjusted by predicting the day-ahead real-time price difference. Therefore, if the day-ahead real-time price difference can be accurately predicted, the profitability of market participants can be improved. On the contrary, if the accuracy of the predicted day-ahead real-time price difference is low, it may cause losses to market participants.

[0018] It should be noted that the day-ahead real-time price difference is the difference between the day-ahead price and the real-time price. The above-mentioned day-ahead price is the price for selling electricity on the prediction day one day before the prediction day, and the above-mentioned real-time price is the price for selling electricity at 5 minutes or 15 minutes before the electricity transaction on the prediction day. In the electricity spot market pilot project in China, due to the short time interval between trading and delivery, the electricity price can sensitively reflect the changes in market supply and demand, and the electricity price will fluctuate over time. Therefore, the current day-ahead trading market usually uses 15 minutes or 60 minutes as a trading period. Among them, in each trading period, the market operation agency will determine and announce the real-time price according to the supply and demand situation for market participants to refer to.

[0019] In the related art, given that there is a certain regularity in the day-ahead real-time price difference, most rely on the influencing factors of the day-ahead real-time price difference to construct prediction models such as the k-nearest neighbor model or the elastic net for price difference prediction. However, these methods are often restricted by the data distribution, prone to overfitting, which thus affects the accuracy of the prediction. In recent years, the decision tree model has received extensive attention due to its need not to perform cumbersome data standardization processing and anti-overfitting ability. However, a single decision tree model is prone to underfitting when dealing with complex data, which will also affect the low accuracy of the day-ahead real-time price difference.

[0020] In view of this, one or more embodiments of the present application provide a method, system and device for predicting the electricity price difference, which can solve the above problems and improve the accuracy of predicting the day-ahead real-time price difference.

[0021] Please refer to Figure 1 , an embodiment of the present application provides a method for predicting the electricity price difference, and the method may include the following steps:

[0022] S1: Obtain a data set for the day to be measured, where the data set includes multi-category boundary condition data, multi-category weather condition data, and day-ahead real-time price difference data. The multi-category weather condition data is the ratio of the weather condition difference to the reference value, and the day-ahead real-time price difference data is the difference between the day-ahead price and the real-time price.

[0023] Since there is a certain regularity in the day-ahead real-time price difference, its change trend is affected by the boundary condition data and the weather condition data. The data categories of the above multi-category boundary condition data include electric load data, wind power data, photovoltaic power data, and tie line data. The data categories of the above multi-category weather condition data include temperature data, light intensity data, humidity data, and wind speed data.

[0024] Exemplarily, temperature fluctuations will affect the demand for electricity. For example, the increased use of air conditioners and heating equipment leads to an increase in electricity demand, an increase in the real-time price, which thus affects the day-ahead real-time price difference. Exemplarily, the light intensity data will affect the output of photovoltaic power generation, the humidity data will affect the output of hydropower generation, and the wind speed data will affect the efficiency of wind power generation, which in turn affects the day-ahead real-time price difference. Similarly, the electric load data reflects the real-time change of electricity demand, and the electricity price will change in real time according to the electricity demand. Similarly, the volatility and unpredictability of wind power and photovoltaic power will affect the power supply, which in turn causes changes in the day-ahead real-time price difference. Therefore, obtaining the multi-category boundary condition data and multi-category weather condition data of the prediction day can predict the change trend of the day-ahead real-time price difference, and the change trend of the day-ahead real-time price difference can provide a reference for the declaration of the day-ahead price.

[0025] In this embodiment, the uncertainties of various weather conditions are used as weather condition data, and the uncertainty of the above weather conditions is the ratio of the difference between the maximum and minimum values of the weather conditions to the reference value. Exemplarily, the highest temperature in a period is 30°, the lowest temperature is 10°, and the reference value of the preset temperature is 10°. Then the difference between the maximum and minimum values of the current temperature data is 20°, and the uncertainty of the current temperature data is 2. In addition, the above day-ahead real-time price difference data includes multiple day-ahead real-time price differences. The above day-ahead implementation price difference is the difference between the day-ahead price and the real-time price. The day-ahead price and the real-time price of the predicted day are provided by the power market, and the above real-time price includes the predicted prices of each trading period.

[0026] S3: Send the data set to the price difference prediction model. The price difference prediction model randomly extracts the data in the data set to determine multiple sample sets. Each sample set includes multiple condition data, and the type of the condition data is one of multi-category boundary condition data, multi-category weather condition data, and day-ahead real-time price difference data.

[0027] Each of the above sample sets contains multi-category boundary condition data, multi-category weather condition data, and day-ahead real-time price difference data. Among them, the above multi-category boundary condition data includes multiple electric load data, multiple wind power data, multiple photovoltaic power data, and multiple tie-line data. The above multi-category weather condition data includes multiple temperature data, multiple light intensity data, multiple humidity data, and multiple wind speed data. The above day-ahead real-time price difference data includes multiple day-ahead real-time price differences. Among them, each day-ahead real-time price difference has its corresponding multi-category boundary condition data and multi-category weather condition data.

[0028] S5: Traverse the multiple sample sets, and iteratively divide the current sample set according to the condition data in the sample set to form multiple subsets of the sample set, obtain the target values of each subset, and determine the prediction result according to the target values of each subset.

[0029] In this embodiment, when iteratively dividing the current sample set according to the condition data in each sample set, the multiple condition data in the current sample set can be iteratively divided according to the average value of the day-ahead real-time price difference data, or the multiple condition data in the current sample set can be iteratively divided according to the mean square deviation of the day-ahead real-time price difference data from the preset standard value.

[0030] In this embodiment, the target values of the respective subsets are obtained. The target value is an output result of the subset in the price difference prediction model. Based on the above target values, the day-ahead real-time price differences for each time period of the prediction day can be determined. Optionally, the average value of the output results of the respective subsets is used as the output result of the current sample set, and the average value of the output results of the respective sample sets is used as the prediction result of the data set. Optionally, the median value or the weighted average value can also be used as the prediction result. The prediction result of the above data set is used as the day-ahead real-time price difference for a certain time period of the prediction day.

[0031] In one embodiment, when making a prediction for the day to be measured, it is necessary to first obtain the data set for the three days prior to the day to be measured, including boundary condition data, weather condition data, and day-ahead real-time price difference data. Specifically, the prediction condition data for each trading time period of the three days prior to the day to be measured is obtained. The above prediction condition data includes boundary condition prediction data, weather condition prediction data, and day-ahead real-time price difference prediction data. The data categories of the above boundary condition prediction data include electric load prediction data, wind power prediction data, photovoltaic power prediction data, and tie line prediction data. The data categories of the above weather condition prediction data include temperature prediction data, light intensity prediction data, humidity prediction data, and wind speed prediction data.

[0032] Further, for each type of data in the boundary condition prediction data, sampling is performed every 15 minutes, that is, the values of each type of data in each trading time period of 15 minutes are obtained and used as a sampling value, and the average value of four adjacent sampling values is used as the boundary condition data for a corresponding data category. For example, the data for each trading time period of the electric load prediction data is sampled, and four adjacent sampling values are used as an electric load data in the data set.

[0033] Similarly, for each type of data in the weather condition prediction data, sampling is performed every 60 minutes, that is, the values of each type of data in each trading time period of 60 minutes are obtained and used as a sampling value, and each sampling value is used as the weather condition data for a corresponding data category. For example, the data for each trading time period of the temperature prediction data is sampled, and each sampling value is used as a temperature data in the data set.

[0034] Similarly, sampling is performed every 15 minutes for the day-ahead real-time price difference prediction data, that is, the values of the day-ahead real-time price difference prediction data in each trading time period of 15 minutes are obtained and used as a sampling value, and each sampling value is used as the day-ahead real-time price difference data in the data set.

[0035] In this embodiment, the boundary condition prediction data and the day-ahead real-time price difference prediction data are sampled every 15 minutes, and are aligned with the weather condition prediction data sampled every 60 minutes by taking the average value. Data disorder can be avoided through the unification in the time dimension.

[0036] In one embodiment, based on step S5, traverse the multiple sample sets, and iteratively divide the current sample set according to the condition data in the sample set to form multiple subsets of the sample set, including:

[0037] S501: Divide the current sample set into two regions, obtain the minimum value of the sum of the minimum mean square errors of the target values and the predicted values of the respective condition data in the two regions, and determine the pair (j, s) of the cut-off point according to the minimum value of the sum of the minimum mean square errors. Wherein, the target value of the condition data is the day-ahead real-time price difference corresponding to each condition data, and the predicted value of the condition data is the average value of each target value in the current region.

[0038] S503: Re-segment the current training sample set according to the pair (j, s) of the cut-off point to form a first subset R 1 (j, s) and a second subset R 2 (j, s), where

[0039] R 1 (j, s) = {x|x {j} ≤ s}, R 2 (j, s) = {x|x {j} ≤ s}, where j is a feature variable and s is a cut-off feature.

[0040] S505: Iteratively divide the regions of the first subset and the second subset to form multiple subsets.

[0041] In this embodiment, first perform regional pre-division, and by obtaining the pair (j, s) of the cut-off point, re-perform regional division according to the feature variable j and the cut-off feature s of the cut-off point. Specifically, the feature variable j represents one type of condition data among the electric load prediction data, wind power prediction data, photovoltaic power prediction data, and tie line prediction data in the boundary condition prediction data, and the temperature prediction data, light intensity prediction data, humidity prediction data, and wind speed prediction data in the weather condition prediction data. The cut-off feature s represents the condition data with the minimum mean square error between the target value and the predicted value under the feature variable in the current region. According to the pair (j, s) of the cut-off point represented by the current condition data, re-divide the region of the current sample set to form a first subset R 1 (j, s) and a second subset R 2 (j, s). Where R 1(j, s) = {x|x {j} > s}, R 2 (j, s) = {x|x {j} ≤ s}, which represents dividing each conditional data in the sample set according to the value of the segmentation feature s for the segmentation feature value of the feature variable j. The first subset represents that the segmentation feature of the conditional data of the current feature variable is greater than s, and the first subset represents that the segmentation feature of the conditional data of the current feature variable is less than or equal to s.

[0042] Exemplarily, calculate the minimum value of the sum of the minimum mean square errors of the target values and the predicted values of each sample in two regions. Among them, taking the feature data characterized by the minimum mean square error in the region with the minimum sum as the temperature data as an example, divide all the conditional data in the current sample set according to the temperature. Among them, j is the temperature data, and s represents that the uncertainty of the temperature is 2. Divide the conditional data in the current sample set with the temperature data greater than 2 into the subset R 1 , and divide the conditional data in the current sample set with the temperature data less than or equal to 2 into the subset R 2 .

[0043] Furthermore, continue to perform region segmentation on the subset R 1 and the subset R 2 according to steps S501 and S503 to form multiple subsets. Exemplarily, obtain the second segmentation point and perform region division on the subset R1 according to the second segmentation point to form the subsets R 3 and R 4 .

[0044] Among them, in step S501, obtaining the minimum value of the sum of the minimum mean square errors of the target values and the predicted values of each sample in the two regions can be calculated according to the following formula: The minimum value of the sum of the minimum mean square errors is x i is the feature vector, y i is the target value of the i-th conditional data, c 1 is the predicted value of the first region, and c 2 is the predicted value of the second region.

[0045] Among them, the feature vector x i represents a conditional data in each region, y i represents the target value of the i-th conditional data, which can be understood as the day-ahead real-time price difference corresponding to each conditional data, c 1 is the first predicted value in a region, preferably the average value of each target value in the current region, and c 2 is the second predicted value in another region, preferably the average value of each target value in the current region.

[0046] Specifically, in the pre-partitioned regions R 1 and R2 calculate the predicted value c of the current region 1 or c 2 For any region, calculate the difference between the target value and the predicted value of all conditional data within the current region, sum up the differences between the target value and the predicted value of the above various types of conditional data, and take the sum of the differences between the target value and the predicted value of the smallest type of conditional data, compare it with the sum value of another region and take the minimum value. For the region with the smallest sum value above, determine the conditional data with the smallest difference between the target value and the predicted value under the type of conditional data with the smallest sum of the differences between the target value and the predicted value, use the current conditional data as the pair (j, s) of the cut-off point, and re-divide the current sample set according to the above cut-off point.

[0047] Furthermore, obtain the target values of the respective subsets. Preferably, the above target value can be the day-ahead real-time price difference corresponding to the conditional data with the smallest difference between the target value and the predicted value within the current subset, calculate the average value of the target values of the respective subsets, use the average value of the target values of the respective subsets as the output value of the current sample set, calculate the average value of the output values of the respective sample sets, and use the average value of the output values of the respective sample sets as the predicted price difference.

[0048] Please refer to Figure 2 In one embodiment, send the data set to a price difference prediction model. The above price difference prediction model can be obtained through training. Specifically, the price difference prediction model can be trained according to the following steps:

[0049] S301: Obtain a historical data set and randomly sample with replacement from the historical data set to generate a plurality of training sample sets. The historical data set includes boundary condition historical data, weather condition historical data, and day-ahead real-time price difference historical data.

[0050] S303: For any training sample set, iteratively divide the training sample set until a preset stop condition is met to form a plurality of sub-regions, and generate a corresponding decision tree for each training sample set.

[0051] S305: Obtain a plurality of target values of the respective sub-regions, determine the output value of the current decision tree according to the plurality of target values of the respective sub-regions, and determine the final output value according to the output values of the respective decision trees.

[0052] In this embodiment, the above historical data set includes boundary condition historical data, weather condition historical data, and day-ahead real-time price difference historical data. The data categories of the above boundary condition historical data include electric load historical data, wind power historical data, photovoltaic power historical data, and tie-line historical data. The data categories of the above weather condition historical data include temperature historical data, light intensity historical data, humidity historical data, and wind speed historical data. It should be noted that the data in the above historical data set are all determined real data, and the above day-ahead real-time price difference historical data are the data collected according to historical real situations in the day-ahead trading market.

[0053] In this embodiment, for any training sample set, a decision tree of the current training sample set is established. Specifically, the training sample set is iteratively divided according to the steps from S501 to S505 until a preset stop condition is met to form a decision tree. Optionally, the preset stop condition may include that the number of the decision trees reaches a first preset value, the depth of each decision tree reaches a second preset value, and the number of samples included in a leaf node of each decision tree is less than or equal to a third preset value. Among them, the number of the above decision trees is the number of the training sample sets, the depth of the decision tree is the number of times of iterative segmentation of the current training sample set, and the number of samples included in a leaf node is the number of sub-regions divided in the last iteration.

[0054] In this embodiment, multiple target values of each sub-region are obtained, and the average value of the multiple target values of each sub-region is used as the output value of the decision tree, and the final output value is determined according to the output values of each decision tree.

[0055] Specifically, multiple target values of each sub-region are obtained, the average value of the multiple target values of each sub-region is calculated, and the average value of the multiple target values of each sub-region is used as the output value of the current decision tree. Among them, for the decision tree the output value of the current decision tree I is the identity matrix, M is the number of sub-regions, N m is the number of samples in the current sub-region, R m (j, s) is the m-th sub-region. Specifically, all sub-regions are traversed, and the average value c m of the target values of each sub-region R m is calculated, and c m is used as the output value of the decision tree. By taking the average value method, the accuracy of the current decision tree for the target value can be improved.

[0056] Further, the average value of the output values of each decision tree is calculated, and the average value of the output values of each decision tree is used as the final output value, and the final value k is the number of decision trees. By calculating the average of the output values of multiple decision trees, overfitting of the model can be avoided.

[0057] In one embodiment, after training to obtain the price spread prediction model, the price spread prediction model can be verified. Specifically, a verification set including historical boundary condition data, historical weather condition data, and historical day-ahead real-time price spread data is obtained, and the above verification set is input into the price spread prediction model, and the price spread prediction model outputs a prediction result. Further, the mean square error between the above prediction result and the historical day-ahead real-time price spread data is calculated, and the preset stop condition of the price spread prediction model is dynamically adjusted according to the magnitude of the mean square error. Wherein, the mean square error where n is the number of samples, y i is the day-ahead real-time price spread of the i-th sample, y i ^ is the prediction result of the i-th sample.

[0058] In this embodiment, the preset stop condition of the price spread prediction model can be modified according to the prediction result of the verification set, including the number of decision trees, the depth of the decision tree, and the number of samples included in a leaf node of the decision tree. For each modified price spread prediction model, the mean square error between its prediction result and the historical day-ahead real-time price spread data is calculated through the verification set, so that the value of the above mean square error continuously decreases until it reaches the minimum value that may occur. Taking the stop condition at this time as the final preset stop condition of the price spread prediction model can make the prediction result of the price spread prediction model closer to the actual historical day-ahead real-time price spread, thereby improving the accuracy of the price spread prediction model for predicting the day-ahead real-time price spread. In addition, a test set can also be used to test the prediction accuracy of the price spread prediction model,

[0059] Exemplarily, the present application provides an embodiment of training a price spread prediction model. The model is trained by obtaining data of a certain trading center in 2024. Specifically, the data from January 1, 2024 to August 31, 2024 of a certain trading center is used as the training set, the data from September 1, 2024 to September 30, 2024 is used as the verification set, and the data from October 1, 2024 to October 31, 2024 is used as the test set. Specifically:

[0060] First, a data set from January 1, 2024 to October 31, 2024 is obtained, including boundary condition data such as electric load data, wind power data, photovoltaic power data, and tie line data, weather condition data such as temperature data, light intensity data, humidity data, and wind speed data, and the day-ahead real-time price spread. Among them, the weather condition data is the uncertainty between the difference of the maximum and minimum values of the weather condition and the reference value. The reference values of the above temperature data, light intensity data, humidity data, and wind speed data are 25°C, 1000w / m 3, 150% and 1.10 m / s.

[0061] Divide the above dataset into a training set, a validation set, and a test set. Input the data of the training set into the spread prediction model for training. Input the data of the validation set into the trained spread prediction model. Dynamically adjust the parameters of the spread prediction model, i.e., the preset stop condition, according to the mean square error between the day-ahead real-time spread output by the spread prediction model and the true day-ahead real-time spread, so as to reduce the value of the above mean square error and improve the prediction accuracy of the spread prediction model for the day-ahead real-time spread. Subsequently, test the spread prediction model through the test set, using the minimum mean square error function and the direction accuracy rate as the measurement indicators. The above direction accuracy rate is the positive and negative directions of the day-ahead real-time spread, so as to ensure the accuracy of the prediction of the day-ahead real-time spread.

[0062] This application also provides an embodiment applying the above method for predicting the electricity price spread. This embodiment is carried out according to the following steps:

[0063] Step 1: Obtain the prediction condition data for the day to be measured, including electric load prediction data, wind power prediction data, photovoltaic power prediction data, tie line prediction data, temperature prediction data, light intensity prediction data, humidity prediction data, wind speed prediction data, and day-ahead real-time prediction spread. The above day-ahead real-time prediction spread data is the difference between the day-ahead predicted price and the real-time predicted price, and the above day-ahead predicted price and real-time predicted price are announced in advance by the market operator of the day-ahead trading market.

[0064] The above prediction condition data is the data of each trading period for the day to be measured and the two days after the day to be measured. Among them, taking 60 minutes as a trading period as an example, since the electric load prediction data, wind power prediction data, photovoltaic power prediction data, tie line prediction data, and day-ahead real-time prediction spread fluctuate greatly within 60 minutes, sample the above electric load prediction data, wind power prediction data, photovoltaic power prediction data, tie line prediction data, and day-ahead real-time prediction spread data every 15 minutes, and take the average value of four adjacent sampling values as an electric load data, wind power data, photovoltaic power data, tie line data, and day-ahead real-time spread data of the corresponding type. Sample the above temperature prediction data, light intensity prediction data, humidity prediction data, and wind speed prediction data every 60 minutes, and take the above sampling values as the temperature data, light intensity data, humidity data, and wind speed data of the corresponding type. Put the above electric load data, wind power data, photovoltaic power data, tie line data, temperature data, light intensity data, humidity data, wind speed data, and day-ahead real-time spread data into the dataset.

[0065] It should be noted that the above-mentioned day-ahead real-time price difference data includes the day-ahead real-time price differences of each trading period, and each of the above-mentioned day-ahead real-time price differences has its uniquely corresponding electric load data, wind power data, photovoltaic power data, tie-line data, temperature data, light intensity data, humidity data, and wind speed data.

[0066] Step 2: Input the data set into the price difference prediction model. The price difference prediction model divides the data set into multiple sample sets, conducts iterative region division for any sample set to form a decision tree for the current sample set, and outputs the target value of the current decision tree. Take the average of the target values of all decision trees as the output result of the price difference prediction model, which is the day-ahead real-time price difference of the optimized prediction day.

[0067] Specifically, for any decision tree, select the pair (j, s) of the optimal splitting point as the selection feature of the decision tree node. Here, j is the variable feature, which can be understood as one of the electric load data, wind power data, photovoltaic power data, tie-line data, temperature data, light intensity data, humidity data, and wind speed data, and s is the splitting feature, which can be understood as a conditional data determined according to the principle of minimum mean square error under the current variable feature. Divide the current sample set into two subsets according to the pair (j, s) of the optimal splitting point, and continue to find the splitting points of each subset according to the principle of minimum mean square error to conduct iterative region division to form multiple subsets.

[0068] Furthermore, obtain the average of the day-ahead real-time price differences of each conditional data within each subset, determine the conditional data with the minimum mean square error between the day-ahead real-time price difference within the current subset and the average of the day-ahead real-time price differences, and take the day-ahead real-time price difference of this conditional data as the target value of the current subset. Obtain the average of the target values of each subset within the current sample set, take it as the output value of the current sample set, and take the average of the output values of each sample set as the output value of the price difference prediction model, which is the day-ahead real-time price difference of the prediction day adjusted by the price difference prediction model.

[0069] Step 3: Declare the day-ahead price of the prediction day according to the above-mentioned adjusted day-ahead real-time price difference of the prediction day. Among them, according to the above-mentioned day-ahead real-time price difference and the predicted real-time price provided by the market operation agency, the predicted values and change trends of the day-ahead price in each trading period can be determined, and the day-ahead price declaration is carried out accordingly.

[0070] In this embodiment, multiple subsets are divided according to the random forest algorithm and the least mean square algorithm, and the average value of the subset output values is calculated as the output value of the current decision tree. Then, the average value of multiple decision trees is used to determine the final output result. Compared with the traditional single decision tree model, the random forest algorithm using multiple decision trees can capture the linear relationship in the data, avoid the underfitting situation that may occur when the single decision tree model processes complex problems, and thus improve the accuracy of the prediction of the day-ahead real-time price difference.

[0071] Meanwhile, through random sample set selection and random feature selection, the risk of overfitting is further reduced, and it has better generalization ability. In addition, by combining weather data from multiple sources, the dependence on a single meteorological source is eliminated, the uncertainty of weather condition data is introduced, and the training data set is enriched to improve the accuracy of the prediction of the day-ahead real-time price difference.

[0072] Please refer to Figure 3 , this application also provides a prediction system for electricity price difference, and the system includes:

[0073] A data processing unit 100, configured to obtain a data set of a to-be-measured day, where the data set includes multi-category boundary condition data, multi-category weather condition data, and day-ahead real-time price difference data, the multi-category weather condition data is the ratio of the weather condition difference to a reference value, and the day-ahead real-time price difference data is the difference between the day-ahead price and the real-time price.

[0074] A model prediction unit 200, configured to send the data set to a price difference prediction model, the price difference prediction model randomly extracts the data in the data set to determine multiple sample sets, each sample set includes multiple condition data, and the type of the condition data is one of multi-category boundary condition data, multi-category weather condition data, and day-ahead real-time price difference data. Traverse the multiple sample sets, and iteratively divide the current sample set according to the condition data in the sample set to form multiple subsets of the sample set, obtain the target values of the respective subsets, and determine the prediction result according to the target values of the respective subsets.

[0075] Among them,

[0076] The data processing unit 100 is specifically configured to obtain a data set of a to-be-measured day, where the data set includes multi-category boundary condition data, multi-category weather condition data, and day-ahead real-time price difference data, the multi-category weather condition data is the ratio of the difference between the maximum and minimum values of the weather condition to a reference value, and the day-ahead real-time price difference data is the difference between the day-ahead price and the real-time price.

[0077] Further, obtain the predicted condition data of the day to be measured. The predicted condition data includes boundary condition prediction data, weather condition prediction data, and real-time price difference prediction data before the day. Traverse each data category in the boundary condition prediction data, and sample the boundary condition prediction data of the current data category every 15 minutes. Based on the sampling results, determine the average value of the adjacent four sampling values as the boundary condition data of a corresponding data category. Traverse each data category in the weather condition prediction data, and sample the weather condition prediction data of the current data category every 60 minutes. Based on the sampling results, determine the sampling value as the weather condition data of a corresponding data category. Sample the real-time price difference prediction data before the day every 15 minutes, and based on the sampling results, determine the average value of the adjacent four sampling values as a real-time price difference data before the day.

[0078] The model prediction unit 200 is specifically configured to traverse the multiple sample sets, and iteratively divide the current sample set according to the condition data in the sample set to form multiple subsets of the sample set. Obtain the target values of the respective subsets, calculate the average value of the target values of the respective subsets, use the average value of the target values of the respective subsets as the output value of the current sample set, calculate the average value of the output values of the respective sample sets, and use the average value of the output values of the respective sample sets as the predicted price difference.

[0079] The further function descriptions of the above respective modules and units are the same as those in the corresponding above embodiments, and will not be elaborated herein.

[0080] The prediction system for electricity price difference in the embodiments of the present application is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, or other devices that can provide the above functions.

[0081] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of a computer device provided by the embodiments of the present application. As Figure 4As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if needed, multiple processors and / or multiple buses can be used together with multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (such as an array of servers, a set of blade servers, or a multi-processor system). Figure 4 In the figure, one processor 10 is taken as an example.

[0082] The processor 10 can be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 can further include a hardware chip. The above hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above programmable logic device can be a complex programmable logic device, a field programmable gate array, a generic array logic, or any combination thereof.

[0083] Among them, the memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiments.

[0084] The memory 20 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device. In addition, the memory 20 can include a high-speed random access memory and can also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 can optionally include a memory remotely set relative to the processor 10, and these remote memories can be connected to the computer device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0085] The memory 20 can include a volatile memory, such as a random access memory; the memory can also include a non-volatile memory, such as a flash memory, a hard disk, or a solid-state drive; the memory 20 can also include a combination of the above types of memories.

[0086] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or communication networks.

[0087] The embodiments of the present application also provide a computer-readable storage medium. The methods according to the embodiments of the present application can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code that is originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the methods described herein can be stored as such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the methods shown in the above embodiments are implemented.

[0088] The embodiments of the present application provide a computer program product, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods of any embodiment of the present application.

[0089] The systems or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0090] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0091] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0092] This application is described with reference to the flowcharts and / or block diagrams of methods, systems, and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0093] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0094] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0095] It should also be noted that the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such a process, method, commodity, or device. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, commodity, or device including the said element.

[0096] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment.

[0097] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

[0098] Although the embodiments of the present application have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present application, and such modifications and variations fall within the scope defined by the appended claims.

Claims

1. A method for predicting electricity price difference, characterized in that: The method comprises: Acquire a data set of a test day, the data set including multi-category boundary condition data, multi-category weather condition data and day-ahead real-time price difference data, the multi-category weather condition data is a ratio of a difference between a maximum weather condition value and a reference value, and the day-ahead real-time price difference data is a difference between a day-ahead price and a real-time price; The data set is sent to a price difference prediction model, and the price difference prediction model randomly extracts data from the data set to determine a plurality of sample sets, wherein the sample sets include a plurality of condition data, and the type of the condition data is one of multi-category boundary condition data, multi-category weather condition data, and day-ahead real-time price difference data; The multiple sample sets are traversed, and the current sample set is iteratively divided according to the conditional data in the sample set to form multiple subsets of the sample set, a target value of each subset is obtained, and a prediction result is determined according to the target value of each subset.

2. The method according to claim 1, characterized in that The data categories of the multi-category boundary condition data include electric load data, wind power data, photovoltaic power data and tie line data, and the data categories of the multi-category weather condition data include temperature data, light intensity data, humidity data and wind speed data; Obtain a data set for the day to be tested, the data set including boundary condition data, weather condition data and day-ahead real-time price difference data, specifically including: Obtain forecast condition data for the day to be tested, the forecast condition data including boundary condition forecast data, weather condition forecast data and day-ahead real-time price difference forecast data, the data categories of the boundary condition forecast data including electric load forecast data, wind power forecast data, photovoltaic power forecast data and tie line forecast data, the data categories of the weather condition forecast data including temperature forecast data, light intensity forecast data, humidity forecast data and wind speed forecast data; Traversing each data category in the boundary condition prediction data, sampling the boundary condition prediction data of the current data category once every 15 minutes, and determining the average value of four adjacent sampling values ​​as the boundary condition data of a corresponding data category based on the sampling result; Traversing each data category in the weather condition forecast data, and sampling the weather condition forecast data of the current data category once every 60 minutes, and determining the sampled value as weather condition data of a corresponding data category based on the sampling result; The day-ahead real-time price difference forecast data is sampled once every 15 minutes, and based on the sampling result, the average value of four adjacent sampling values ​​is determined as a day-ahead real-time price difference data.

3. The method according to claim 1, characterized in that Traversing the multiple sample sets and iteratively segmenting the current sample set according to the condition data in the sample set to form multiple subsets of the sample set includes: Divide the current sample set into two regions, obtain the minimum value of the sum of the minimum mean square errors of the target value and the predicted value of each condition data in the two regions, and determine the pair of split points (j, s) according to the minimum value of the sum of the minimum mean square errors, wherein the target value of the condition data is the day-ahead real-time price difference corresponding to each condition data, and the predicted value of the condition data is the average value of each target value in the current region; The current training sample set is re-divided into regions according to the pair of segmentation points (j, s) to form a first subset R1(j, s) and a second subset R2(j, s), wherein: R1(j,s)={x|x {j} >s}, R2(j,s)={x|x {j} ≤s}, where j is the feature variable and s is the segmentation feature; Iterative region segmentation is performed on the first subset and the second subset to form a plurality of subsets.

4. The method according to claim 3, characterized in that Obtaining the minimum value of the sum of the minimum mean square error of the target value and the predicted value of each condition data of the two regions includes: The minimum value of the sum of the minimum mean square errors is x i is the feature vector, y i is the target value of the i-th conditional data, c1 is the predicted value of the first region, and c2 is the predicted value of the second region.

5. The method according to any one of claims 1, 3 or 4, characterized in that: Obtaining target values ​​for each subset and determining prediction results according to the target values ​​for each subset includes: Obtaining target values ​​of the respective subsets, calculating an average value of the target values ​​of the respective subsets, and taking the average value of the target values ​​of the respective subsets as the output value of the current sample set; An average value of the output values ​​of the sample sets is calculated, and the average value of the output values ​​of the sample sets is used as the predicted price difference.

6. The method according to claim 1, characterized in that The training method of the price difference prediction model includes: Acquire a historical data set and perform random sampling with replacement from the historical data set to generate multiple training sample sets, wherein the historical data set includes boundary condition historical data, weather condition historical data, and day-ahead real-time price difference historical data; For any training sample set, the training sample set is iteratively divided until a preset stop condition is met to form multiple sub-areas, and a corresponding decision tree is generated for each training sample set; A plurality of target values ​​of each sub-region is obtained, an output value of the current decision tree is determined according to the plurality of target values ​​of each sub-region, and a final value of the output is determined according to the output value of each decision tree.

7. The method according to claim 6, characterized in that Acquiring multiple target values ​​of each sub-region, taking an average value of the multiple target values ​​of each sub-region as an output value of a decision tree, and determining a final output value according to the output values ​​of each decision tree includes: Obtain multiple target values ​​for each sub-region, calculate the average value of the multiple target values ​​for each sub-region, and use the average value of the multiple sub-regions as the output value of the current decision tree, wherein the decision tree The output value of the current decision tree I is the identity matrix, M is the number of sub-regions, N m is the number of samples in the current sub-region, R m (j,s) is the mth sub-region; Calculate the average value of the output values ​​of each decision tree, and use the average value of the output values ​​of each decision tree as the final value of the output. k is the number of decision trees.

8. The method according to claim 6, characterized in that The preset stop conditions include: The number of decision trees reaches a first preset value, the depth of each decision tree reaches a second preset value, and the number of samples contained in a leaf node of each decision tree is less than or equal to a third preset value.

9. The method according to claim 6 or 8, characterized in that: After the price difference prediction model is trained, the price difference prediction model is verified. The method includes: Acquire a validation set including boundary condition historical data, weather condition historical data, and day-ahead real-time price difference historical data, input the validation set into the price difference prediction model, and the price difference prediction model outputs a prediction result; Calculate the mean square error between the prediction result and the previous real-time price difference historical data. Where n is the number of samples, y i is the day-ahead real-time price difference of the i-th sample, y i ^ is the prediction result of the i-th sample; The preset stop condition of the price difference prediction model is dynamically adjusted according to the size of the mean square error.

10. A system for predicting electricity price difference, characterized in that: The system comprises: A data processing unit, used to obtain a data set of a test day, wherein the data set includes multi-category boundary condition data, multi-category weather condition data and day-ahead real-time price difference data, wherein the multi-category weather condition data is a ratio of a weather condition difference value to a reference value, and the day-ahead real-time price difference data is a difference between a day-ahead price and a real-time price; A model prediction unit is used to send the data set to a price difference prediction model, the price difference prediction model randomly extracts data from the data set to determine multiple sample sets, the sample sets include multiple condition data, the type of the condition data is one of multi-category boundary condition data, multi-category weather condition data and day-ahead real-time price difference data, traverse the multiple sample sets, and iteratively divide the current sample set according to the condition data in the sample set to form multiple subsets of the sample set, obtain target values ​​of each subset, and determine prediction results according to the target values ​​of each subset.