Strategy optimization method and system based on medium and long term contract and spot coupling, and medium
By constructing a coupled model based on multi-source heterogeneous data and a stochastic two-level game model, a medium- and long-term contract with dynamically adjusted options is generated, which solves the problem of the disconnect between medium- and long-term contracts and the spot market, realizes the intelligence of contracts and the endogenization of risks, and improves the overall efficiency and economy of scheduling plans.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-13
- Publication Date
- 2026-03-27
AI Technical Summary
In existing technologies, the disconnect between medium- and long-term contracts and spot market trading mechanisms leads to low contract fulfillment rates, day-ahead scheduling is prone to local optima and overall inefficiency, lacks effective dynamic adjustment means, cannot accurately match supply and demand relationships, and increases scheduling costs and risks.
Based on multi-source heterogeneous data, a coupled model and a stochastic two-level game model are constructed to generate medium- and long-term contracts with dynamically adjusted options. By combining agent user load forecasts and day-ahead market clearing results, a rolling optimization scheduling plan is performed. A two-factor dynamic weight decomposition model is used for settlement, and the strategy is updated through a hierarchical reinforcement learning framework.
It has enabled the intelligentization and risk endogenization of medium- and long-term contracts, enhanced the flexibility and risk resistance of contracts, improved the overall efficiency and economy of day-ahead scheduling plans, reduced scheduling costs and risks, and promoted market stability.
Smart Images

Figure CN121745989A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of power market transaction, in particular to a strategy optimization method and system based on coupling of medium and long-term contracts and spot market and a medium. BACKGROUND
[0002] Under the background of deepening of power market reform, the transaction mechanism of medium and long-term contracts and spot market in parallel has become the core architecture of the modern power market system. Among them, the medium and long-term contract aims to stabilize the supply and demand relationship, lock the basic price and avoid market risk; while the spot market realizes the accurate, efficient and economic dispatch of power resources through competitive clearing in short periods (such as day-ahead, real-time). However, in the current market practice, there are serious disconnections and lack of effective linkage between the two transaction mechanisms, which makes it difficult for market participants (such as power generation enterprises, power selling companies and large users) to develop a globally optimal transaction and operation strategy.
[0003] The core defect of the prior art scheme is that the signing of the medium and long-term contract and the operation and dispatch of the spot market are separated into two independent decision-making processes. Specifically, (1) the traditional medium and long-term contract is mostly a rigid power price agreement, which fails to fully embed the risk hedging mechanism for future spot market price fluctuations and new energy output uncertainty. The price prediction model relied on when signing the contract lacks sufficient accuracy and fails to quantify uncertainty (i.e. probability distribution), resulting in the contract itself being unable to adapt to subsequent drastic changes in the spot market; (2) the initial day-ahead dispatch plan developed based on the above rigid medium and long-term contract only considers the physical load balancing, and fails to optimize the financial attributes of the medium and long-term contract and the real-time price signal of the spot market in coordination. After the day-ahead market clearing result is announced, the market participants usually can only passively execute the original plan or make limited and non-systematic adjustments. This linear mode of "signing a contract first and then dispatching" makes the final dispatch scheme neither effectively utilize the medium and long-term contract to avoid risks nor flexibly capture arbitrage opportunities in the spot market, often resulting in: the actual dispatch power and the medium and long-term contract decomposition power are seriously mismatched, resulting in high deviation assessment fees, the sum of physical operation costs (such as unit start-stop and fuel consumption) and financial settlement costs (such as high-price purchase or low-price sale) is not globally optimal, in extreme price events, there is a lack of effective dynamic adjustment means, resulting in financial losses beyond expectations.
[0004] The existing dispatch plan is only based on short-term marginal cost to develop a plan, which may result in waste of resources such as curtailment of medium and long-term dimension, does not combine the price lock clause in the contract, increases the risk of cost fluctuations in dispatch execution, and in the absence of guidance from medium and long-term contracts, day-ahead dispatch is prone to "local optimum, global inefficiency".
[0005] In addition, in the electricity market, multiple stakeholders, including power generation companies, users, and grid companies, establish their rights and obligations through medium- and long-term contracts. Existing dispatch plans are unable to take into account the demands of all parties and cannot accurately match the supply and demand relationship stipulated in the contracts. This leads to situations such as over-generation / under-generation by some stakeholders and inability to guarantee electricity demand. The lack of a conflict coordination mechanism under contractual constraints reduces the dispatch stability of the entire power system.
[0006] Therefore, existing technologies have failed to establish a closed-loop mechanism of "contract-scheduling-settlement-feedback." The fundamental reason lies in the lack of a systematic method capable of integrating multi-source heterogeneous data for high-precision, probabilistic price forecasting, and on this basis, generating intelligent medium- and long-term contracts with dynamic adjustment capabilities, thereby driving the coordinated optimization of physical scheduling and financial settlement. If this problem remains unresolved, market participants will continue to struggle between rigid contractual constraints and the ever-changing spot market, making it difficult to achieve the dual goals of economic benefits and risk control. Summary of the Invention
[0007] To address the problems of low contract fulfillment rates due to insufficient consideration of electricity allocation requirements in existing technologies, and the tendency for day-ahead scheduling to fall into local optima and global inefficiency due to the lack of guidance from medium- and long-term contracts, this invention proposes a scheduling plan determination method based on medium- and long-term contracts, comprising: Medium- to long-term contracts with dynamically adjusted options are generated based on multi-source heterogeneous data combined with a pre-built coupled model and a stochastic two-level game model. Based on historical data of agent users, predict the short-term load of agent users and formulate an initial scheduling plan in combination with medium and long-term contracts; obtain the day-ahead market clearing results, and based on the clearing results, the dynamic adjustment options in the medium and long-term contracts, and the physical-financial collaborative optimization objectives, perform rolling optimization on the initial scheduling plan to obtain the final scheduling plan.
[0008] Preferably, the step of performing rolling optimization on the initial scheduling plan based on the clearing results, the dynamic adjustment options in medium- and long-term contracts, and the physical-financial collaborative optimization objective to obtain the final scheduling plan includes: The total cost is calculated based on physical operating costs, the electricity volume allocated from medium- and long-term contracts, the actual dispatched electricity volume in the clearing results, and the expected spot settlement price difference. With the goal of minimizing the total cost, the initial scheduling plan is continuously optimized under the constraints to obtain the final scheduling plan.
[0009] Preferably, the generation of medium- to long-term contracts containing dynamically adjusted options based on multi-source heterogeneous data combined with a pre-built coupled model and a stochastic two-level game model includes: By substituting multi-source heterogeneous data into a pre-built spot price prediction model, the probability distribution of spot market prices is obtained. Based on the spot market price probability distribution, the historical output curve of the new energy power generator, the historical power consumption curve of the market user, the installed capacity of the new energy power generator and the market rules, a long-term contract containing a dynamic adjustment option is generated by combining a pre-constructed stochastic bi-level game model; The pre-constructed spot price prediction model is a spatio-temporal graph constructed based on multiple meteorological stations, new energy power generation nodes and key load centers in the province, combined with a graph convolution network or a graph attention network and a long short-term memory network.
[0010] Preferably, the multi-source heterogeneous data is substituted into the pre-constructed spot price prediction model to obtain a spot market price probability distribution, which comprises: The multiple meteorological stations, new energy power generation nodes and key load centers in the province are regarded as nodes in the graph, and the geographical distance or the power grid topological relationship is regarded as the edge to construct a spatio-temporal graph. Based on the spatio-temporal graph, the information of the neighbor nodes is aggregated at each time step by using a graph convolution network or a graph attention network to extract a spatial feature vector; The spatial feature vector is spliced, and the spliced spatial feature vector is input into a long short-term memory network to capture the long-term trend and periodic pattern in the time series data of each node, thereby obtaining a high-dimensional feature vector fused with spatial correlation and time evolution; The high-dimensional feature vector is generated into a spot market price probability distribution by a generative model.
[0011] Preferably, the long-term contract containing a dynamic adjustment option is generated by combining a pre-constructed stochastic bi-level game model based on the spot market price probability distribution, the historical output curve of the new energy power generator, the historical power consumption curve of the market user, the installed capacity of the new energy power generator and the market rules, which comprises: The stochastic bi-level game model is converted into a single-layer linear model and solved by using KKT condition and strong duality theory, thereby obtaining two groups of candidate contract curves containing electricity price and power; The weights of the two groups of candidate contract curves are determined by using an entropy weight method, and the basic contract electricity price and power curve are calculated according to the weights; The dynamic adjustment option is embedded in the basic contract electricity price to form a long-term contract containing a dynamic adjustment option.
[0012] Further, the construction of the spot price prediction model comprises: A spatio-temporal graph is constructed based on multiple meteorological stations, new energy power generation nodes and key load centers in the province by using a graph neural network; The spot price prediction model is constructed from the spatio-temporal graph combined with a graph convolution network or a graph attention network and a long short-term memory network.
[0013] Further, based on multi-source heterogeneous data, a pre-constructed coupling model and a random double-layer game model are combined to generate a medium and long-term contract containing a dynamic adjustment option, and the medium and long-term contract further includes: Calculate the real-time convergence degree between the provincial new energy output curve, the social electricity load curve and the historical spot price curve; Based on the real-time convergence degree, an online learning algorithm is used to dynamically adjust the contribution weight of new energy output, electricity load and weather information to the spot price prediction model; According to the dynamically adjusted weight and the identified market state, the spot market price probability distribution is nonlinearly calibrated.
[0014] In another aspect, the application also provides a strategy optimization method based on medium and long-term contract and spot coupling, comprising: Based on the final dispatching plan and the actual spot market price, a double-factor dynamic weight decomposition model is used to settle the medium and long-term contract, and a risk report is generated; Based on the risk report and the hierarchical reinforcement learning framework, the medium and long-term contract strategy is updated; The final dispatching plan is determined by the dispatching plan determination method based on the medium and long-term contract.
[0015] Preferably, the method for settling the medium and long-term contract based on the final dispatching plan and the actual spot market price, and generating a risk report, comprises: Based on the characteristics of the medium and long-term contract subject, the load or output proportion of the kth hour is calculated based on the historical load or output curve as the basic weight of the kth hour, wherein k is the time value; Based on the deviation of the actual spot market price of the kth hour from the daily average price, the spot price correction weight of the kth hour is calculated; Based on the basic weight of the kth hour and the spot price correction weight of the kth hour, the total decomposition weight of the kth hour is calculated by combining the respective coefficients; Based on the total decomposition weight of the kth hour multiplied by the total contract power, the settlement power of the kth hour is calculated; Based on the settlement power of the kth hour, the difference settlement amount and the risk exposure are calculated; According to the market subject's medium and long-term power purchase contract, the spot price prediction model and the risk exposure, the risk-reward Pareto optimal frontier is solved; Based on the preset risk preference level, the optimal solution is selected from the risk-reward Pareto optimal frontier to generate a risk report containing strategy suggestions.
[0016] Preferably, the difference settlement amount = Σ (kth hour contract electricity price × settlement electricity quantity) - Σ (kth hour actual spot market price × actual deviation electricity quantity) (if the electricity purchasing party, the deviation electricity quantity is negative, indicating that the high price needs to be purchased, increasing the cost).
[0017] Preferably, updating the long-term contract strategy based on the risk report and the hierarchical reinforcement learning framework, including: Encoding the risk report, long-term market trend information and historical contract execution data into a high-dimensional state vector; Outputting specific actions by the policy network of the high-level agent in the hierarchical reinforcement learning framework according to the state vector, the actions including adjusting the matching weight of the trading counterparty, modifying the optimization boundary of the contract quantity and price, or generating an option purchase instruction; Ensuring that the learning goal of the agent is consistent with the long-term business goal of the market subject through the reward function of the high-level agent; the reward function includes the expected return of the medium and long-term contract, the risk exposure reduction amount and the contract fulfillment rate improvement amount; Receiving the spot market price probability distribution, market state and medium and long-term holding constraints issued by the high-level agent by the bottom-level agent, and outputting specific day-ahead market bidding strategies and real-time market pricing strategies according to the current market state and holding constraints, combined with the reward function of the bottom-level agent, wherein the reward function of the bottom-level agent includes the immediate return of spot market transactions, deviation assessment fees and the contribution degree of medium and long-term contract risk hedging; Calculating a meta-reward signal for the degree of achievement of the long-term goal of the high-level agent; The high-level agent and the bottom-level agent train through sharing the experience replay pool, and coordinate strategies through the meta-reward signal.
[0018] In another aspect, the application also provides a strategy optimization system based on the coupling of medium and long-term contracts and spot markets, including: A report generation module for settling the medium and long-term contracts based on the final dispatch plan combined with the actual spot market price, using a double-factor dynamic weight decomposition model, and generating a risk report; A strategy updating module for updating the medium and long-term contract strategy based on the risk report and the hierarchical reinforcement learning framework; Wherein, the final dispatch plan is determined by using the above-mentioned dispatch plan determination method based on the medium and long-term contract.
[0019] Preferably, the report generation module is specifically used for: Based on the characteristics of the medium and long-term contract subject, using the historical load or output curve to calculate the load or output proportion of the kth hour as the basic weight of the kth hour, wherein k is the time value; The spot price adjustment weight for the k-th hour is calculated based on the deviation between the actual spot market price and the daily average price in the k-th hour. The total decomposition weight of the k-th hour is calculated by combining the base weight of the k-th hour and the spot price adjustment weight of the k-th hour with their respective coefficients. The settlement electricity volume for the k-th hour is calculated by multiplying the total decomposition weight of the k-th hour by the total contract electricity volume. Calculate the price difference settlement amount and risk exposure based on the settlement electricity volume in the k-th hour; Based on the market participants' medium- and long-term power purchase contracts, spot price forecasting models, and risk exposures, the risk-return Pareto optimal frontier is solved. Based on a preset risk preference level, the optimal solution is selected from the risk-return Pareto optimal frontier to generate a risk report containing strategy recommendations.
[0020] Preferably, the strategy update module is specifically used for: The risk report, long-term market trend information, and historical contract execution data are encoded into a high-dimensional state vector; The policy network of the high-level agent in the hierarchical reinforcement learning framework outputs specific actions based on the state vector. These actions include adjusting the matching weights of the trading counterparty, modifying the optimization boundary of the contract quantity and price, or generating option purchase instructions. The reward function of the high-level intelligent agent ensures that the learning objectives of the intelligent agent are consistent with the long-term business objectives of the market entity; the reward function includes the expected return of medium and long-term contracts, the amount of reduction in risk exposure, and the amount of improvement in contract performance rate. The underlying intelligent agent receives the probability distribution of spot market prices, market status, and medium- to long-term position constraints issued by the higher-level intelligent agent. Based on the current market status and position constraints, and combined with the reward function of the underlying intelligent agent, it outputs specific daily market volume reporting strategies and real-time market quotation strategies. The reward function of the underlying intelligent agent includes the immediate profit of spot market transactions, deviation assessment fees, and contribution to medium- to long-term contract risk hedging. Calculate the meta-reward signal for the achievement of the long-term goals of the high-level intelligent agent; High-level agents and low-level agents are trained by sharing an experience replay pool and coordinate policies through meta-reward signals.
[0021] In another aspect, this application also provides an electronic device, comprising: at least one processor and a memory; the memory and the processor are connected via a bus; The memory is used to store one or more programs; When the one or more programs are executed by the at least one processor, a strategy method based on medium- and long-term contracts and spot market linkage transactions as described above is implemented.
[0022] In still another aspect, the application also provides a computer readable storage medium, which has an execution program stored thereon, and the execution program, when executed, implements a strategy method based on linkage transaction of medium and long-term contract and spot.
[0023] Compared with the prior art, the application has the following beneficial effects: A scheduling plan determination method based on medium and long-term contract, comprising: generating a medium and long-term contract containing dynamic adjustment options based on multi-source heterogeneous data in combination with a pre-constructed coupling model and a stochastic double-layer game model; predicting the short-term load of an agent user based on historical data of the agent user, and formulating an initial scheduling plan in combination with the medium and long-term contract; obtaining a day-ahead market clearing result, and performing rolling optimization on the initial scheduling plan based on the clearing result, the dynamic adjustment options in the medium and long-term contract, and a physical-financial synergistic optimization target to obtain a final scheduling plan. The application fully considers future market uncertainty, realizes the intellectualization and risk endogenization of the medium and long-term contract signing process, significantly enhances the contract flexibility and risk resistance, and also solves the problems of disconnection between contract constraints and day-ahead scheduling, and imbalance between the economy and feasibility of the day-ahead scheduling plan. BRIEF DESCRIPTION OF DRAWINGS
[0024] Figure 1 A scheduling plan determination method based on medium and long-term contract according to the application is shown in the flowchart. Figure 2 A strategy optimization method based on coupling of medium and long-term contract and spot according to the application is shown in the flowchart. Figure 3 A strategy method based on linkage transaction of medium and long-term contract and spot according to the embodiment of the application is shown in the flowchart. Figure 4 An electronic device structure according to the application is shown in the schematic diagram. DETAILED DESCRIPTION
[0025] In order to better understand the application, the content of the application is further described below in combination with the drawings and examples of the specification.
[0026] Example 1: A scheduling plan determination method based on medium and long-term contract, as shown in Figure 1 , comprising: generating a medium and long-term contract containing dynamic adjustment options based on multi-source heterogeneous data in combination with a pre-constructed coupling model and a stochastic double-layer game model; Based on historical data of the agent user, a short-term load of the agent user is predicted, and an initial scheduling plan is formulated in combination with a medium and long-term contract; a day-ahead market clearing result is obtained, and based on the clearing result, a dynamic adjustment option in the medium and long-term contract, and a physical-financial collaborative optimization target, the initial scheduling plan is subjected to rolling optimization to obtain a final scheduling plan.
[0027] The present application provides a scheduling plan determination method based on a medium and long-term contract, which comprises the following steps: S1. Based on a spot price prediction scenario generated by a coupling model fusing provincial day-ahead new energy output, provincial day-ahead social electricity load and real-time regional meteorological information, a medium and long-term contract containing a dynamic adjustment option is generated by a random double-layer game model optimization; S2. Based on the medium and long-term contract and load prediction, an initial day-ahead scheduling plan is formulated; S3. A day-ahead market clearing result is obtained, and based on the clearing result, the dynamic adjustment option in the medium and long-term contract and the physical-financial collaborative optimization target, the initial day-ahead scheduling plan is subjected to rolling re-optimization to obtain a final scheduling plan; Preferably, the generation of the spot price prediction scenario in the S1 step further comprises: S1.1. Input provincial day-ahead new energy output, provincial day-ahead social electricity load and real-time regional meteorological information; S1.2. Build a spatio-temporal graph neural network and long short-term memory network coupling model to process the input multi-source heterogeneous data; specifically, the spatio-temporal graph neural network (ST-GNN) part: multiple meteorological stations, new energy generation nodes and key load centers in the province are regarded as "nodes" in the graph, and geographical distance or power grid topological relationship is regarded as "edge". Through a graph convolution network (GCN) or a graph attention network (GAT), the information of the neighbor nodes is aggregated at each time step to extract spatial dependence features (such as the influence of upstream wind speed on downstream wind power output).
[0028] The long short-term memory network (LSTM) part: processes the time series data of each node (such as the hourly output of a certain wind farm and the hourly load of a certain region), and uses its gating mechanism to capture long-term trends and periodic patterns.
[0029] Coupling mode: after the spatial feature vectors extracted by the ST-GNN at each time step are spliced, they are taken as the input sequence of the LSTM. The LSTM further captures the dynamic rules of the evolution of these spatial features over time. The final output is a high-dimensional feature vector that fuses "spatial correlation" and "temporal evolution".
[0030] S1.3. Based on the output of the coupling model, use the generative model to calculate the spot market price probability distribution of the future time period. Specifically, the generative model refers to a class of machine learning models that can learn the data distribution and generate new samples. The conditional diffusion model or conditional variational autoencoder mentioned in this paper belongs to this category. They take the feature vector output by the coupling model as a condition, start from a simple noise distribution (such as a standard normal distribution), and gradually generate future price samples that conform to historical price rules.
[0031] Relationship with the coupling model: The coupling model is responsible for extracting deep features of the input data, and the generative model is based on these features to model and generate the probability distribution of future prices. The two are a two-stage model that connects the front and back: the coupling model is a feature extractor, and the generative model is a distribution generator The preferred solution of the present application further comprises: S1.4. Calculate the real-time convergence degree between the provincial new energy output curve, the social electricity load curve, and the historical spot price curve; S1.5. Based on the real-time convergence degree, use an online learning algorithm to dynamically adjust the contribution weight of new energy output, electricity load, and weather information to the spot price prediction model; S1.6. According to the dynamically adjusted weight and the identified market state, nonlinearly calibrate the spot market price probability distribution obtained in S1.3. Specifically, the state recognition module built into the system (which can be based on clustering algorithms, hidden Markov models, or deep classification networks) automatically identifies, and makes a comprehensive judgment based on the real-time convergence degree calculated in S1.4-S1.5, the dynamic weight, and the feature vector output by the coupling model.
[0032] Market states usually include: (1) High volatility-low supply (new energy output drops, load is high); (2) Stable-high demand (temperature drops sharply, heating load rises); (3) New energy generation-low load (sufficient light + weekend, risk of zero price); (4) Cold wave warning (strong cold air southward, expected load peak); (5) Regular operation (supply and demand balance, stable price).
[0033] The preferred solution of the present application further comprises: S1.7. Input the historical output curve of new energy power generators, the historical electricity curve of market users, the installed capacity of new energy power generators, market rules, and the calibrated spot market price probability distribution obtained in S1.6; S1.8. Constructing a random bi-level game model with new energy power generator and user combination as participants, wherein the objective functions of both parties of the game contain expected revenue and risk terms, and the calculation of the risk terms is based on the calibrated spot market price probability distribution; S1.9. Transforming the random bi-level game model into a single-level linear model and solving it by using KKT conditions and strong duality theory, to obtain two groups of candidate contract curves containing electricity price and power; S1.10. Determining the weights of the two groups of candidate contract curves by using entropy weight method, and obtaining the basic contract electricity price and power curve according to the weights; S1.11. Embedding a dynamic adjustment option in the basic contract to form a final medium and long-term contract, wherein the dynamic adjustment option stipulates that a certain proportion of subsequent contract electricity can be adjusted under certain spot market conditions.
[0034] Preferably, after the step S1.11, the method further comprises: S1.12. Using historical or simulated spot market data to perform stress testing on the medium and long-term contract obtained in the step S1.11, to verify the robustness of the medium and long-term contract under different market scenarios, and returning to the step S1.9 to adjust the parameters of the game model if the verification fails.
[0035] Preferably, in the step S3, the physical-financial coordinated optimization objective function is: minimizing (physical operation cost + absolute value of difference between medium and long-term contract decomposed electricity and actual dispatched electricity × expected spot settlement price difference); wherein the physical operation cost includes coal consumption cost and start-stop cost of conventional units, and the expected spot settlement price difference is determined based on the calibrated spot market price prediction obtained in the step S1.6.
[0036] Based on the obtained final dispatching plan, the application further provides a strategy optimization method based on coupling of medium and long-term contract and spot, as shown in Figure 2 , which comprises: S4. Based on the final dispatching plan and actual spot market price, using a two-factor dynamic weight decomposition model to settle the medium and long-term contract; S5. Generating a risk report based on the settlement result, and updating the medium and long-term contract strategy by using a hierarchical reinforcement learning framework with the risk report as input.
[0037] Preferably, in the step S4, the two-factor dynamic weight decomposition model comprises: S4.1. Calculating the basic weight of the kth hour: based on the characteristics of the contract subject, using its historical load or output curve to calculate the load or output proportion of the kth hour; S4.2. Calculate the spot price correction weight of the kth hour: calculate based on the deviation of the actual spot market price of the kth hour from the daily average price; S4.3. Calculate the total decomposition weight of the kth hour: the total decomposition weight is equal to the base weight multiplied by the first coefficient, plus the spot price correction weight multiplied by the second coefficient, and the sum of the first coefficient and the second coefficient is one; S4.4. Calculate the settlement electricity of the kth hour: the settlement electricity of the kth hour is equal to the total decomposition weight of the kth hour multiplied by the total electricity of the contract.
[0038] Preferably, in the scheme, the generation of the risk report in the S5 step comprises: S5.1. Based on the settlement results of the S4 step, calculate the difference settlement amount and the risk exposure; The calculation is the basis for generating the risk report and is the core indicator of the actual financial performance of the market subject, which is used to evaluate the profitability of the trading strategy.
[0039] Specifically, the difference settlement amount = Σ (kth hour contract electricity price × settlement electricity) - Σ (kth hour actual spot market price × actual deviation electricity) (if it is a power purchase party, the deviation electricity is negative, indicating that it needs to be purchased at a high price, increasing the cost).
[0040] Risk exposure: can be defined as the potential maximum loss under the most unfavorable market scenario, for example: using conditional value at risk: taking the average loss under the worst percentage of the settlement result distribution or based on the stress test results, calculating the maximum loss amount that may be generated under extreme prices.
[0041] S5.2. According to the medium and long term power purchase contract of the market subject, the spot price prediction model obtained in the S1.6 step, and the risk exposure in the S5.1 step, solve the risk-reward Pareto optimal frontier; Specifically, the "spot price prediction model" is the general term of the entire prediction system, which includes all components from S1.1 to S1.6 (data input → coupled model → generative model → dynamic calibration). And the "calibrated spot market price probability distribution" is the final output result of the model after running at a specific time. The two are the relationship between system and product.
[0042] And the spot price prediction model is an integrated hybrid prediction system, which consists of the following parts: (1) Multi-source heterogeneous data input layer; (2) Spatio-temporal graph neural network and LSTM coupled feature extraction layer; (3) Probability distribution generation layer based on generative model (such as diffusion model); (4) Online learning driven dynamic weight adjustment and non-linear calibration layer, which outputs the spot price probability distribution of multiple time periods in the future to support risk-aware decision-making.
[0043] S5.3. Based on the preset risk preference level, an optimal solution is selected from the Pareto optimal frontier to generate a risk report containing strategy suggestions.
[0044] Preferably, in the S5 step, updating the medium and long-term contract strategy through the hierarchical reinforcement learning framework includes: being composed of a high-level agent and a bottom-level agent, both of which are trained cooperatively through a shared experience replay pool and a meta-reward signal. The high-level agent is the top decision-making unit of the framework, responsible for formulating long-term strategies. It includes a high-level and a bottom-level; the high-level is used to receive macro states (risk reports, trends, etc.), and output strategic actions (such as adjusting counterparty weights, modifying optimization boundaries, and purchasing options); the bottom-level is used to receive high-level instructions and market states, and execute specific tactics (day-ahead reporting, real-time pricing), and manage spot market operations.
[0045] S5.4. Receiving the risk report generated in the S5.3 step, the risk report contains strategy suggestions generated based on the risk-reward Pareto optimal frontier; S5.5. Encoding the risk report, long-term market trend information, and historical contract execution data into a state vector; S5.6. The strategy network of the high-level agent outputs actions according to the state vector, the actions include adjusting the matching weight of the trading counterparty, modifying the optimization boundary of the contract quantity and price, or generating option purchase instructions. The high-level agent is a strategic decision-making module in the hierarchical reinforcement learning framework, usually a deep neural network; S5.7. The reward function of the high-level agent includes the expected return of the medium and long-term contract, the amount of risk exposure reduction, and the amount of contract fulfillment rate improvement. Specifically, the reward function of the high-level agent includes the expected return, the amount of risk exposure reduction, and the amount of contract fulfillment rate improvement. Its role is to guide the high-level agent to learn how to formulate the optimal long-term strategy, and ensure that its actions can effectively improve the overall performance.
[0046] The reward function of the bottom-level agent includes the immediate return, the deviation cost, and the risk hedging contribution. Its role is to encourage the bottom-level agent to make operations in the spot market that are beneficial to reducing the overall risk, not only pursuing profits, but also serving the medium and long-term position management. The two work together to continuously optimize the system in "exploration" and "utilization", and realize the closed-loop update of the strategy.
[0047] Preferably, the hierarchical reinforcement learning framework further includes: S5.8. The bottom-layer agent receives the calibrated spot market price probability distribution obtained in step S1.6, the market state, and the medium- and long-term position constraint issued by the high-layer agent; specifically, the high-layer agent generates the medium- and long-term position constraint based on the following factors: (1) Strategy suggestions in the risk report (such as "suggest increasing long position in a certain period"); (2) Preset risk preference level (conservative users limit short position size); (3) Long-term market trend (expecting price rise, then suggest establishing long position); (4) Historical contract execution bias (if there is often power shortage in a certain period, then increase position in advance). This constraint is a "hard requirement" for the bottom-layer agent, ensuring that spot trading serves the overall risk management goal.
[0048] The basis for the high-layer agent to generate the medium- and long-term position constraint is the comprehensive analysis of multi-dimensional information; First, the risk report provides key indicators such as current contract execution risk exposure, margin settlement amount, and optimal risk-reward balance point, which allows the high-layer agent to determine the future risk management direction. Second, the market participants' preset risk preference level, such as conservative, balanced, or aggressive, directly affects the tightness of the position constraint. For example, conservative users will strictly control short positions, while aggressive users will allow large risk exposure in certain periods to seek high returns. In addition, long-term market trend information, such as seasonal electricity price trends and new energy output prediction trends, is also used to judge future supply and demand patterns, thereby laying out the position structure in advance. Meanwhile, the system also references historical contract execution data to analyze past deviation power, deviation assessment fees, etc., to identify periods where problems frequently occur and set stronger position safeguard requirements in these periods. Finally, spot market price prediction results and market state identification results also affect constraint generation, for example, in periods where extreme high prices are predicted, the system will require sufficient long positions to be established to hedge risks. By synthesizing the above factors, the high-layer agent outputs specific position requirements through its strategy network and converts them into executable mathematical constraint conditions, such as the net position in a certain period not being lower than a certain value, which is then issued to the bottom-layer agent for execution.
[0049] S5.9. The strategy network of the bottom-layer agent outputs day-ahead market bidding strategies and real-time market pricing strategies based on the current market state and the medium- and long-term position constraint; S5.10. The reward function of the bottom-layer agent includes immediate income from spot market trading, deviation assessment fees, and contribution to medium- and long-term contract risk hedging; S5.11. The high-level agent and the bottom-level agent are trained through sharing an experience replay pool, and strategy coordination is achieved through a meta-reward signal, which is calculated based on the degree of achievement of the long-term goal of the high-level agent by the execution result of the bottom-level agent.
[0050] Specifically, at the end of each trading day, the system packages the complete decision-making and execution process of the day into a number of experience units and writes them into the shared experience replay pool. In the offline training phase, the system randomly selects a batch of experience samples from the replay pool to update the policy networks of the high-level and bottom-level agents, respectively. The training uses a deep reinforcement learning algorithm to adjust the network parameters by calculating the policy gradient, so that the agent is more inclined to choose actions that bring higher rewards in the future. Since the high-level and the bottom-level share the same replay pool, they can learn each other's decision-making logic, thereby improving the overall system's synergy. The target network mechanism is also introduced during training to improve stability, and the online running strategy model is updated regularly to ensure that the learning results can be applied to actual trading decisions in a timely manner.
[0051] The meta-reward signal is essentially a performance scale that measures whether the current contract strategy is successful, and it converts the execution experience in the spot market into direct guidance for future contract design. Through this mechanism, the system realizes a complete closed loop from "signing contracts - execution scheduling - settlement evaluation - strategy updating", truly embodying the main idea of the invention of "intelligently updating medium and long-term contract strategies".
[0052] Specifically, the training process is as follows: (1) The high-level and bottom-level agents perform actions in the environment to generate experiences (state, action, reward, next state); (2) Store the experience in the shared experience replay pool; (3) During training, a batch of experiences are randomly sampled from the pool to update the policy networks of the high-level and bottom-level agents, respectively; (4) Coordinate the goals of the two through the meta-reward signal to achieve co-evolution.
[0053] The shared experience replay pool stores: s: state vector (risk report + trend + historical data); a: agent action (such as "increase A-type counterparty weight"); r: corresponding reward (high-level: comprehensive income, bottom-level: spot performance); s': next state.
[0054] Specifically, the "execution result" refers to the actual trading performance of the bottom-level agent in the spot market, which specifically includes: (1) The bid amount and the cleared amount of the previous day's market; (2) The bid price and the transaction price of the real-time market; (3) the deviation amount of the actual physical scheduling plan from the medium and long-term contract; (4) the generated deviation assessment fee; the net income of spot market transaction. These results are aggregated to calculate their contribution to the long-term goals of the upper intelligent agent (such as annual total income and average risk level), and then generate a meta-reward signal to complete the strategy coordination.
[0055] The advantages of the present application are: 1. By fusing the provincial day-ahead new energy output, provincial day-ahead social electricity load and real-time regional meteorological information, and constructing a spatio-temporal graph neural network and long short-term memory network coupled model for processing, the system can capture the spatio-temporal correlation characteristics and time series dynamic characteristics that affect the spot market price, realize more accurate probabilistic prediction of the spot market price, and significantly improve the prediction accuracy and robustness. The design overcomes the defects of traditional single models or simple weighted models, such as difficulty in fully fusing multi-source heterogeneous data, neglecting spatial dependence, or being unable to effectively handle long-period dependence, and provides a high-quality input basis for subsequent medium and long-term contract optimization and risk control.
[0056] 2. By constructing a random double-layer game model based on the spot price prediction scenarios generated by the above coupled model to optimize the generation of medium and long-term contracts containing dynamic adjustment options, the new energy generator and market user can reach a contract scheme that takes into account both their expected income and risk preference through the game mechanism, realizing the intelligentization and risk endogenization of the medium and long-term contract signing process, and significantly enhancing the contract flexibility and risk resistance while ensuring the contract economy. The design changes the traditional medium and long-term contract "one size fits all", lack of flexibility, and difficulty in adapting to the sharp fluctuations of the spot market, and upgrades the contract from a simple physical electricity quantity agreement to an "intelligent contract" containing financial options.
[0057] 3. By embedding dynamic adjustment options in the medium and long-term contract, and after obtaining the day-ahead market clearing result, the initial day-ahead scheduling plan is rolled and re-optimized in combination with the options and physical-financial coordinated optimization objectives, so that market participants can flexibly adjust the physical operation plan according to the latest market certainty information to minimize the total cost (the sum of physical operation cost and financial deviation cost), realizing the refinement and dynamization of day-ahead scheduling decision, and improving the overall operation economy and risk hedging efficiency. The design uses the flexibility of options to closely link the medium and long-term contract with the spot market transaction, avoiding high deviation assessment fees or missing market opportunities due to rigid planning.
[0058] 4. By adopting a two-factor dynamic weight decomposition model for settlement of medium and long-term contracts, i.e. combining the basic weight based on historical electricity / output characteristics and the correction weight based on the deviation degree of the current market price on the same day, the decomposition of the contract settlement electricity quantity respects the inherent physical operation mode of the market subject and reflects the overall price signal of the market on the same day, realizes the fairness and rationality of the settlement mechanism, and achieves the effects of balancing the physical and financial properties, reducing settlement disputes, and promoting market stability. The design solves the "inferior coin drives out good coin" or incentive distortion problem that may be caused by the traditional settlement method (such as according to the contract curve or according to the actual curve).
[0059] Embodiment 2 The application will be described in further detail below with reference to the accompanying drawings.
[0060] A strategy optimization method based on the coupling of medium and long-term contracts and spot, as shown in Figure 3 , includes the following steps: S1. Based on the spot price prediction scenario generated by the coupling model of fusing provincial day-ahead new energy output, provincial day-ahead social electricity load and real-time regional meteorological information, a medium and long-term contract containing dynamic adjustment options is generated by optimizing a random double-layer game model; aiming to generate a medium and long-term contract that can reflect market game and effectively manage future risks; S1.1 Input data: First, the system collects and inputs key multi-source heterogeneous data, including: provincial day-ahead new energy output prediction (covering wind power, photovoltaic, etc.), provincial day-ahead social electricity load prediction, and real-time regional meteorological information (such as wind speed, light intensity, temperature, humidity, etc.) covering key areas in the future period (such as the next day).
[0061] S1.2 Build a coupling model: In order to accurately predict future spot prices, the application builds an advanced spatio-temporal graph neural network and long short-term memory network coupling model. The model first uses a spatio-temporal graph neural network (ST-GNN) to process data with spatial correlation. For example, multiple meteorological stations and key load nodes within a province are regarded as "nodes" in the graph, and the geographical distance or power grid topological relationship between nodes is regarded as an "edge". ST-GNN can effectively capture the spatial propagation effect of different regional meteorological conditions (such as wind speed) and its joint influence on surrounding new energy output and load. At the same time, the long short-term memory network (LSTM) is used to process the time series data of each node (such as the load of a certain key node, the output of a certain wind farm, the wind speed of a certain meteorological station) to capture its long-term trend and periodic characteristics. Finally, the deep spatio-temporal features extracted by ST-GNN and the time series features extracted by LSTM are fused as the input of the subsequent model.
[0062] S1.3 Generating Price Probability Distribution: Based on the output of the coupling model in step S1.2, the system utilizes a generative model (e.g., a diffusion-based model or a conditional generative adversarial network) to generate a probability distribution of the spot market price for multiple future time periods (e.g., 24 hours). Unlike traditional point forecasts, the probability distribution (e.g., a distribution curve for each hour's price) quantifies the uncertainty of the forecast, providing critical information for subsequent risk management. For example, the model might output that there is a 70% probability that the spot price at hour 10 will be between 300-400 yuan / MWh, a 20% probability that it will be below 300 yuan / MWh, and a 10% probability that it will be above 500 yuan / MWh.
[0063] S1.4 Calculating Real-Time Convergence: To further improve the accuracy of the forecast, the system calculates the real-time convergence between the provincial new energy output curve, the social electricity load curve, and the historical spot price curve. This can be achieved through the Dynamic Time Warping (DTW) algorithm. DTW can measure the minimum distance between two time series after nonlinear stretching or compression on the time axis, thereby quantifying their similarity in shape. For example, when the superimposed shape of the new energy output curve and the load curve is highly similar to the shape of a historical "high price" event, the system will increase the probability of a high price occurring in the future.
[0064] S1.5 Dynamically Adjusting Weights: Based on the real-time convergence calculated in step S1.4, the system uses an online learning algorithm (such as adaptive filtering or online gradient descent) to dynamically adjust the contribution weights of the new energy output, electricity load, and weather information data sources to the final price prediction model. For example, when the system detects that the current market state is highly similar to a historical "new energy output sudden drop leading to price surge" state, it will automatically increase the weight of the "new energy output" data source, making it have a greater impact on the final prediction.
[0065] S1.6 Nonlinear Calibration: Taking into account the dynamic weights obtained in step S1.5 and the market state identified by the system (such as "high volatility-low supply," "stable-high demand," etc.), the system performs nonlinear calibration on the spot market price probability distribution generated in step S1.3. For example, in the "high volatility-low supply" state, the system may amplify the upper tail (high price part) of the price distribution to more accurately reflect the risk of extreme events.
[0066] S1.7 Inputting the Game Model: The calibrated spot market price probability distribution obtained in step S1.6, along with the historical output curves of new energy generators, the historical electricity consumption curves of market users, the installed capacity of new energy generators, market rules, and other information, are input into a stochastic bi-level game model.
[0067] S1.8 Build the game model: the participants of this game model are new energy generators and market users (or their agents, power selling companies). The upper model represents one party (such as users), and the lower model represents the other party (such as power generators), and both parties aim to maximize their own interests. The objective functions of both parties contain two parts: expected revenue and risk term. The expected revenue is calculated based on the contract power and the electricity price. The risk term is directly calculated using the spot price probability distribution obtained in S1.6, for example, the conditional value at risk (CVaR) can be used to quantify the expected loss in the worst case, prompting participants to consider potential market risks in negotiations.
[0068] S1.9 Model transformation and solution: due to the difficulty of solving the bi-level game model, the present application uses the KKT (Karush-Kuhn-Tucker) condition and strong duality theory to transform the original bi-level game model into a single-level linear model. Specifically, the optimality condition (KKT condition) of the lower problem is used as the constraint of the upper problem, and the inequality constraint in the lower problem is processed using the duality theory, and finally a mixed integer linear programming (MILP) problem is formed, which can be efficiently solved by using commercial solvers (such as Gurobi, CPLEX).
[0069] S1.10 Determine the candidate contract: after solving the single-level linear model obtained in S1.9, two sets of candidate contract curves (each curve contains 24 period electricity prices and power) are obtained. These two curves represent the optimal contract solutions that both parties may accept under the current market scenario.
[0070] S1.11 Calculate the basic contract: the entropy weight method is used to determine the weight of the two sets of candidate contract curves. The entropy weight method is an objective weighting method that determines the weight of each candidate solution based on its dispersion in different indicators (such as the revenue of each party and the risk level). The greater the dispersion (the smaller the information entropy), the higher the weight. According to the calculated weight, the two sets of candidate contract curves are weighted and averaged to obtain a basic contract electricity price and power curve.
[0071] S1.12 Embed dynamic adjustment options: in the basic contract obtained in S1.11, embed dynamic adjustment options. The option stipulates that when the real-time price of the future spot market exceeds a certain threshold (such as 500 yuan / MWh) for N consecutive hours, the buyer has the right to require the contract power to be adjusted upwards by a certain percentage (such as 20%) within the next M hours (such as 24 hours) to lock in part of the high-priced spot market revenue. Conversely, when the price is continuously below a certain threshold, the seller has the right to require downward adjustment. This provides flexibility for both parties to deal with extreme price risks.
[0072] S1.13 Stress test: Stress test the S1.12 generated long-term contract with options using historical or simulated spot market data, simulate the contract performance and both parties' profits under various extreme market scenarios (e.g. continuous gloomy day, cold wave). If the test result shows that the contract is too risky for one party under a certain scenario (fails the robustness test), go back to S1.9, adjust the risk preference parameters in the game model, and solve again.
[0073] In this embodiment, based on the long-term contract and load forecasting, an initial day-ahead scheduling plan is developed. After signing a long-term contract, market participants need to develop a preliminary operation plan based on contract obligations and their own needs. Market participants (such as power sellers) combine the long-term contract obtained in S1 (which specifies the amount and price of electricity to be purchased from the market in the next 24 hours) and the load forecast of their proxy users to develop a preliminary initial day-ahead scheduling plan. This plan may include: planning to purchase how much electricity from the day-ahead market, the start-stop and output plan of the controllable power source (if any), etc. This plan is the basis for subsequent optimization.
[0074] In this embodiment, S3. Obtain the day-ahead market clearing result, and based on the clearing result, the dynamic adjustment option in the long-term contract, and the physical-financial coordinated optimization objective, rollingly re-optimize the initial day-ahead scheduling plan to obtain the final scheduling plan. After the day-ahead market is cleared, the latest market information is used to fine-tune the initial plan.
[0075] The system obtains the day-ahead market clearing result, including the clearing price and transaction volume of each period.
[0076] Based on the clearing result, the signed long-term contract and the dynamic adjustment option contained therein, the system takes the physical-financial coordinated optimization objective as the core to rollingly re-optimize the initial plan of S2 step.
[0077] The optimization objective function is: minimize (physical operation cost + absolute value of difference between long-term contract decomposed electricity and actual dispatched electricity × expected spot settlement price difference).
[0078] Physical operation cost: refers to the operation cost of the market participant's controllable unit (such as gas unit, energy storage), mainly including coal consumption cost (or gas cost) and start-stop cost. This part is the cost on the physical layer.
[0079] | Long-term contract decomposed electricity - actual dispatched electricity |: This represents the deviation of electricity due to the mismatch between the long-term contract and the actual dispatch. This deviation needs to be bought and sold in the spot market (real-time / balancing market), which will generate additional cost or income.
[0080] Expected spot price spread: refers to the difference between the expected buying or selling price and the long-term contract price when deviated electricity quantity is traded in real-time / balancing market. This value is determined based on the calibrated spot market price prediction obtained in step S1.6.
[0081] The core idea of this objective function is not only to pursue low cost of physical operation, but also to minimize the financial risk (deviation assessment risk) caused by the mismatch between contract and dispatch. By solving this optimization problem, a more optimal final dispatch plan is obtained.
[0082] In this embodiment, S4. Based on the final dispatch plan and the actual spot market price, a two-factor dynamic weight decomposition model is used to settle the long-term contract.
[0083] After the trading day ends, the long-term contract is settled according to the actual execution and market price.
[0084] S4.1 Calculate the base weight: Calculate the base weight of the kth hour. This weight is based on the inherent characteristics of the contract subject (user or power supplier) and uses its historical load or output curve to calculate the load or output proportion in the kth hour. For example, an industrial user's daytime electricity consumption accounts for 60% of its total daily electricity consumption, so its daytime,kis 0.6. This reflects the typical electricity consumption pattern of the user.
[0085] S4.2 Calculate the spot price correction weight: Calculate the spot price correction weight of the kth hour. This weight is calculated based on the deviation of the actual spot market price in the kth hour from the daily average price.
[0086] S4.3 Calculate the total decomposition weight: Calculate the total decomposition weight of the kth hour. The total decomposition weight is equal to the base weight multiplied by the first coefficient, plus the spot price correction weight multiplied by the second coefficient, and the sum of the first coefficient and the second coefficient is one. That is: total decomposition weight = (base weight x a) + (spot price correction weight x b), where a + b = 1; The first coefficient a and the second coefficient b can be set according to market rules or mutual agreement to balance the importance of "historical mode" and "market state" in settlement.
[0087] S4.4 Calculate the settlement electricity quantity: Calculate the settlement electricity quantity of the kth hour. The settlement electricity quantity of the kth hour is equal to the total decomposition weight of the kth hour multiplied by the total electricity quantity of the contract.
[0088] In this embodiment, S5. Based on the settlement result, generate a risk report and use the risk report as input to update the long-term contract strategy through a hierarchical reinforcement learning framework; feedback the transaction results to guide future decision-making, forming an intelligent closed loop.
[0089] S5.1 Calculate risk exposure: Based on the settlement results from S4, calculate the margin settlement amount and risk exposure. Risk exposure can be quantified as the potential maximum loss due to price fluctuation.
[0090] S5.2 Solve Pareto frontier: Based on the market participants' long-term power purchase contracts, the spot price prediction model obtained in S1.6, and the risk exposure calculated in S5.1, the system solves a risk-reward Pareto optimal frontier. This frontier is a set of solutions, where any solution cannot improve one objective without compromising the other. It provides decision-makers with all possible optimal trade-off solutions.
[0091] S5.3 Generate strategy recommendations: Based on the pre-set risk preference level (such as conservative, balanced, aggressive), select an optimal solution from the Pareto optimal frontier obtained in S5.2, and generate a risk report containing specific strategy recommendations. For example, for conservative users, it is recommended to increase the coverage rate of long-term contracts; for aggressive users, it is recommended to increase option purchases in specific periods.
[0092] S5.4 Receive risk report: The high-level agent in the hierarchical reinforcement learning framework receives the risk report generated in S5.3, which contains strategy recommendations based on the Pareto frontier. S5.5 Encode state vector: Encode the received risk report, long-term market trend information (such as annual price trend) and historical contract execution data (such as deviation rate in the past year) into a high-dimensional state vector as the input of the high-level agent.
[0093] S5.6 Output action: The strategy network (such as deep neural network) of the high-level agent outputs specific actions based on the input state vector. These actions include: adjusting the matching weight of the trading counterparties (preferentially trading with certain types of counterparties when looking for matches), modifying the optimization boundary of the contract quantity and price (such as adjusting the revenue target or risk tolerance upper limit in the game model S1.8), or generating option purchase instructions (clearly recommending adding what type of options in new contracts).
[0094] S5.7 Set reward function: The reward function of the high-level agent is designed to be multi-dimensional, including the expected revenue of long-term contracts, the amount of risk exposure reduction, and the amount of contract compliance rate improvement. This ensures that the learning goal of the agent is highly consistent with the long-term business goals of market participants.
[0095] S5.8 Receive constraints and information: The bottom agent in the hierarchical reinforcement learning framework receives the calibrated spot market price prediction, market state, and medium-long term position constraints (e.g., the high-level agent decides to hold a certain amount of long position in the future period) from the S1.6 step.
[0096] S5.9 Execute trading strategy: The strategy network of the bottom agent outputs specific day-ahead market order strategies and real-time market quotation strategies according to the current market state and received medium-long term position constraints.
[0097] S5.10 Set bottom reward: The reward function of the bottom agent includes the immediate income of spot market trading, deviation assessment fees, and the contribution to the risk hedging of medium-long term contracts. In particular, the "contribution to the risk hedging of medium-long term contracts" innovatively regards spot trading behavior as a risk management tool for medium-long term positions, encouraging the bottom agent to conduct trading in the spot market that is beneficial to reducing overall risk.
[0098] S5.11 Realize strategy coordination: The high-level agent and the bottom agent train through sharing the experience replay pool (storing the history state, action, and reward of both parties) and realize strategy coordination through meta-reward signals. The meta-reward signal is calculated based on the achievement of the high-level agent's long-term goal (such as annual total income, annual average risk exposure) by the bottom agent's execution results, ensuring that the two agents have consistent goals and co-evolve in the learning process.
[0099] In this embodiment, first, advanced data fusion and prediction techniques are used in S1 to generate probabilistic scenarios for future spot market prices (rather than a single prediction value), which provides a foundation for subsequent risk management. Then, it models the signing process of medium-long term contracts as a game process, allowing power suppliers and users to negotiate under full consideration of future market uncertainty (risk), ultimately generating a contract that meets the interests of both parties and has risk flexibility (through options). After the system is started, first, S1.1-S1.6 is executed to generate spot price prediction scenarios. Then, S1.7-S1.11 is executed to build and solve the stochastic double-layer game model, generating the base contract. Finally, S1.12 is executed for stress testing, and dynamic adjustment options are embedded in the contract (S1.11).
[0100] In this embodiment, in S2, after obtaining the medium-long term contract, the market subject needs to plan its physical operation in the next 24 hours. This step is a preliminary, static plan based on the contract and its own needs. The subject (such as a power selling company) compares the load forecast curve of its agent users with the power purchase curve specified in the medium-long term contract. If the load is greater than the contract power, the difference is planned to be purchased in the day-ahead market; if the load is less than the contract power, the excess power is planned to be sold. At the same time, considering the operation plan of its controllable resources (such as energy storage and gas units), a preliminary day-ahead market declaration plan is formed.
[0101] In this embodiment, in S3, the system obtains the day-ahead market clearing result. Then, taking the initial plan in S2 as the initial solution, a rolling optimization model is established. The optimization objective of this model is a physical-financial coordinated optimization objective, and the constraint conditions include grid safety constraints, unit operation constraints, energy storage charging and discharging constraints, and trigger conditions of dynamic adjustment options. By solving this optimization problem, a more optimal final dispatch plan is obtained to guide real-time operation.
[0102] In this embodiment, in S4, after the trading day ends, the medium-long term contract needs to be settled according to the actual execution and market prices. The traditional "settlement according to contract curve" or "settlement according to actual curve" has drawbacks. The two-factor model of the present invention combines the inherent power consumption characteristics of users and the market state of the day to more fairly and reasonably determine the settlement power.
[0103] In this embodiment, in S5, the experience (settlement result) of a single transaction is converted into guidance for future decision-making, realizing the self-learning and evolution of the system. The risk report is a summary of experience, and the hierarchical reinforcement learning framework is the engine of learning.
[0104] In this embodiment, in S1.1, the system obtains these data in real time or quasi-real time through data interfaces (such as SCADA systems, weather bureau APIs). For example, new energy output prediction values are published by provincial dispatch centers, social load forecasts are provided by load forecasting systems, and weather information is provided by regional weather station networks.
[0105] In this embodiment, in S1.2, the spatio-temporal graph neural network (ST-GNN): the power system is a physical network, and the weather and load of adjacent areas have spatial correlation. ST-GNN divides the grid or geographical area into multiple nodes (such as substations, weather stations, and load centers), and the connections between nodes form a graph. ST-GNN aggregates the information of neighboring nodes through a message passing mechanism, thereby learning the spatial dependence. For example, strong winds upstream can enhance the output of downstream wind farms.
[0106] Long Short-Term Memory (LSTM) Networks: Electricity load, renewable energy output, and meteorological parameters all exhibit strong periodicity (daily cycle, weekly cycle) and trends. LSTM is a special type of recurrent neural network (RNN) that effectively captures long-term dependencies and avoids the gradient vanishing problem through its "gate" mechanism (forget gate, input gate, output gate). Coupling: Spatial features extracted by ST-GNN (such as the integrated meteorological-load-output status of a certain region) are used as input to LSTM, which is responsible for capturing the time-series patterns of these spatial features as they evolve over time. This coupling method can capture both spatiotemporal features simultaneously. The model is built using deep learning frameworks such as PyTorch and TensorFlow. The ST-GNN part can be implemented using Graph Convolutional Networks (GCNs) or Graph Attention Networks (GATs). The LSTM part uses standard LSTM layers. The connection between the two can be achieved by concatenating the outputs of the ST-GNN at each time step and feeding them into the LSTM, or by using a more complex attention mechanism for fusion.
[0107] In this embodiment, in S1.3, a generative model is trained. Its input is the output features of the coupled model in S1.2, and its output is the price samples for each hour of the next 24 hours. By generating a large number of samples (e.g., 10,000), a histogram of the price distribution for each hour can be statistically derived, or a distribution function (e.g., normal distribution, Gaussian mixture distribution) can be fitted. The output of the coupled model is a high-dimensional feature vector containing an understanding of the current spatiotemporal state. The generative model (e.g., a conditional diffusion model or a conditional variational autoencoder CVAE) uses this feature vector as a condition to learn a mapping from a simple noise distribution (e.g., a standard normal distribution) to a complex target distribution (spot price distribution). It can generate a large number of future price samples that conform to historical patterns, thereby constructing a probability distribution.
[0108] Example 3 Based on the same inventive concept, this invention also provides a strategy optimization system based on the coupling of medium- and long-term contracts and spot trading, comprising: The report generation module is used to settle the medium- and long-term contracts based on the final scheduling plan and the actual spot market price, using a two-factor dynamic weight decomposition model, and to generate a risk report. The strategy update module is used to update the medium- and long-term contract strategy based on the risk report and the hierarchical reinforcement learning framework. The final scheduling plan is determined using a scheduling plan determination method based on medium- and long-term contracts as described above.
[0109] Preferably, the report generation module is specifically used for: The load or output proportion of the kth hour is calculated based on the historical load or output curve as the basic weight of the kth hour, wherein k is the time value; The spot price correction weight of the kth hour is calculated based on the deviation of the actual spot market price of the kth hour from the daily average price; The total decomposition weight of the kth hour is calculated based on the basic weight of the kth hour and the spot price correction weight of the kth hour combined with respective coefficients; The settlement electricity of the kth hour is calculated based on the total electricity of the contract multiplied by the total decomposition weight of the kth hour; The difference settlement amount and risk exposure are calculated based on the settlement electricity of the kth hour; The risk-reward Pareto optimal frontier is solved based on the medium and long-term electricity purchase contract of the market subject, the spot price prediction model and the risk exposure; Based on the preset risk preference level, the optimal solution is selected from the risk-reward Pareto optimal frontier to generate a risk report containing strategy suggestions.
[0110] Preferably, the strategy updating module is specifically used for: The risk report, long-term market trend information and historical contract execution data are encoded into a high-dimensional state vector; The specific action is output by the policy network of the high-level agent in the hierarchical reinforcement learning framework according to the state vector, and the action includes adjusting the matching weight of the trading counterparty, modifying the optimization boundary of the contract quantity and price, or generating an option purchase instruction; The learning goal of the agent is ensured to be consistent with the long-term business goal of the market subject through the reward function of the high-level agent; the reward function includes the expected income of the medium and long-term contract, the risk exposure reduction amount and the contract fulfillment rate improvement amount; The spot market price probability distribution, market state and medium and long-term holding constraints issued by the high-level agent are received by the bottom-level agent, and the specific day-ahead market bidding strategy and real-time market pricing strategy are output according to the current market state and holding constraints combined with the reward function of the bottom-level agent, wherein the reward function of the bottom-level agent includes the immediate income of the spot market transaction, the deviation assessment fee and the contribution degree of the risk hedging of the medium and long-term contract; The meta-reward signal is calculated according to the achievement degree of the long-term goal of the high-level agent; The high-level agent and the bottom-level agent are trained through sharing the experience replay pool, and the strategy coordination is performed through the meta-reward signal.
[0111] Embodiment 4 As Figure 4As shown, the present application further provides an electronic device, which can be a computer device, a single-chip microcomputer device, a smart mobile device, etc. The electronic device in the embodiment can include a processor, a memory, a transceiver component, etc. The memory, the processor and the transceiver component are connected through a bus; the memory can be used to store an execution program, and the exemplary execution program can include instructions; the processor is used to execute the instructions stored in the memory. The memory can also be used to store data, which can be called and / or modified when the instructions are executed.
[0112] The processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc., which are the computing core and control core of the terminal, and are suitable for implementing one or more instructions, and are specifically suitable for loading and executing one or more instructions in the storage medium to implement a corresponding method flow or a corresponding function, so as to implement the steps of the scheduling plan determination method based on the medium and long-term contract or the strategy optimization method based on the coupling of the medium and long-term contract and spot in the embodiment.
[0113] Embodiment 5 Based on the same inventive concept, the present application further provides a readable storage medium, specifically an electronic device readable storage medium (Memory), which is a memory device in the electronic device and is used to store programs and data. It can be understood that the storage medium herein can include the built-in storage medium in the electronic device, and of course can also include the expansion storage medium supported by the electronic device. The storage medium provides a storage space, and the storage space stores the operating system of the terminal. Moreover, one or more instructions suitable for being loaded and executed by the processor are also stored in the storage space, and the instructions can be one or more execution programs (including program codes). It should be noted that the storage medium herein can be a high-speed RAM memory, or a non-volatile memory, such as at least one disk memory. The loading and execution of one or more instructions stored in the storage medium by the processor can implement the steps of the scheduling plan determination method based on the medium and long-term contract or the strategy optimization method based on the coupling of the medium and long-term contract and spot in the embodiment.
[0114] Those skilled in the art will appreciate that embodiments of the present application can be readily used as a method, a system, or a computer program product. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer readable program code.
[0115] The present application is described in reference to the flowchart and / or block diagrams of the method, apparatus (system) and computer program product according to embodiments of the application. It will be understood that each block of the flowchart and / or block diagrams, and combinations of blocks in the flowchart and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processing system or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks.
[0116] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks.
[0117] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more functions specified in the flowchart and / or block diagram block or blocks.
[0118] The above merely provides an embodiment of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall fall within the scope of the present application.
Claims
1. A method for determining a scheduling plan based on medium- and long-term contracts, characterized in that, include: Medium- to long-term contracts with dynamically adjusted options are generated based on multi-source heterogeneous data combined with a pre-built coupled model and a stochastic two-level game model. Based on historical data of agent users, predict the short-term load of agent users and formulate an initial scheduling plan in combination with medium and long-term contracts; Obtain the day-ahead market clearing results, and based on the clearing results, the dynamic adjustment options in medium- and long-term contracts, and the physical-financial collaborative optimization objective, perform rolling optimization on the initial scheduling plan to obtain the final scheduling plan.
2. The method as described in claim 1, characterized in that, Based on the clearing results, the dynamic adjustment options in medium- and long-term contracts, and the physical-financial collaborative optimization objective, the initial scheduling plan is continuously optimized to obtain the final scheduling plan, including: The total cost is calculated based on physical operating costs, the electricity volume allocated from medium- and long-term contracts, the actual dispatched electricity volume in the clearing results, and the expected spot settlement price difference. With the goal of minimizing the total cost, the initial scheduling plan is continuously optimized under the constraints to obtain the final scheduling plan.
3. The method as described in claim 1, characterized in that, The method for generating medium- to long-term contracts containing dynamically adjusted options based on multi-source heterogeneous data combined with a pre-built coupled model and a stochastic two-level game model includes: By substituting multi-source heterogeneous data into a pre-built spot price prediction model, the probability distribution of spot market prices is obtained. Based on the probability distribution of spot market prices, the historical output curves of renewable energy generators, the historical electricity consumption curves of market users, the installed capacity of renewable energy generators and market rules, a medium- and long-term contract containing dynamic adjustment options is generated by combining a pre-constructed stochastic two-layer game model. The pre-built spot price prediction model is based on a spatiotemporal graph constructed using graph neural networks from multiple meteorological stations, new energy power generation nodes, and key load centers within the province, combined with graph convolutional networks or graph attention networks and long short-term memory networks.
4. The method as described in claim 3, characterized in that, The step of substituting multi-source heterogeneous data into a pre-built spot price prediction model to obtain the probability distribution of spot market prices includes: Multiple meteorological stations, new energy power generation nodes, and key load centers within the province are regarded as nodes in the graph, and a spatiotemporal graph is constructed using geographical distance or power grid topology as edges. Based on the spatiotemporal graph, a graph convolutional network or a graph attention network is used to aggregate the information of neighboring nodes at each time step and extract spatial feature vectors. The spatial feature vectors are concatenated and then input into a long short-term memory network to capture long-term trends and periodic patterns in the time series data of each node, resulting in a high-dimensional feature vector that integrates spatial correlation and temporal evolution. The high-dimensional feature vectors are used to generate a probability distribution of spot market prices using a generative model.
5. The method as described in claim 3, characterized in that, The process of generating medium- to long-term contracts containing dynamically adjusted options based on the probability distribution of spot market prices, the historical output curves of renewable energy generators, the historical electricity consumption curves of market users, the installed capacity of renewable energy generators, and market rules, combined with a pre-constructed stochastic two-layer game model, includes: The stochastic two-level game model was transformed into a single-level linear model by using KKT conditions and strong duality theory and then solved to obtain two sets of candidate contract curves containing electricity price and power. The entropy weight method is used to determine the weights of the two sets of candidate contract curves, and the basic contract electricity price and power curve are calculated based on the weights. By embedding a dynamic adjustment option into the basic contract electricity price, a medium- to long-term contract containing the dynamic adjustment option is formed.
6. A strategy optimization method based on the coupling of medium- and long-term contracts and spot trading, characterized in that, include: Based on the final scheduling plan and actual spot market prices, a two-factor dynamic weight decomposition model is used to settle the medium- and long-term contracts and generate risk reports. Update medium- and long-term contract strategies based on the aforementioned risk report and hierarchical reinforcement learning framework; The final scheduling plan is determined using a scheduling plan determination method based on medium- and long-term contracts as described in any one of claims 1-5.
7. The method as described in claim 6, characterized in that, The settlement of the medium- and long-term contracts is based on the final scheduling plan and actual spot market prices, using a two-factor dynamic weight decomposition model, and a risk report is generated, including: Based on the characteristics of the main body of medium and long-term contracts, the load or output ratio of the k-th hour is calculated using historical load or output curves, which serves as the basic weight for the k-th hour, where k is the time value; The spot price adjustment weight for the k-th hour is calculated based on the deviation between the actual spot market price and the daily average price in the k-th hour. The total decomposition weight of the k-th hour is calculated by combining the base weight of the k-th hour and the spot price adjustment weight of the k-th hour with their respective coefficients. The settlement electricity volume for the k-th hour is calculated by multiplying the total decomposition weight of the k-th hour by the total contract electricity volume. Calculate the price difference settlement amount and risk exposure based on the settlement electricity volume in the k-th hour; Based on the market participants' medium- and long-term power purchase contracts, spot price forecasting models, and risk exposures, the risk-return Pareto optimal frontier is solved. Based on a preset risk preference level, the optimal solution is selected from the risk-return Pareto optimal frontier to generate a risk report containing strategy recommendations.
8. A strategy optimization system based on the coupling of medium- and long-term contracts and spot trading, characterized in that, include: The report generation module is used to settle the medium- and long-term contracts based on the final scheduling plan and the actual spot market price, using a two-factor dynamic weight decomposition model, and to generate a risk report. The strategy update module is used to update the medium- and long-term contract strategy based on the risk report and the hierarchical reinforcement learning framework. The final scheduling plan is determined using a scheduling plan determination method based on medium- and long-term contracts as described in any one of claims 1-5.
9. An electronic device, characterized in that, include: At least one processor and memory; The memory and processor are connected via a bus; The memory is used to store one or more programs; When the one or more programs are executed by the at least one processor, a scheduling plan determination method based on medium- and long-term contracts as described in any one of claims 1 to 5 is implemented, or a strategy optimization method based on the coupling of medium- and long-term contracts and spot markets as described in any one of claims 6 to 7 is implemented.
10. A readable storage medium, characterized in that, It contains an execution program, which, when executed, implements a scheduling plan determination method based on medium- and long-term contracts as described in any one of claims 1 to 5, or a strategy optimization method based on the coupling of medium- and long-term contracts and spot markets as described in any one of claims 6 to 7.