Method and device for generating day-ahead market report and quotation strategy of thermal power generating unit
By generating day-ahead market quotation strategies for thermal power units using an intelligent decision-making framework based on reinforcement learning, the limitations of existing technologies that rely on human experience and static cost models are overcome. This enables global optimization and self-adaptation in complex power market environments, thereby improving the profitability and competitiveness of thermal power units.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA RESOURCES POWER TECH RES INST CO LTD
- Filing Date
- 2025-12-26
- Publication Date
- 2026-05-26
AI Technical Summary
Existing day-ahead market pricing strategies for thermal power units rely on manual experience or static cost models, making it difficult to achieve global optimization in complex electricity market environments. They lack adaptability, cannot effectively cope with market uncertainties and price fluctuations, and lack a joint optimization mechanism for returns and risks.
An intelligent decision-making framework based on reinforcement learning is adopted. By acquiring and preprocessing market data, a state space is constructed and input into a pre-trained policy network to generate the optimal volume and price strategy. The parameters are optimized using a deep reinforcement learning network to generate a segmented price curve that satisfies the monotonically increasing constraint.
It has improved the profitability and market competitiveness of thermal power units in complex market environments, enhanced the robustness and interpretability of bidding strategies, realized the transformation from experience-driven to data-driven intelligent decision-making, and improved profitability and bidding robustness in volatile electricity markets.
Smart Images

Figure CN122089428A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power market strategy generation technology, and in particular to a method and apparatus for generating day-ahead market quotation strategies for thermal power units. Background Technology
[0002] The electricity market has become increasingly important for the operation and profitability of power generation companies. Power generation companies are required to submit segmented bidding curves for each generating unit one day in advance. These curves serve as the basis for market clearing calculations, and must satisfy both unit operating constraints and balance profitability with bidding risks. A reasonable bidding strategy can help power generation companies obtain greater profits in market competition while ensuring the stable operation of the power system. This is of great significance for improving the economic efficiency of power generation companies, optimizing the allocation of power resources, and promoting the healthy development of the electricity market.
[0003] Currently, thermal power unit pricing primarily employs either manual experience-based methods or static cost model methods. Manual methods rely on traders' experience, who formulate pricing strategies based on their knowledge and past trading experience. Cost model methods, on the other hand, consider only marginal costs and safety margins, determining prices by calculating unit costs and setting a certain safety range.
[0004] However, in the complex day-ahead electricity market environment, the high-dimensional and coupled boundary conditions lead to a non-convex and severely discretized feasible region. Traditional static optimization methods often can only search within a local region, making global optimization difficult. Furthermore, existing strategies lack adaptability to the uncertainties of the day-ahead market, mostly relying on fixed parameters or single predicted values, failing to consider the random volatility of market prices. This makes the strategies prone to missing bidding opportunities in high-volatility scenarios and unable to proactively adjust in low-volatility scenarios. Moreover, existing methods generally lack a joint optimization mechanism for revenue, quantity preservation, and risk, making it difficult to meet the multi-objective decision-making needs of different enterprises with varying operational preferences. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a method and apparatus for generating day-ahead market quotation strategies for thermal power units, aiming to solve at least one of the above-mentioned technical problems.
[0006] The technical solution of the present invention to solve the above-mentioned technical problems is as follows: Firstly, this application provides a method for generating day-ahead market quotation strategies for thermal power units, employing the following technical solution: A method for generating a day-ahead market quotation strategy for thermal power units, comprising: S1, Obtain preprocessed current market data, which includes current day-ahead forecast electricity price, current real-time forecast electricity price, current fuel cost data, unit operation constraints, and medium- and long-term contract information, which includes contracted electricity volume and contracted electricity price; S2, Based on the preprocessed current market data, construct the current state space, which includes market environment prediction information and the unit's own state parameters; S3, input the current state space into the pre-trained policy network to obtain the optimal volume quotation policy curve; The pre-trained policy network is trained based on historical market data through a reinforcement learning framework. In the reinforcement learning framework, thermal power units are modeled as agents, and the market clearing process and revenue calculation are modeled as the environment. Through continuous interaction between the agent and the environment, the parameters of the policy network are optimized to maximize the cumulative reward, thus obtaining the pre-trained policy network.
[0007] The beneficial effects of this invention are as follows: By constructing an intelligent decision-making framework based on reinforcement learning, complex market game theory and unit constraint optimization problems are transformed into a data-driven self-learning process, thereby fundamentally overcoming the limitations of traditional methods that rely on human experience or static cost models. By acquiring preprocessed current market data and constructing a state space, and then inputting the state space into a pre-trained policy network based on the reinforcement learning framework, the optimal quantity and price quotation strategy curve can be generated. This solves the problem that existing thermal power units rely on human experience and lack intelligent optimization in quantity and price quotation under complex spot market environments, improves the expected returns and market competitiveness of power generation companies, and enhances the robustness and interpretability of the quotation strategy, realizing a shift from "experience-driven" to "data and model-driven" intelligent decision-making.
[0008] Based on the above technical solution, the present invention can be further improved as follows.
[0009] Furthermore, the training method for the policy network includes: S11, Obtain the original historical market data and preprocess the original historical market data to obtain the preprocessed historical market data sequence; S12, retrieve the state space of the current iteration's historical day from the historical market data sequence in chronological order; S13, input the state space of the historical day into the current policy network to be trained, and generate a bidding policy action. The bidding policy action represents the segmented quantity bidding curve to be evaluated. The segmented quantity bidding curve includes the bid price and bid volume of each segment. S14, input the pricing strategy action into the pre-built market simulation environment, determine the expected return corresponding to the pricing strategy action and the next state space of the historical day's state space, and determine the reward space of the current iteration based on the expected return, wherein the reward space characterizes the degree to which the pricing strategy action obtains a good return. S15: Based on the reward space of the current iteration, the state space of the historical day, the bidding policy action, and the next state space of the state space of the historical day, the policy gradient is calculated using the reinforcement learning network to update the parameters of the current policy network. S16. Repeat steps S12 to S15 until the historical data sequence has been traversed or the performance index of the policy network has converged, and the trained policy network is obtained.
[0010] The beneficial effects of adopting the above-mentioned further scheme are as follows: After acquiring and preprocessing the original historical market data, the bidding strategy action is generated using the historical daily state space and the strategy network to be trained. The expected return and reward space are determined through the market simulation environment. Then, the strategy network parameters are updated using a reinforcement learning network. After multiple iterations of training, a strategy network that can generate the optimal bidding strategy curve for thermal power units is obtained. This improves the profitability and bidding robustness of thermal power units in the volatile electricity market, enables thermal power units to form a more adaptable and economical bidding strategy under multi-market linkage conditions, and enhances the competitiveness of units in the dynamic market environment.
[0011] Furthermore, the preprocessing of the original historical market data to obtain a preprocessed historical market data sequence includes: Outlier detection and removal, as well as missing value processing, are performed on the original historical market data to obtain the first historical market data. Align the data from different sources in the first historical market data according to the set timestamps to obtain the second historical market data; Based on the second historical market data, the constraint parameters of the thermal power unit operating boundary are extracted; Based on the constraint parameters of the thermal power unit's operating boundary, the second historical market data, and the preset time sequence, the preprocessed historical market data sequence is determined.
[0012] The beneficial effects of adopting the above-mentioned further solutions are as follows: outlier detection and removal, as well as missing value processing, can remove errors and missing parts of the data and improve data quality; aligning data from different sources according to a set timestamp ensures that the data corresponds to the same time period, making subsequent analysis more accurate; extracting constraint parameters of the thermal power unit's operating boundary can provide constraints on unit operation for subsequent strategy generation; and determining the preprocessed historical market data sequence based on constraint parameters, second historical market data, and a preset time sequence provides a high-quality, orderly data foundation that conforms to unit operation constraints for subsequent strategy network training, which helps to generate more accurate and effective volume and price quotation strategies.
[0013] Furthermore, the step of inputting the pricing strategy action into a pre-built market simulation environment and determining the expected return corresponding to the pricing strategy action includes: The preprocessed historical market data sequence is input into a preset probability prediction model to obtain the probability distribution information of electricity prices; Based on the day-ahead electricity price in the historical market data sequence, the probability distribution information of the electricity price, and the random sampling method, multiple day-ahead market electricity price scenarios and real-time market electricity price scenarios are generated, and each set of electricity price scenarios corresponds to a predicted electricity price; For each electricity price scenario, the bidding strategy action and the electricity price scenario are input into a preset market clearing model for evaluation to obtain the winning bid volume of the bidding strategy action under the electricity price scenario; For each electricity price scenario, the revenue under the electricity price scenario is determined based on the winning bid volume, the day-ahead electricity price and real-time electricity price corresponding to the electricity price scenario, and the preset revenue calculation model. Based on the revenue under each of the aforementioned electricity price scenarios, the expected revenue corresponding to the pricing strategy action is determined.
[0014] The beneficial effects of adopting the above-mentioned further scheme are as follows: Preprocessed historical market data is input into a probabilistic prediction model to obtain electricity price probability distribution information. Then, multiple electricity price scenarios are generated by combining this with random sampling, simulating various electricity price fluctuation scenarios and considering market uncertainty. Under each electricity price scenario, the bidding strategy actions are evaluated to obtain the winning bid volume. Then, the revenue is determined by combining this with a revenue calculation model, ultimately yielding the expected revenue. Based on electricity price probability prediction and a randomized market clearing mechanism, the scheme can simulate the winning bid results and revenue performance of thermal power units under different electricity price fluctuation scenarios. Deep reinforcement learning algorithms are used to generate bidding strategies with stable revenue performance in multiple scenarios, improving the profitability and bidding robustness of thermal power units in a volatile electricity market.
[0015] Furthermore, the step of inputting the bidding strategy action and the electricity price scenario into a preset market clearing model for evaluation, to obtain the winning bid volume of the bidding strategy action under the electricity price scenario, includes: The bid prices for each segment of the bidding strategy action are compared with the predicted electricity prices under the electricity price scenario to obtain the comparison results; Multiple winning segments were determined based on the comparison results; Based on the declared electricity volume of each of the winning bid segments, the winning bid volume of the bidding strategy action under the electricity price scenario is determined; The determination of multiple winning segments based on comparison results includes: For any segment in the bidding strategy action, if the bid price of the segment is not higher than the predicted electricity price under the electricity price scenario, then the segment is determined to be the winning segment. For any segment in the bidding strategy action, if the bid price of the segment is higher than the predicted electricity price under the electricity price scenario, then the segment is determined not to be the winning segment.
[0016] The beneficial effects of adopting the above-mentioned further scheme are: the winning bid segment is determined by comparing the bid price of each segment in the bidding strategy with the predicted electricity price under the electricity price scenario, and then the winning bid volume is determined. This can simulate the winning bid results of thermal power units under different electricity price fluctuation scenarios, provide an accurate basis for calculating expected revenue and optimizing the bidding strategy, and improve the profitability and bidding stability of thermal power units in the volatile electricity market.
[0017] Furthermore, the preset revenue calculation model is as follows: ; Where i is the time period index; n is the number of time periods divided into that day; , These represent the medium- and long-term contract electricity price and contract electricity volume corresponding to the i-th time period, respectively. , These represent the clearing price and the electricity volume declared or won by generating units in the i-th period of the day-ahead market, respectively. , This represents the settlement price and actual output or power volume of the electricity in the i-th time period of the real-time market. This represents the average marginal generation cost of the thermal power unit in the i-th time period, which is calculated based on the fuel unit price, unit thermal efficiency, and auxiliary power consumption.
[0018] The beneficial effects of adopting the above-mentioned further scheme are: to unify the various revenues and costs of thermal power units under the framework of medium and long-term contracts, day-ahead declarations and real-time deviation settlements into a single mathematical expression, and to provide clear optimization objectives and quantifiable benefit evaluation indicators for the optimization of the reporting and pricing strategy based on reinforcement learning.
[0019] Furthermore, the preset probability prediction model is a Monte Carlo simulation model or an XGBoost machine learning model, and the method for constructing the preset probability prediction model is as follows: The Monte Carlo simulation model or XGBoost machine learning model is trained based on historical cleared electricity prices, cleared electricity volume and load demand data to obtain a trained probabilistic prediction model.
[0020] The beneficial effects of adopting the above-mentioned further solutions are: using Monte Carlo simulation models or XGBoost machine learning models as preset probabilistic prediction models, and training them based on historical clearing electricity prices, clearing electricity volume, and load demand data, it is possible to accurately predict electricity prices, obtain the probability distribution information of electricity prices, and provide the necessary price information basis for subsequently generating multiple electricity price scenarios and determining the expected revenue corresponding to the bidding strategy actions.
[0021] Furthermore, the formula for calculating the reward space is as follows: ; in, , , Configurable weight parameters; Used to measure the degree to which the target of maintaining quantity is achieved; This indicator is used to describe the risk of bids deviating from predicted electricity prices, bids being too high leading to non-transactions, or price fluctuations. To predict electricity prices.
[0022] The beneficial effects of adopting the above-mentioned further scheme are: by using configurable weight parameters α, β, and λ, and combining the measurement of the degree of completion of the guaranteed quantity target and the characterization of the bidding risk, a reward space calculation formula can be constructed. This enables the reinforcement learning agent to learn the optimal bid price-output pair that meets the preferences under different business needs and around the profit target, thereby achieving joint optimization of the profit, risk, and execution results of the bidding strategy, so as to meet the different operating preferences and multi-objective decision-making needs of power generation companies.
[0023] Furthermore, the reinforcement learning network is a deep reinforcement learning network, which includes a policy network, a value network, and a target network. The policy network generates bid-ask policies and actions based on the input state space. The policy network employs a multilayer perceptron structure and integrates an attention mechanism to generate segmented bid-ask curves that satisfy monotonically increasing constraints. The value network evaluates the long-term value of state-action pairs. The value network employs a deep network structure with dual-path encoding of states and actions. The target network includes a policy target network and a value target network. The target network updates the parameters of the policy network and the value network based on a soft update mechanism.
[0024] The beneficial effects of adopting the above-mentioned further scheme are as follows: The policy network adopts a multilayer perceptron structure and integrates an attention mechanism, which can generate a segmented price quotation curve that satisfies the monotonically increasing constraint based on the input state space, thus helping to automatically generate a pricing strategy that conforms to market rules; the value network adopts a deep network structure with dual-path encoding of state and action, which can accurately evaluate the long-term value of state-action pairs; the target network updates the parameters of the policy network and the value network based on a soft update mechanism, which can provide a relatively stable training target for pricing strategy learning, enabling the agent to gradually converge to the optimal strategy in a dynamically changing market environment. Overall, such a deep reinforcement learning network can improve the accuracy and stability of pricing strategy generation, thereby enhancing the competitiveness and profitability of thermal power units in the market.
[0025] Secondly, this application provides a device for generating day-ahead market quotation strategies for thermal power units, which adopts the following technical solution: A device for generating a day-ahead market quotation strategy for thermal power units includes: The acquisition module is used to acquire preprocessed current market data, which includes the current day-ahead forecast electricity price, the current real-time forecast electricity price, current fuel cost data, unit operation constraints, and medium- and long-term contract information, which includes contracted electricity volume and contracted electricity price. The construction module is used to construct the current state space based on the preprocessed current market data. The current state space includes market environment prediction information and the unit's own state parameters. The volume reporting and pricing strategy generation module is used to input the current state space into a pre-trained strategy network to obtain the optimal volume reporting and pricing strategy curve. The pre-trained policy network is trained based on historical market data through a reinforcement learning framework. In the reinforcement learning framework, thermal power units are modeled as agents, and the market clearing process and revenue calculation are modeled as the environment. Through continuous interaction between the agent and the environment, the parameters of the policy network are optimized to maximize the cumulative reward, thus obtaining the pre-trained policy network.
[0026] Additional aspects and advantages of this application will be set forth in part in the description which follows, and will become apparent from the description or may be learned by practice of this application. Attached Figure Description
[0027] Figure 1 A flowchart illustrating a method for generating market pricing strategies for thermal power units, provided as an embodiment of the present invention; Figure 2 This is a schematic diagram of a thermal power unit market pricing strategy generation device provided in one embodiment of the present invention. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0029] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article, unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship.
[0030] This application provides a method for generating market quotation strategies for thermal power units. This method can be executed by an electronic device, which can be a server or a mobile terminal device. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services. The mobile terminal device can be a laptop computer, a desktop computer, etc., but is not limited to these.
[0031] To make the technical solution of this invention clearer, the core terms used in this specification are defined below: Reinforcement learning is an artificial intelligence decision optimization algorithm that uses a "state-action-reward" feedback mechanism to allow the agent to continuously try and learn the optimal decision strategy in the environment. It is widely used in complex dynamic optimization scenarios. In the current market, which represents the electricity spot market, market participants submit output-price declaration curves or quantity declaration curves to the trading center the day before the trading day. The market determines the winning generating units and output plans through a clearing algorithm. The day-ahead quotation strategy represents the price-output combination decision made by power generation companies based on information such as the operating constraints of thermal power units, fuel costs, and market forecasts, and is used to participate in the day-ahead market clearing.
[0032] like Figure 1 As shown, a method for generating market quotation strategies for thermal power units includes: S1, Obtain preprocessed current market data, which includes current day-ahead forecast electricity price, current real-time forecast electricity price, current fuel cost data, unit operation constraints, and medium- and long-term contract information, which includes contracted electricity volume and contracted electricity price; S2, Based on the preprocessed current market data, construct the current state space, which includes market environment prediction information and the unit's own state parameters; S3, input the current state space into the pre-trained policy network to obtain the optimal volume quotation policy curve; The pre-trained policy network is trained based on historical market data through a reinforcement learning framework. In the reinforcement learning framework, thermal power units are modeled as agents, and the market clearing process and revenue calculation are modeled as the environment. Through continuous interaction between the agent and the environment, the parameters of the policy network are optimized to maximize the cumulative reward, thus obtaining the pre-trained policy network.
[0033] This method constructs an intelligent decision-making framework based on reinforcement learning, transforming complex market game theory and unit constraint optimization problems into a data-driven self-learning process. This fundamentally overcomes the limitations of traditional methods that rely on human experience or static cost models. By acquiring preprocessed current market data and constructing a state space, and then inputting the state space into a pre-trained policy network based on the reinforcement learning framework, the optimal quantity and price quotation strategy curve can be generated. This solves the problem that existing thermal power units rely on human experience and lack intelligent optimization in quantity and price quotation under complex spot market environments. It improves the expected returns and market competitiveness of power generation companies, enhances the robustness and interpretability of the quotation strategy, and realizes the transformation from "experience-driven" to "data and model-driven" intelligent decision-making.
[0034] In this application, the reinforcement learning network is a deep reinforcement learning network, which includes a policy network, a value network, and a target network. The strategy network is used to generate bidding strategy actions based on the input state space. The strategy network employs a multilayer perceptron structure and integrates an attention mechanism to generate segmented bid-ask curves that satisfy monotonically increasing constraints. It learns a nonlinear mapping from market state to the optimal price-output bid curve (optimal PQ bid curve) through a multilayer neural network. The network can automatically generate K bid segments, each containing the cumulative bid output. and declared price The compliance of the PQ curve is ensured through a monotonically increasing constraint mechanism. The policy function expression is as follows: , represents the deterministic mapping for generating quote action A in state St, where These are the policy network parameters.
[0035] The value network is used to evaluate the long-term value of state-action pairs. It employs a deep network structure with dual-path encoding of states and actions, mapping market states and bid actions to value assessment results. To reduce overestimation, a dual-network architecture can be used. The state-value function is... , indicating state Long-term value valuation; state-action value function is , indicating the state The following is an estimate of the long-term cumulative return of taking pricing action A, where and These are the parameters of the value network.
[0036] The target network includes a policy target network and a value target network. The target network updates the parameters of the policy network and the value network based on a soft update mechanism. By periodically synchronizing the main network parameters through the soft update mechanism, the agent gradually converges to the optimal policy in a dynamically changing market environment. The target network includes the policy target network. and value target network Parameter updates are adopted Form, in which To update the coefficients. The formula for calculating the target value of the time-series difference is as follows: ,in This is the discount factor.
[0037] To address the differentiated impact of electricity price fluctuations at different times on pricing strategies in the electricity market, an attention mechanism is employed to identify and extract key time-period features. The time-series attention module calculates the importance weights for the electricity price sequences of the i-th time period. Generate weighted feature representation The feature attention module calculates the importance weights of each dimension of the state vector through a channel attention mechanism, enabling adaptive focusing on key factors such as high electricity price periods and tightly constrained scenarios.
[0038] Optional methods for training the policy network include: S11, Obtain the original historical market data and preprocess the original historical market data to obtain the preprocessed historical market data sequence; S12, retrieve the state space of the current iteration's historical day from the historical market data sequence in chronological order; S13, input the state space of the historical day into the current policy network to be trained, and generate a bidding policy action. The bidding policy action represents the segmented quantity bidding curve to be evaluated. The segmented quantity bidding curve includes the bid price and bid volume of each segment. S14, input the pricing strategy action into the pre-built market simulation environment, determine the expected return corresponding to the pricing strategy action and the next state space of the historical day's state space, and determine the reward space of the current iteration based on the expected return, wherein the reward space characterizes the degree to which the pricing strategy action obtains a good return. S15: Based on the reward space of the current iteration, the state space of the historical day, the bidding policy action, and the next state space of the state space of the historical day, the policy gradient is calculated using the reinforcement learning network to update the parameters of the current policy network. S16. Repeat steps S12 to S15 until the historical data sequence has been traversed or the performance index of the policy network has converged, and the trained policy network is obtained.
[0039] After acquiring and preprocessing the original historical market data, the bidding strategy actions are generated using the historical daily state space and the strategy network to be trained. The expected returns and reward space are determined through a market simulation environment. Then, the strategy network parameters are updated using a reinforcement learning network. After multiple iterations of training, a strategy network that can generate the optimal bidding strategy curve for thermal power units is obtained. This improves the profitability and bidding robustness of thermal power units in the volatile electricity market, enabling thermal power units to form more adaptive and economical bidding strategies under multi-market linkage conditions, and enhancing the competitiveness of units in dynamic market environments.
[0040] In this embodiment, a revenue calculation model for thermal power units in the spot market needs to be pre-constructed. Based on the bidding method for thermal power in the spot electricity market, a revenue calculation formula for thermal power units is selected. According to the revenue calculation formula, various parameters required for market operation are continuously collected, including day-ahead and real-time forecast electricity prices in the spot market, fuel cost data, and unit operation data. Daily revenue of thermal power units is calculated. It can be described as: ; Where i is the time period index; n is the number of time periods divided into that day; , These represent the medium- and long-term contract electricity price and contract electricity volume corresponding to the i-th time period, respectively. , These represent the clearing price and the electricity volume declared or won by generating units in the i-th period of the day-ahead market, respectively. , This represents the settlement price and actual output or power volume of the electricity in the i-th time period of the real-time market. This represents the average marginal generation cost of the thermal power unit in the i-th time period, which is calculated based on the fuel unit price, unit thermal efficiency, and auxiliary power consumption.
[0041] A probabilistic prediction model for day-ahead and real-time clearing electricity prices in the electricity market is pre-constructed. The probabilistic prediction model is a Monte Carlo simulation model, an XGBoost machine learning model or its variants, or other machine learning models capable of outputting interval information. The method for constructing the pre-constructed probabilistic prediction model is as follows: The Monte Carlo simulation model or XGBoost machine learning model is trained based on historical cleared electricity prices, cleared electricity volume and load demand data to obtain a trained probabilistic prediction model.
[0042] Using historical market operation data such as cleared electricity prices, cleared electricity volume, and load demand, a forecasting method capable of characterizing uncertainty is employed to model electricity prices for each time period, obtaining the expected value μ and uncertainty parameter σ of the predicted electricity price. μ represents the central level of the predicted electricity price, and σ reflects the possible fluctuation range of the predicted result; its acquisition method is not limited and can be achieved through any method capable of describing price dispersion or range. The μt and σt output by the forecasting model provide necessary inputs for the randomized clearing simulation model, providing a necessary price information foundation for subsequent pricing strategy optimization and electricity market simulation.
[0043] Pre-construct the state space Action space A and reward space Rt.
[0044] The state space St contains market environment information and operating parameters of the thermal power unit in time period t, specifically including: day-ahead electricity price forecast sequence and its uncertainty parameters, real-time electricity price forecast sequence and its uncertainty parameters, medium- and long-term contract electricity volume and price, current minimum / maximum output limits of the unit, fuel cost, ramp-up constraints, historical bidding success rate, historical average revenue, and other features. The state vector on day t is... It can be represented as: ; Where t represents day t; The day-ahead market forecast electricity price sequence for day t; Let be the real-time market forecast electricity price sequence for day t; It is the medium- to long-term contract electricity price on day t; This refers to the medium- to long-term contract electricity volume on day t. This is the fuel cost parameter for day t; It is the historical success rate; It is the historical average return; This is the lower limit of the unit's output. This is the upper limit of the generator unit's output; It is the lower limit of the market price; It is the upper limit of market price; Historical reported electricity volume Historical declared prices.
[0045] Through the aforementioned state space design, the agent can comprehensively consider multi-dimensional information such as electricity price expectations, contractual constraints, unit characteristics, and market competition, providing sufficient decision-making basis for generating the optimal pricing strategy. The state vector, after standardization, is input into the policy network and value network for subsequent calculations.
[0046] Pre-construct an action space A, defining the action space as: For any time interval i: ; Constraints: ; ; In the formula, K is the number of price segments for that period, which is generally 3 to 10 segments and can be configured according to actual market rules; The cumulative declared output corresponding to segment j is within the range of the upper and lower limits of thermal power unit output. ; To correspond to the declared price, the price range is the upper and lower limits of the declared price as stipulated by the market at the present time. .
[0047] Design Reward Space Rt: in, , , Configurable weight parameters; Used to measure the degree of achievement of the guaranteed power generation target, such as the proportion of historical winning bids for the current month to the current month's power generation target; This indicator is used to characterize risks such as bid deviations from predicted electricity prices, excessively high bids leading to non-transactions, or price fluctuations. Because returns are randomized due to the influence of sampled electricity price scenarios, the reward function is reflected during training as a statistical measure of returns under multiple scenarios (such as expected return or expected return after penalty). To predict electricity prices.
[0048] For different business scenarios, this invention constructs various scenario-based reward functions by adjusting the forms of α, β, λ and related functions, including but not limited to: Prioritizing Quantity in Bidding Scenarios: This applies to situations where it is essential to guarantee the awarded electricity volume, such as when undertaking supply obligations or needing to fulfill high contracted volumes. In this case, increasing the β weight and decreasing the λ weight encourages the agent to use lower bids to increase the success rate without incurring significant losses.
[0049] Profit-first scenario: When there is no mandatory requirement to guarantee the quantity and only the pursuit of maximizing the revenue per kilowatt-hour is to increase the weight of α and decrease β, so that the agent pays more attention to the marginal revenue during high electricity price periods and increases the bidding level within an acceptable risk range.
[0050] Enhanced risk constraints scenario: Higher penalties are imposed on behaviors such as bids deviating from predicted electricity prices and excessively aggressive bids leading to large-scale non-transactions, increasing λ to ensure that the agent balances robustness and predictability when making bids.
[0051] Different scenarios can correspond to different sets of fixed weight configurations, or a "scenario identifier" feature can be added to the state space, allowing a single reinforcement learning model to adaptively output different PQ policies under different scenario identifiers.
[0052] In this embodiment of the application, the original historical market data is preprocessed to obtain a preprocessed historical market data sequence, including: Outlier detection and removal, as well as missing value processing, are performed on the original historical market data to obtain the first historical market data. Align the data from different sources in the first historical market data according to the set timestamps to obtain the second historical market data; Based on the second historical market data, the constraint parameters of the thermal power unit operating boundary are extracted; Based on the constraint parameters of the thermal power unit's operating boundary, the second historical market data, and the preset time sequence, the preprocessed historical market data sequence is determined.
[0053] In the above implementation, the collected market data undergoes preprocessing, including removing outlier data, aligning time dimensions, organizing unit constraint parameters, and structuring samples. Basic outlier removal includes simple anomaly detection of key data such as day-ahead clearing prices, system load, actual unit output, and fuel costs, removing obvious erroneous and missing values. Time dimension alignment refers to uniformly aligning variables such as medium- and long-term contract electricity volume, day-ahead clearing electricity volume, actual unit output, and fuel costs according to time indices.
[0054] Unit constraint parameter processing refers to extracting static parameters such as minimum output, maximum output, technical constraints (such as ramping), and marginal cost of the unit for reinforcement learning state construction; sample structuring involves constructing training samples for each time period based on cleaned and aligned data to form structured input feature vectors, which are used to support subsequent day-ahead price-output pair (PQ) strategy optimization learning.
[0055] In this embodiment of the application, the step of inputting the pricing strategy action into a pre-built market simulation environment and determining the expected return corresponding to the pricing strategy action includes: The preprocessed historical market data sequence is input into a preset probability prediction model to obtain the probability distribution information of electricity prices; Based on the day-ahead electricity price in the historical market data sequence, the probability distribution information of the electricity price, and the random sampling method, multiple day-ahead market electricity price scenarios and real-time market electricity price scenarios are generated, and each set of electricity price scenarios corresponds to a predicted electricity price; For each electricity price scenario, the bidding strategy action and the electricity price scenario are input into a preset market clearing model for evaluation to obtain the winning bid volume of the bidding strategy action under the electricity price scenario; For each electricity price scenario, the revenue under the electricity price scenario is determined based on the winning bid volume, the day-ahead electricity price and real-time electricity price corresponding to the electricity price scenario, and the preset revenue calculation model. Based on the revenue under each of the aforementioned electricity price scenarios, the expected revenue corresponding to the pricing strategy action is determined.
[0056] In the above implementation, a simplified day-ahead market clearing model is constructed based on the output of the electricity price forecasting model to simulate the relationship between "bid price - winning bid volume". The forecasted electricity price is regarded as a volatile quantity. Based on the expected value μ of the forecasted electricity price and the uncertainty parameter σ, multiple possible electricity price scenarios are generated through random sampling to reflect the impact of forecast error on the clearing result.
[0057] Specifically, the step of inputting the bidding strategy action and the electricity price scenario into a preset market clearing model for evaluation, and obtaining the winning bid volume of the bidding strategy action under the electricity price scenario, includes: The bid prices for each segment of the bidding strategy action are compared with the predicted electricity prices under the electricity price scenario to obtain the comparison results; Multiple winning segments were determined based on the comparison results; Based on the declared electricity volume of each of the winning bid segments, the winning bid volume of the bidding strategy action under the electricity price scenario is determined; The determination of multiple winning segments based on comparison results includes: For any segment in the bidding strategy action, if the bid price of the segment is not higher than the predicted electricity price under the electricity price scenario, then the segment is determined to be the winning segment. For any segment in the bidding strategy action, if the bid price of the segment is higher than the predicted electricity price under the electricity price scenario, then the segment is determined not to be the winning segment.
[0058] In the above implementation, it is assumed that the market adopts a marginal clearing mechanism, and the segmented declarations submitted by the unit in each time period i before the day are... If the declared price for this segment The predicted electricity price shall not exceed that under a certain sampling scenario s. That is, satisfying: ; If the bid price for that segment is considered successful, the corresponding declared electricity volume will be included in the unit's winning bid volume. If the bid price is higher than the predicted electricity price under this scenario, the segment will not be accepted. The total winning bid volume of the unit in time period i corresponds to the accepted bid segment volume.
[0059] The winning bid volumes obtained under different electricity price scenarios are pre-constructed into the revenue calculation model of thermal power units in the spot market to obtain the revenue of multiple electricity price scenarios corresponding to the bidding strategy.
[0060] Furthermore, for real-time market revenue, based on the obtained real-time electricity price forecast and its uncertainty parameters, the real-time electricity price scenario is obtained by combining it with the unit's deviation electricity volume in the real-time phase, using a random sampling method similar to the day-ahead market, and the corresponding real-time deviation revenue is calculated. This revenue, together with the day-ahead market revenue, constitutes the total revenue of the revenue calculation model, providing complete reward feedback for the reinforcement learning strategy.
[0061] By inputting preprocessed historical market data into a probabilistic prediction model to obtain electricity price probability distribution information, and then combining this with random sampling to generate multiple electricity price scenarios, various electricity price fluctuation scenarios can be simulated, taking into account market uncertainty. Under each electricity price scenario, the bidding strategy actions are evaluated to obtain the winning bid volume, and then the revenue under each scenario is determined by combining the revenue calculation model, ultimately obtaining the expected revenue. Based on electricity price probability prediction and randomized market clearing mechanism, this method can simulate the winning bid results and revenue performance of thermal power units under different electricity price fluctuation scenarios. By using deep reinforcement learning algorithms to generate bidding strategies with stable revenue performance in multiple scenarios, the profitability and bidding robustness of thermal power units in volatile electricity markets can be improved.
[0062] This invention fully leverages the price volatility characteristics of the electricity market. By comprehensively considering day-ahead price forecasts, real-time deviation revenue, and market uncertainties, it generates adaptive pricing strategies through deep reinforcement learning algorithms, achieving revenue optimization and improved strategy robustness under multiple scenarios. Thermal power units can simultaneously balance revenue, risk, and execution requirements in a dynamic market environment, more effectively coping with uncertainties caused by forecast errors, electricity price fluctuations, and load changes, thereby improving economic efficiency and competitiveness in complex markets.
[0063] Figure 2 A schematic diagram of a thermal power unit market quotation strategy generation device 200 is shown.
[0064] like Figure 2 As shown, a market quotation and pricing strategy generation device 200 for thermal power units mainly includes: The acquisition module 201 is used to acquire preprocessed current market data, which includes the current day-ahead forecast electricity price, the current real-time forecast electricity price, the current fuel cost data, unit operation constraints, and medium- and long-term contract information, which includes contracted electricity volume and contracted electricity price. Construction module 202 is used to construct the current state space based on the preprocessed current market data. The current state space includes market environment prediction information and the unit's own state parameters. The volume and price quotation strategy generation module 203 is used to input the current state space into a pre-trained strategy network to obtain the optimal volume and price quotation strategy curve. The pre-trained policy network is trained based on historical market data through a reinforcement learning framework. In the reinforcement learning framework, thermal power units are modeled as agents, and the market clearing process and revenue calculation are modeled as the environment. Through continuous interaction between the agent and the environment, the parameters of the policy network are optimized to maximize the cumulative reward, thus obtaining the pre-trained policy network.
[0065] In one example, the module in any of the above devices may be one or more integrated circuits configured to implement the above methods, such as one or more application-specific integrated circuits (ASICs), or one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs), or a combination of at least two of these integrated circuit forms.
[0066] For example, when modules in a device can be implemented via a processing element scheduler, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processor capable of calling programs. Alternatively, these modules can be integrated together as a system-on-a-chip (SOC).
[0067] In this application, various objects such as messages / information / devices / network elements / systems / apparatus / actions / operations / processes / concepts may be named. It is understood that these specific names do not constitute a limitation on the relevant objects. The names may be changed depending on the scenario, context, or usage habits. The understanding of the technical meaning of the technical terms in this application should be mainly determined from their functions and technical effects embodied / performed in the technical solution.
[0068] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and modules described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0069] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0070] The terms “comprising,” “including,” or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0071] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the foregoing application concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions claimed in this application.
Claims
1. A method for generating a day-ahead market quotation strategy for thermal power units, characterized in that, include: S1, Obtain preprocessed current market data, which includes current day-ahead forecast electricity price, current real-time forecast electricity price, current fuel cost data, unit operation constraints, and medium- and long-term contract information, which includes contracted electricity volume and contracted electricity price; S2, Based on the preprocessed current market data, construct the current state space, which includes market environment prediction information and the unit's own state parameters; S3, input the current state space into the pre-trained policy network to obtain the optimal volume quotation policy curve; The pre-trained policy network is trained based on historical market data through a reinforcement learning framework. In the reinforcement learning framework, thermal power units are modeled as agents, and the market clearing process and revenue calculation are modeled as the environment. Through continuous interaction between the agent and the environment, the parameters of the policy network are optimized to maximize the cumulative reward, thus obtaining the pre-trained policy network.
2. The method for generating a day-ahead market quotation strategy for thermal power units according to claim 1, characterized in that, The training method for the policy network includes: S11, Obtain the original historical market data and preprocess the original historical market data to obtain the preprocessed historical market data sequence; S12, retrieve the state space of the current iteration's historical day from the historical market data sequence in chronological order; S13, input the state space of the historical day into the current policy network to be trained, and generate a bidding policy action. The bidding policy action represents the segmented quantity bidding curve to be evaluated. The segmented quantity bidding curve includes the bid price and bid volume of each segment. S14, input the pricing strategy action into the pre-built market simulation environment, determine the expected return corresponding to the pricing strategy action and the next state space of the historical day's state space, and determine the reward space of the current iteration based on the expected return, wherein the reward space characterizes the degree to which the pricing strategy action obtains a good return. S15: Based on the reward space of the current iteration, the state space of the historical day, the bidding policy action, and the next state space of the state space of the historical day, the policy gradient is calculated using the reinforcement learning network to update the parameters of the current policy network. S16. Repeat steps S12 to S15 until the historical data sequence has been traversed or the performance index of the policy network has converged, and the trained policy network is obtained.
3. The method for generating a day-ahead market quotation strategy for thermal power units according to claim 2, characterized in that, The preprocessing of the original historical market data to obtain a preprocessed historical market data sequence includes: Outlier detection and removal, as well as missing value processing, are performed on the original historical market data to obtain the first historical market data. Align the data from different sources in the first historical market data according to the set timestamps to obtain the second historical market data; Based on the second historical market data, the constraint parameters of the thermal power unit operating boundary are extracted; Based on the constraint parameters of the thermal power unit's operating boundary, the second historical market data, and the preset time sequence, the preprocessed historical market data sequence is determined.
4. The method for generating a day-ahead market quotation strategy for thermal power units according to claim 2, characterized in that, The step of inputting the pricing strategy action into a pre-built market simulation environment and determining the expected return corresponding to the pricing strategy action includes: The preprocessed historical market data sequence is input into a preset probability prediction model to obtain the probability distribution information of electricity prices; Based on the day-ahead electricity price in the historical market data sequence, the probability distribution information of the electricity price, and the random sampling method, multiple day-ahead market electricity price scenarios and real-time market electricity price scenarios are generated, and each set of electricity price scenarios corresponds to a predicted electricity price; For each electricity price scenario, the bidding strategy action and the electricity price scenario are input into a preset market clearing model for evaluation to obtain the winning bid volume of the bidding strategy action under the electricity price scenario; For each electricity price scenario, the revenue under the electricity price scenario is determined based on the winning bid volume, the day-ahead electricity price and real-time electricity price corresponding to the electricity price scenario, and the preset revenue calculation model. Based on the revenue under each of the aforementioned electricity price scenarios, the expected revenue corresponding to the pricing strategy action is determined.
5. The method for generating a day-ahead market quotation strategy for thermal power units according to claim 4, characterized in that, The step of inputting the bidding strategy action and the electricity price scenario into a preset market clearing model for evaluation, and obtaining the winning bid volume of the bidding strategy action under the electricity price scenario, includes: The bid prices for each segment of the bidding strategy action are compared with the predicted electricity prices under the electricity price scenario to obtain the comparison results; Multiple winning segments were determined based on the comparison results; Based on the declared electricity volume of each of the winning bid segments, the winning bid volume of the bidding strategy action under the electricity price scenario is determined; The determination of multiple winning segments based on comparison results includes: For any segment in the bidding strategy action, if the bid price of the segment is not higher than the predicted electricity price under the electricity price scenario, then the segment is determined to be the winning segment. For any segment in the bidding strategy action, if the bid price of the segment is higher than the predicted electricity price under the electricity price scenario, then the segment is determined not to be the winning segment.
6. The method for generating a day-ahead market quotation strategy for thermal power units according to claim 4, characterized in that, The preset revenue calculation model is as follows: ; Where i is the time period index; n is the number of time periods divided into that day; , These represent the medium- and long-term contract electricity price and contract electricity volume corresponding to the i-th time period, respectively. , These represent the clearing price and the electricity volume declared or won by generating units in the i-th period of the day-ahead market, respectively. , This represents the settlement price and actual output or power volume of the electricity in the i-th time period of the real-time market. This represents the average marginal generation cost of the thermal power unit in the i-th time period, which is calculated based on the fuel unit price, unit thermal efficiency, and auxiliary power consumption.
7. The method for generating a day-ahead market quotation strategy for thermal power units according to claim 4, characterized in that, The preset probability prediction model is a Monte Carlo simulation model or an XGBoost machine learning model, and the method for constructing the preset probability prediction model is as follows: The Monte Carlo simulation model or XGBoost machine learning model is trained based on historical cleared electricity prices, cleared electricity volume and load demand data to obtain a trained probabilistic prediction model.
8. The method for generating a day-ahead market quotation strategy for thermal power units according to claim 2, characterized in that, The formula for calculating the reward space is: ; in, , , Configurable weight parameters; Used to measure the degree to which the target of maintaining quantity is achieved; This indicator is used to describe the risk of bids deviating from predicted electricity prices, bids being too high leading to non-transactions, or price fluctuations. To predict electricity prices.
9. A method for generating a day-ahead market quotation strategy for thermal power units according to claim 2, characterized in that, The reinforcement learning network is a deep reinforcement learning network, which includes a policy network, a value network, and a target network. The policy network generates bid-ask policies and actions based on the input state space. The policy network employs a multilayer perceptron structure and integrates an attention mechanism to generate segmented bid-ask curves that satisfy monotonically increasing constraints. The value network evaluates the long-term value of state-action pairs. The value network employs a deep network structure with dual-path encoding of states and actions. The target network includes a policy target network and a value target network. The target network updates the parameters of the policy network and the value network based on a soft update mechanism.
10. A device for generating a day-ahead market quotation strategy for thermal power units, characterized in that, include: The acquisition module is used to acquire preprocessed current market data, which includes the current day-ahead forecast electricity price, the current real-time forecast electricity price, current fuel cost data, unit operation constraints, and medium- and long-term contract information, which includes contracted electricity volume and contracted electricity price. The construction module is used to construct the current state space based on the preprocessed current market data. The current state space includes market environment prediction information and the unit's own state parameters. The volume reporting and pricing strategy generation module is used to input the current state space into a pre-trained strategy network to obtain the optimal volume reporting and pricing strategy curve. The pre-trained policy network is trained based on historical market data through a reinforcement learning framework. In the reinforcement learning framework, thermal power units are modeled as agents, and the market clearing process and revenue calculation are modeled as the environment. Through continuous interaction between the agent and the environment, the parameters of the policy network are optimized to maximize the cumulative reward, thus obtaining the pre-trained policy network.