Optimal dispatch method for combined operation of wind power and pumped storage system

The decision-making model constructed by multi-agent reinforcement learning and deep neural networks solves the problems of model dependence, computational complexity and decision lag in the operation optimization and scheduling of wind power-pumped storage combined systems. It realizes dynamic adaptive decision-making, improves system stability and energy utilization, reduces power generation costs and enhances the system's risk resistance.

CN121965823BActive Publication Date: 2026-07-21内蒙古电力(集团)有限责任公司内蒙古电力经济技术研究院分公司
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
内蒙古电力(集团)有限责任公司内蒙古电力经济技术研究院分公司
Filing Date
2026-04-03
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing technologies for the operation optimization and scheduling of wind power-pumped storage combined systems suffer from problems such as strong model dependence, high computational complexity, decision lag and static nature, rigid risk handling, and difficulty in collaborative optimization. These problems result in uneven equipment utilization, low energy utilization, high wind curtailment rate, high power generation cost, poor system stability, and weak risk resistance.

Method used

A decision-making model is constructed using multi-agent reinforcement learning and deep neural networks. Real-time data and risk preference coefficients are used to achieve dynamic adaptive decision-making. Wind turbines and pumped storage power stations are modeled as independent decision-making agents through multi-agent reinforcement learning. Combined with deep neural networks for online learning, they autonomously discover the optimal collaborative decision-making strategy. Power grid models and constraints are introduced to ensure the physical feasibility and market accuracy of the decision.

Benefits of technology

It enables efficient and superior decision-making in complex and ever-changing market environments, improves the efficiency and accuracy of electricity market application results, reduces power generation costs, increases power generation efficiency, enhances system stability and equipment utilization, and strengthens the ability to respond to sudden loads and resist risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121965823B_ABST
    Figure CN121965823B_ABST
Patent Text Reader

Abstract

The application provides an optimization scheduling method of a wind power and pumped storage combined operation system, and the method comprises the following steps: obtaining a decision value set comprising an electricity market bid quantity of a power generation unit, a frequency regulation capacity bid quantity and a frequency regulation mileage bid quantity based on a real-time data set and a risk preference coefficient, and scheduling the combined operation system; a target function of the decision model aims to minimize the total power purchase cost of the electricity market and the frequency regulation auxiliary service market; the decision model is constructed by using multi-agent reinforcement learning and a deep neural network based on a historical data set, a power market model and a combined operation system model. The method has dynamic adaptability and strong robustness, is beneficial to improving the declaration efficiency and result accuracy of the power market, can make various types of energy be cooperatively scheduled, improve the system stability, equipment utilization and energy utilization, reduce the power generation cost and improve the power generation benefit, and improve the response capability to sudden load growth and the anti-risk capability of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power system operation and control technology, specifically to an optimized scheduling method for a wind power and pumped storage combined operation system. Background Technology

[0002] Currently, research on operational optimization and scheduling of wind power-pumped storage or other types of "new energy-energy storage" combined systems, based on participation in online decision-making in the electricity market, mainly focuses on traditional optimization methods based on mathematical programming. These methods typically construct a two-level optimization model or a stochastic programming / distributed bar optimization model.

[0003] In a typical bilevel optimization model, the upper-level model aims to maximize the total revenue of the joint system and determines its online electricity price decision strategy in the market; the lower-level model simulates the market clearing process by which market operators aim to minimize the overall electricity purchase cost for society. To solve this bilevel optimization model, researchers typically use the Karush-Kuhn-Tucker conditions or strong duality theory to transform the bilevel problem into a single-level mixed-integer linear programming problem. To handle the uncertainty of wind power output, existing technologies usually introduce risk metrics such as conditional value at risk (VaR) as part of the objective function or constraints to quantify and manage market risk.

[0004] However, these existing technologies have several inherent limitations: 1. Strong Model Dependency and High Complexity: This method heavily relies on accurate mathematical models of market clearing mechanisms, power flow, and equipment operating characteristics. Constructing these models is a significant challenge in itself, and any simplification or bias in the assumptions can lead to deviations from optimal decision results. Transforming the two-layer model into a solvable single-layer model is complex, and the computational load increases dramatically with system size.

[0005] 2. Decision-making lag and static nature: Optimization-based methods are typically used to formulate "day-ahead" plans. Their solution process is time-consuming and cannot adapt to rapid price fluctuations in intraday or real-time markets. The resulting strategies are essentially static plans based on forecast information, lacking the ability to adapt online.

[0006] 3. Rigid Risk Management: Although methods such as CVaR (Conditional Value at Risk) have been introduced to quantify risk, decision-makers' risk preferences are usually pre-set with a fixed coefficient. This static risk attitude cannot be flexibly adjusted according to dynamic changes in market conditions, resulting in poor strategy performance in volatile market environments.

[0007] 4. Difficulty in synergistic optimization: In traditional models, the complex coupling relationship between wind power and pumped storage needs to be expressed through strict constraints, which further increases the complexity of the model and the difficulty of solving it, making it difficult to efficiently explore the synergistic potential between the two.

[0008] The above limitations lead to low accuracy in the generated market declaration schemes. As a result, when the joint system is operated and scheduled, it will lead to uneven equipment utilization, low wind energy resource utilization, high wind curtailment rate, and low hydropower conversion efficiency. This results in both energy waste and economic losses, making it difficult to maximize the reduction of power generation costs and the improvement of power generation efficiency. At the same time, the lack of coordinated scheduling of various energy sources leads to a decrease in system stability. Furthermore, unreasonable operation and scheduling can also lead to insufficient backup units, making it impossible to cope with sudden load increases. As a result, the system stability is compromised and its risk resistance is weak. Summary of the Invention

[0009] This invention addresses the problems existing in the prior art by providing an optimized scheduling method for a combined wind power and pumped storage system. This method exhibits strong robustness against prediction errors and uncertainties in individual power generation units. It can instantly output the optimal strategy based on real-time market information and system status, possessing dynamic adaptability. Therefore, it maintains high efficiency and superiority in decision-making even in complex and ever-changing market environments. This improves the efficiency and accuracy of electricity market applications, maximizes the reduction of power generation costs and increases power generation efficiency, providing a competitive advantage for power plants. Furthermore, based on this method, the scheduling of the combined system enables coordinated scheduling of various energy sources, improving system stability, equipment utilization, and energy efficiency, maximizing the reduction of power generation costs and increasing power generation efficiency. It also enhances the system's ability to cope with sudden load increases and its resilience to risks.

[0010] To achieve the above objectives, the technical solution adopted by the present invention is as follows: an optimized scheduling method for a wind power and pumped storage combined operation system, comprising: Obtain real-time datasets and risk preference coefficients; the real-time dataset represents the set of input data for the real-time target period, including market information data, equipment status data, and forecast data; the risk preference coefficient represents the weighting factor between the returns and risks of the joint operation system; The decision value set is obtained by solving the decision model based on real-time dataset and risk preference coefficient; the decision value set includes the winning bid amount in the power market, the winning bid amount in the frequency regulation capacity, and the winning bid amount in the frequency regulation mileage for each power generation unit, and the power generation unit represents wind turbine or pumped storage power station; the joint operation system is scheduled according to the decision value set; The decision model includes an objective function that aims to minimize the total electricity purchase cost in the market, which includes the electricity market and the frequency regulation ancillary service market. The decision model is constructed using multi-agent reinforcement learning and deep neural networks based on historical datasets, electricity market models, and joint operation system models. The historical dataset represents the set of input data for the historical target period.

[0011] This application proposes an optimized scheduling method for a wind power and pumped storage combined operation system. This method can break free from the dependence on precise mathematical models and has online learning and adaptive capabilities. Since the risk preference coefficient is a key bridge connecting market uncertainty and the decision-making behavior of decision-making agents, and is an adjustable hyperparameter embedded in the reward function of multi-agent reinforcement learning, decision-making agents with different risk profiles can be "trained" by adjusting the risk preference coefficient, thereby achieving dynamic risk control.

[0012] In some embodiments, the total electricity purchase cost represents the linear superposition of a first total electricity purchase cost and a second total electricity purchase cost. The first total electricity purchase cost represents the total electricity purchase cost in the electricity market, and the second total electricity purchase cost represents the total electricity purchase cost in the frequency regulation ancillary service market. The objective function constructed based on this total electricity purchase cost provides the core of the environment simulator for subsequent multi-agent reinforcement learning and the basis for providing accurate reward signals for decision-making agents.

[0013] In some embodiments, the joint operation system model includes a power grid model, at least two decision agents, and at least one conventional unit model; the at least two decision agents include at least one wind turbine model and at least one pumped storage power station model. The wind turbine model is constructed based on the uncertainty of wind turbine output prediction; The pumped storage power station model is constructed based on the reservoir capacity status of the pumped storage power station, taking into account the continuous and discrete characteristics of the pumped storage power station's capacity and its efficiency.

[0014] In some embodiments, the decision model also includes constraints, including power balance and line transmission capacity constraints, unit operation constraints, and frequency regulation market constraints. Power balance and line transmission capacity constraints characterize the coupling constraints between node power load and inter-node line transmission capacity based on the power grid model. Unit operating constraints include internal control constraints of the joint operation system and state constraints of the pumped storage power station; The frequency regulation market constraint characterizes the coupling constraint between demand and clearing based on the frequency regulation ancillary services market, at least one wind turbine model, and at least one pumped storage power station model.

[0015] The decision-making agent learns decision-making strategies that naturally meet the physical constraints of power grid and unit operation, avoiding violations in actual operation, while accurately reflecting the actual value of frequency regulation services, improving resource allocation efficiency, and providing accurate price signals for market participants.

[0016] In some embodiments, the internal control constraints of the joint operation system are characterized by the control coupling constraints based on the electricity market, at least one wind turbine model, and at least one pumped storage power station model. By modeling the internal control constraints of the joint operation system, wind turbines can charge pumped storage power stations during periods of low prices or high wind curtailment risk, achieving energy time-shifting, improving overall profitability. The internal charging amount can be dynamically adjusted according to market prices, enabling cross-period arbitrage and flexible response to market signals. Furthermore, energy storage at pumped storage power stations reduces wind curtailment, improves renewable energy utilization, and thus enhances wind power absorption capacity.

[0017] In some embodiments, the pumped storage power station state constraints include pumped storage power station operating condition constraints, pumped storage power station start-stop state switching constraints, and pumped storage power station minimum start-stop time constraints. The operational constraints of pumped storage power stations are characterized by the coupling constraints between the charging and discharging power and the charging and discharging state of at least one pumped storage power station model based on the electric energy market. This provides a feasibility guarantee for the final decision and guides the decision-making agent to learn compliant strategies. The start-stop state switching constraint of pumped storage power station characterizes the start-stop state constraints of at least one pumped storage power station model under power generation and pumping conditions, respectively, to avoid the additional costs caused by frequent start-stop, reduce unnecessary start-stop times, and extend equipment life. The minimum start-up and shutdown time constraint of a pumped storage power station characterizes the coupling constraints between the minimum continuous power generation time, minimum continuous shutdown time, power generation operation state, and pumping operation state of at least one pumped storage power station model. By modeling the minimum start-up and shutdown time constraint of the pumped storage power station, the decision agent learns to plan long-term operation strategies and avoids unnecessary losses caused by frequent start-ups and shutdowns.

[0018] In some of these embodiments, risk preference coefficients are obtained when constructing decision models using multi-agent reinforcement learning and deep neural networks based on historical datasets, electricity market models, and joint operation system models. The overall risk preference coefficients are more realistic and have dynamic adjustment capabilities.

[0019] In some embodiments, a decision-making model is constructed using multi-agent reinforcement learning and deep neural networks based on historical datasets, electricity market models, and joint operation system models, including: Obtain initialization parameters and initial risk coefficients; divide the historical dataset to obtain training and test sets; Training models are constructed based on deep neural networks, initialization parameters, and at least two decision agents. Each training model includes a set of Actor networks; the set of Actor networks includes at least two Actor networks; and each of the at least two Actor networks corresponds one-to-one with at least two decision agents. A first cache set is obtained based on the training set, the electricity market model, the joint operation system model, and the Actor network set of the training model. The first cache set represents the set of experience tuples for each decision agent in each unit time period during the historical target period, obtained based on the training set. Each experience tuple includes the current state, output action, return reward, and next state corresponding to the decision agent. The current state represents the state information of the corresponding decision agent in the current unit time period, obtained based on the training set, the electricity market model, and the joint operation system model. The output action represents the action information obtained by the corresponding Actor network based on the current state. The return reward represents the reward information of the corresponding decision agent obtained based on the simulation clearing result. The simulation clearing result represents the clearing information obtained by the electricity market model based on the output action. The next state represents the state information of the next unit time period obtained by the joint operation system model based on the output action. Update the Actor network set, Critic network set, and target network set of the trained model according to the first cache set; The first average total reward for each decision agent is obtained based on the returned reward. The first average total reward represents the average total reward of the decision agent in each historical target period based on the training set. When the first average total revenue does not converge, repeat the process of obtaining the first cache set based on the Actor network set of the training set, the electricity market model, the joint operation system model, and the training model; When the convergence value of the first average total return is higher than the first preset value, the test model is obtained based on the training model corresponding to the highest first average total return. An initial risk coefficient is introduced into the Actor network set of the test model and the Actor network set of the test model is updated. A second cache set is obtained based on the test set, the electricity market model, the joint operation system model and the Actor network set of the test model. The second cache set represents the set of experience tuples of each decision agent in each unit time period based on the test set for the historical target time period. Update the initial risk coefficients and the Actor network set of the test model based on the second cache set; The second average total reward for each decision agent is obtained based on the returned reward. The second average total reward represents the average total reward of the decision agent in each historical target period based on the test set. When the second average total revenue does not converge, repeat the process of obtaining the second cache set based on the Actor network set of the test set, the electricity market model, the joint operation system model, and the test model. When the second average total return converges to a value higher than the second preset value, a decision model is obtained based on the Actor network set, and a risk preference coefficient is obtained based on the updated initial risk coefficient.

[0020] In some embodiments, the equipment status data includes the actual load curve, the actual output of the wind turbine, and the wind speed.

[0021] In some embodiments, the forecast data includes forecast load curves and forecast wind turbine output.

[0022] Compared with the prior art, the present invention has the following beneficial effects: 1. The optimized scheduling method for a wind power and pumped storage combined operation system proposed in this application employs historical data, multi-agent reinforcement learning, and deep neural networks to reinforce the decision-making agent. This enables the decision-making model to autonomously learn how to make the optimal decision under any given state in a simulated environment. This not only makes the decision-making model highly robust to prediction errors and uncertainties of individual power generation units, but also allows the model to instantly output the optimal strategy based on real-time market information and system status. Furthermore, it incorporates the ability to autonomously balance returns and risks, giving the decision-making model dynamic adaptability that traditional methods cannot achieve. Thus, it maintains the efficiency and superiority of decision-making in complex and ever-changing market environments, thereby optimizing the scheduling results.

[0023] 2. The optimized scheduling method for a combined wind power and pumped storage system proposed in this application effectively solves the problem of balancing "collaborative optimization" and "real-time application." On the one hand, through a "centralized training" mechanism, the decision-making agents deeply understand the impact of each other's strategies on the global strategy, achieving a globally optimal strategy that surpasses local optima. On the other hand, once training is complete, online decision-making only requires millisecond-level forward network computation to generate a decision value set, meeting the requirements of the electricity spot market for online decision-making response speed, thereby optimizing the scheduling of the combined operation system. This characteristic of combining deep collaborative optimization capabilities with extreme decision-making efficiency, combined with an adjustable risk preference coefficient, enables the method proposed in this application not only to obtain the optimal decision strategy but also to conform to the authenticity of electricity market transactions. This is conducive to improving the reporting efficiency and accuracy of electricity market results, maximizing the reduction of power generation costs and the improvement of power generation efficiency, providing a competitive advantage for power plants. At the same time, based on this, the scheduling of the combined system can enable the coordinated scheduling of various energy sources, improve system stability, and increase equipment and energy utilization rates, maximizing the reduction of power generation costs and the improvement of power generation efficiency, while also improving the system's ability to cope with sudden load increases and its risk resistance. Attached Figure Description

[0024] Figure 1 This is a flowchart illustrating the optimized scheduling method for the combined wind power and pumped storage system of this application. Figure 2 This is the daily load curve of the jointly operating system in the example of this application; Figure 3 The predicted output curves of wind turbine W1 in 10 typical scenarios in the examples of this application are shown. Figure 4 The predicted output curves of wind turbine W2 in 10 typical scenarios in the examples of this application are shown. Figure 5 This is the bidding situation in the electricity market when considering the risk preference coefficient in the calculation examples of this application; Figure 6 This is the bidding situation for frequency regulation capacity when considering the risk preference coefficient in the calculation examples of this application; Figure 7 This example illustrates the winning bid for frequency regulation mileage when considering risk preference coefficients in the calculations presented in this application. Figure 8 This is a graph showing the percentage of electricity won in the electricity market when considering the risk preference coefficient in the examples of this application; Figure 9 This is a graph showing the percentage of electricity won in the electricity market without considering the risk preference coefficient in the examples provided in this application. Figure 10 This paper presents the daily revenue of pumped storage power station E1 under two scenarios: considering the risk preference coefficient and not considering the risk preference coefficient, in the calculation examples of this application. Detailed Implementation

[0025] Based on the problems raised in the background technology, traditional methods often struggle to find optimal or even feasible solutions when facing the strong uncertainty of wind power output and complex market coupling relationships due to model mismatch and high computational complexity. To fundamentally address the bottlenecks faced by the aforementioned background technologies, this application overturns the traditional static optimization framework that relies on precise mathematical models, pioneering a new path of data-driven intelligent decision-making.

[0026] To clearly illustrate the technical features of this solution, the implementation methods of this application will be described in detail below with reference to the accompanying drawings and embodiments. This will allow for a full understanding and implementation of how this application uses technical means to solve technical problems and achieve corresponding technical effects. The embodiments of this application and the various features within them can be combined with each other without conflict, and the resulting technical solutions are all within the protection scope of this application.

[0027] To achieve the above objectives, the technical solution proposed in this application is as follows: See Figure 1This application proposes an optimized scheduling method for a wind power and pumped storage combined operation system, including: Obtain real-time datasets and risk preference coefficients; real-time datasets represent the set of input data for the real-time target period, including market information data, equipment status data, and forecast data. Market information data includes at least the power generation price declared in the electricity market and the frequency regulation capacity price and frequency regulation mileage price declared in the frequency regulation ancillary service market. Equipment status data includes at least the actual load curve, actual wind turbine output, and wind speed. Forecast data includes at least the forecast load curve and forecast wind turbine output. Risk preference coefficients represent the weighting factors of the revenue and risk of the joint operation system.

[0028] Optionally, to fully consider the impact of price volatility risk on the operation of pumped storage power stations, a Long Short-Term Memory (LSTM) network model from the field of machine learning is used to accurately predict spot market price fluctuations based on market information data in historical datasets. To quantify price volatility risk, the Conditional Value at Risk (CVaR) method is used to obtain a risk preference coefficient, achieving an effective measurement of risk. Furthermore, the adjustability of the risk preference coefficient is achieved through adjusting confidence levels, optimizing the loss function, introducing dynamic weights, combining multi-objective optimization, applying scenario analysis and stress testing, and utilizing big data and machine learning.

[0029] The decision value set is obtained by solving the decision model based on real-time datasets and risk preference coefficients. The decision value set includes the winning bid amount of electricity in the market, the winning bid amount of frequency regulation capacity, and the winning bid amount of frequency regulation mileage for each power generation unit. The power generation unit represents wind turbine or pumped storage power station. In addition, the decision value set also includes the power generation price, frequency regulation capacity price, and frequency regulation mileage price declared by the power generation unit. Among them, the power generation price declared by the pumped storage power station includes the discharge price and the charging price. The joint operation system is scheduled based on the decision value set; The decision-making model includes an objective function that aims to minimize the total cost of electricity purchase in the market. The market includes the electricity market and the frequency regulation ancillary service market. Traditional market clearing models usually only consider the electricity market and ignore the frequency regulation ancillary service market, which makes it impossible to accurately reflect the multiple values ​​of individual power generation units. Even if the frequency regulation ancillary service market is considered, electricity is often optimized separately from frequency regulation capacity and frequency regulation mileage, which makes it impossible to achieve joint clearing and leads to suboptimal solutions.

[0030] When existing technologies model with the goal of minimizing electricity purchase costs, they lack refined modeling of frequency regulation capacity and frequency regulation mileage. In some embodiments, the total electricity purchase cost represents the linear superposition of the first total electricity purchase cost and the second total electricity purchase cost to achieve a unified optimization objective. The first total electricity purchase cost represents the total electricity purchase cost in the electricity market, and the second total electricity purchase cost represents the total electricity purchase cost in the frequency regulation ancillary service market.

[0031] The objective function is: ; In the formula, To minimize the function, For unit time period sequence number, The total number of units within a time period. This refers to the serial number of the pumped storage power station. This represents the total number of pumped storage power stations. For the first The unit time period The discharge power of a pumped storage power station that won a bid in the electricity market. For the first The unit time period The charging power of a pumped storage power station that won a bid in the electricity market. For the first The unit time period The winning bid amount for the frequency regulation capacity of a pumped storage power station. For the first The unit time period The winning bid for the frequency regulation mileage of a pumped storage power station. For the first The unit time period The discharge price declared by a pumped storage power station in the electricity market. For the first The unit time period The charging electricity price declared by each pumped storage power station in the electricity market. For the first The unit time period The price of frequency regulation capacity for a pumped storage power station For the first The unit time period Price of frequency regulation mileage for a pumped storage power station This refers to the serial number of the wind turbine unit. This represents the total number of wind turbine units. For the first The unit time period The volume of electricity generated by each wind turbine unit in the market. For the first The unit time period The price of electricity generated by a single wind turbine unit This is the serial number of a conventional generating unit. This represents the total number of conventional generating units. For the first The unit time period The number of conventional generating units that won bids in the electricity market. For the first The unit time period The winning bid amount for the frequency regulation capacity of a conventional generating unit For the first The unit time period The winning bid amount for the frequency regulation mileage of a conventional unit For the first The unit time period The electricity generation price of a conventional generating unit For the first The unit time period Price of frequency regulation capacity for a conventional generating unit For the first The unit time period Price of frequency regulation mileage for a conventional generating unit.

[0032] When modeling the objective function, three bidding variables are introduced: the winning bid amount in the electricity market, the winning bid amount for frequency regulation capacity, and the winning bid amount for frequency regulation mileage. This achieves joint optimization of the electricity market and the frequency regulation ancillary service market. Furthermore, three bidding variables are simultaneously introduced: the generation price, the frequency regulation capacity price, and the frequency regulation mileage price. This allows individual generators to differentiate their bids, reflecting their marginal costs in different markets. The objective function obtained through this modeling can simultaneously optimize electricity resources with high-frequency capacity and frequency regulation mileage resources, avoiding resource mismatch caused by separate optimization, resulting in more accurate market clearing. Individual generators can differentiate their bids in different markets based on their own technical characteristics, improving market efficiency. This objective function provides the core of the environment simulator for subsequent multi-agent reinforcement learning and the foundation for accurate reward signals for decision-making agents.

[0033] The decision-making model, based on historical datasets, a power market model, and a joint operation system model, is constructed using multi-agent reinforcement learning and deep neural networks. The historical dataset represents the set of input data for the historical target period. Through multi-agent reinforcement learning, wind turbines and pumped-storage power stations are modeled as independent decision-making agents. Deep neural networks then autonomously discover the optimal collaborative decision-making strategy through trial and error in a shared market environment. Each decision-making agent makes decisions based on its own local observations, and the state space comprehensively encompasses the market environment, forecast information, and physical operating status.

[0034] The decision-making model proposed in this application relies on global information to guide the strategy updates of each decision-making agent during the training phase, enabling them to learn to collaborate. During the execution phase, each agent can independently and quickly make optimal decisions based solely on local information, meeting the real-time requirements of market bidding and thus obtaining the preferred scheduling strategy. Simultaneously, a risk preference coefficient is introduced during the execution phase. When making each decision, the pumped storage power station considers the current actual state and expectations of future market conditions to maximize benefits and effectively manage risks. This application proposes an optimized scheduling method for a wind power and pumped storage combined operation system. This method is able to break free from dependence on precise mathematical models and possesses online learning and adaptive capabilities. Since the risk preference coefficient is a key bridge connecting market uncertainty and the decision-making behavior of decision-making agents, and is an adjustable hyperparameter embedded in the reward function of multi-agent reinforcement learning, adjusting the risk preference coefficient can "train" decision-making agents with different risk profiles, thereby achieving dynamic risk control.

[0035] In some embodiments, the joint operation system model includes a power grid model, at least two decision agents, and at least one conventional generator unit model, consistent with real-world joint operation systems. The power grid model is based on a DC optimal power flow model, which achieves a good balance between computational efficiency and accuracy. Furthermore, the DC optimal power flow model is suitable for large-scale training in multi-agent reinforcement learning environments and accurately reflects the power flow distribution characteristics of the power grid. The at least two decision agents include at least one wind turbine generator unit model and at least one pumped storage power station model. The wind turbine model is constructed based on the uncertainty of wind turbine output prediction. Optionally, the wind turbine output prediction is obtained using Monte Carlo method and k-means clustering algorithm. The pumped storage power station model is constructed based on the reservoir capacity status of the pumped storage power station.

[0036] Optionally, the reservoir capacity state expression for the pumped storage power station model is: ; In the formula, For the first The unit time period The remaining capacity of the pumped storage power station For the first The minimum allowable capacity of a pumped storage power station For the first The maximum allowable capacity of a pumped storage power station. For the first The charging efficiency of a pumped storage power station. For the first The discharge efficiency of a pumped storage power station. For the first The final capacity of a pumped storage power station For the first The initial capacity of a pumped storage power station; Problems with existing technologies: Traditional models often simplify the capacity of pumped storage power stations as a continuous variable, ignoring its discrete characteristics and efficiency losses. Capacity boundary constraints, such as maximum / minimum capacity, are often neglected, making strategies infeasible in actual operation. Furthermore, the correlation between initial and final capacity in traditional models is often simplified, failing to achieve optimization for multi-day continuous operation. This paper introduces a capacity dynamic equation into the pumped storage power station model to clarify the relationship between capacity changes and charging / discharging power and efficiency, achieving energy conservation, accurately reflecting energy conversion losses, and improving strategy economy. It also introduces capacity boundary constraints to ensure that the capacity does not exceed physical limits, avoiding strategy infeasibility. Finally, it introduces the correlation between initial and final capacity to achieve collaborative optimization for multi-day continuous operation. Through the modeling of the pumped storage power station, the strategies learned by the decision-making agent naturally satisfy capacity boundary constraints, avoiding overcharging or over-discharging in actual operation, thus ensuring strategy feasibility. Moreover, the correlation between initial and final capacity enables the decision-making agent to plan multi-day operation strategies, achieving cross-day arbitrage through multi-day collaborative optimization.

[0037] In some embodiments, the decision model also includes constraints, including power balance and line transmission capacity constraints, unit operation constraints, and frequency regulation market constraints. Power balance and line transmission capacity constraints characterize the coupling constraints between nodal power loads and inter-node line transmission capacities based on power grid models. Traditional decision-making strategy studies often neglect grid-related physical constraints, assuming that all power generation resources can be transmitted without limit, making the strategies infeasible in actual power grids. Even when considering grid-related physical constraints, simplified models, such as single-node models, are often used, failing to reflect the spatial differences in marginal electricity prices at nodes. The power balance and line transmission capacity constraints are: ; ; In the formula, For the first A set of wind turbine units on each node This refers to the node number in the power grid. For the first A collection of conventional units on each node For the first A collection of pumped storage power stations on a node. For the first The total number of electrical loads on each node. For electrical load, In order to be with the first The set of all nodes connected to a given node. Let n be the electrical load at time t. for The sequence number of the middle node. For the first The node and the first Line admittance between nodes For the first The unit time period The phase angle of each node, For the first The unit time period The phase angle of each node, For the first The transmission capacity limit of each line, For line number, In order to be with the first All lines connected to each node. For the first The dual variables corresponding to the power balance constraints per unit time period For the first The unit time period The corresponding dual variable for the minimum line transmission capacity constraint of each line. For the first The unit time period The corresponding dual variable of the maximum line transmission capacity constraint of each line; Introducing nodal power balance equations into power balance and line transmission capacity constraints ensures that the injected and outflowing power of each node is balanced, reflecting the physical laws of the power grid. The decision-making agent's learned decision-making strategies naturally satisfy the physical constraints of the power grid, avoiding violations in actual operation. Introducing line transmission capacity constraints limits the power flow of each line to its thermal stability limit, avoiding overload. Introducing dual variables provides marginal value signals for the decision-making agent, enabling it to engage in spatial arbitrage and guiding its decision-making strategies.

[0038] Unit operation constraints include internal control constraints of the joint operation system and state constraints of pumped storage power stations. Traditional methods often treat wind power and energy storage as independent entities, ignoring the internal energy flow between them and failing to tap into their synergistic potential. Even when considering internal coupling, fixed ratios or simplified rules are often used, making it impossible to dynamically adjust according to market signals. Furthermore, the uncertainty of wind power forecast output is difficult to effectively reflect in the constraints.

[0039] In some embodiments, the internal control constraints of the joint operation system characterize the control coupling constraints based on the electricity market, at least one wind turbine model, and at least one pumped storage power station model. The internal control constraints of the joint operation system are: ; In the formula, For the first The unit time period The total charging volume of a pumped storage power station in the electricity market. For the first The unit time period The total discharge volume of pumped storage power stations in the electricity market. For the first The unit time period The wind turbine unit is directed towards the first The charging capacity of a pumped storage power station For the first The unit time period Predicted output of each wind turbine unit; By introducing an internal charging variable into the internal control constraints of the joint operation system, it is clarified that wind turbines can directly charge pumped storage power stations, realizing physical energy transfer. Simultaneously, an energy balance relationship is established, where the total output of the wind turbines equals the sum of the grid-connected electricity and the internal charging amount, ensuring energy conservation. Furthermore, a link is established between the total charging and discharging of the pumped storage power station and its internal charging amount, achieving an organic unity of internal coordination and market participation. Through modeling the internal control constraints of the joint operation system, wind turbines can charge pumped storage power stations during periods of low prices or high wind curtailment risk, achieving energy time-shifting and improving overall profitability. The internal charging amount can be dynamically adjusted according to market prices, enabling cross-time arbitrage and flexible response to market signals. Moreover, energy storage through pumped storage power stations reduces wind curtailment, improves the utilization rate of new energy sources, and thus enhances wind power absorption capacity.

[0040] In some embodiments, the pumped storage power station state constraints include pumped storage power station operating condition constraints, pumped storage power station start-stop state switching constraints, and pumped storage power station minimum start-stop time constraints. The operating condition constraints of a pumped storage power station characterize the coupling constraints between the charging and discharging power and the charging and discharging state based on the electricity market and at least one pumped storage power station model; the operating condition constraints of the pumped storage power station are: ; In the formula, For the first The unit time period The maximum charging power of a pumped storage power station For the first The unit time period The maximum discharge power of a pumped storage power station For the first The unit time period The charging power declared by each pumped storage power station in the electricity market. For the first The unit time period The discharge power declared by each pumped storage power station in the electricity market For the first The unit time period The charging status of a pumped storage power station For the first The unit time period The discharge status of a pumped storage power station , These are variables of 0 and 1, respectively. Traditional models often simplify the charging and discharging power of pumped storage power stations as continuous variables, ignoring the discrete operating characteristics where simultaneous charging and discharging is not possible. Furthermore, maximum power constraints are often statically set, failing to reflect actual operational limitations. In addition, introducing 0-1 variables into the traditional operating constraints of pumped storage power stations increases model complexity, making it difficult to solve efficiently using traditional methods.

[0041] In the pumped storage power station operating condition constraints of this application: 0-1 state variables are introduced to clearly distinguish charging and discharging conditions, ensuring that only one condition can be maintained at any given time; a maximum power constraint is introduced to limit the charging and discharging power from exceeding the equipment's rated value, reflecting physical limits; the operating conditions of the pumped storage power station are linked to market variables, meaning that the charging and discharging power declared by the pumped storage power station is subject to operating condition constraints, ensuring the feasibility of market quotations. This pumped storage power station operating condition constraint strategy allows the decision-making agent to learn strategies that naturally satisfy the physical operating characteristics of the pumped storage power station, avoiding violations in actual operation and providing feasibility assurance for the final decision; the introduction of 0-1 state variables for refined modeling accurately reflects the discrete operating condition characteristics of pumping and storage, improving the model's realism; simultaneously, the pumped storage power station operating condition constraints are transformed into penalty terms in multi-agent reinforcement learning, guiding the decision-making agent to learn compliant strategies.

[0042] The start-stop state switching constraint of a pumped storage power station characterizes the start-stop state constraints of at least one pumped storage power station model under power generation and pumping conditions, respectively; the start-stop state switching constraint of the pumped storage power station is: ; ; In the formula, For the first The unit time period The pumped storage power station is in the start-up state under power generation conditions. For the first The unit time period The pumped storage power station is currently in a shutdown state, despite being in operation for power generation. For the first The unit time period The power generation and operation status of a pumped storage power station, if This means , ,like This means , ; For the first The unit time period The pumped storage power station is in the start-up state under pumping operation. For the first The unit time period The pumped storage power station is currently in a shutdown state due to pumping operation. For the first The unit time period The pumping operation status of a pumped storage power station; Traditional models often neglect the start-up and shutdown costs and time constraints of pumped storage power stations. The final decision-making strategy leads to frequent start-ups and shutdowns of pumped storage power stations, which is impractical. Furthermore, the logical relationship between start-up and shutdown states is often ignored, such as the inability to start up and shut down simultaneously, resulting in an unreasonable model. At the same time, the start-up and shutdown costs are difficult to accurately quantify in the objective function.

[0043] By introducing start-stop state variables into the start-stop state switching constraints of pumped storage power stations, the logical relationship between start-up and shutdown is clarified. For example, startup cannot be immediately followed by shutdown, ensuring the rationality of state transitions. This establishes the logical constraints for the start-stop states of pumped storage power stations. Simultaneously, a link is established between start-stop variables and operating condition variables to achieve continuity in state transitions. Through modeling the start-stop state switching constraints of pumped storage power stations, the decision-making agent learns to weigh start-stop costs against market benefits, avoiding the additional costs caused by frequent start-stops, reducing unnecessary start-stop cycles, extending equipment lifespan, ensuring the rationality of the final decision strategy and improving economic efficiency. This ensures that the operating state transitions of pumped storage power stations conform to physical laws and avoids unrealistic decision strategies.

[0044] The minimum start-up and shutdown time constraint for a pumped storage power station characterizes the coupling constraint between the minimum continuous power generation time, minimum continuous shutdown time, power generation operation state, and pumping operation state of at least one pumped storage power station model; the minimum start-up and shutdown time constraint for a pumped storage power station is: ; In the formula, For the first The minimum continuous power generation time of a pumped storage power station For the first The minimum continuous pumping time for a pumped storage power station For the first Minimum continuous downtime of a pumped storage power station For the first The unit time period Power generation and operation status of pumped storage power stations For the first The unit time period The pumping operation status of a pumped storage power station; Traditional models often overlook the minimum start-up and shutdown time constraints of pumped storage power stations, leading to unreasonable frequent start-ups and shutdowns in the final decision-making strategy. Even when considering these constraints, complex integer programming constraints are often used, increasing the difficulty of solving the model. Furthermore, temporal constraints are difficult to embed directly into traditional reinforcement learning frameworks. This application introduces a minimum time constraint for the minimum start-up and shutdown time constraints of pumped storage power stations, requiring the station to operate continuously for a preset minimum time after startup and to remain offline for a preset minimum time after shutdown. It also constrains the temporal correlation between the current state and historical states to ensure the minimum time requirement is met. In addition, the minimum start-up and shutdown time constraints are transformed into a penalty term in multi-agent reinforcement learning, providing a negative reward for violating the minimum time constraint and guiding the decision-making agent to learn compliant strategies. By modeling the minimum start-up and shutdown time constraints of the pumped storage power station, the decision-making agent learns to plan long-term operation strategies, avoid unnecessary losses caused by frequent start-ups and shutdowns, ensure that the equipment has sufficient rest time, extend its service life, and at the same time, stable operation helps to build market credibility, improve medium- and long-term market competitiveness, and ensure the stability of market participation.

[0045] Frequency regulation market constraints characterize the coupling constraints between demand and clearing based on the frequency regulation ancillary services market, at least one wind turbine model, and at least one pumped storage power station model; frequency regulation market constraints are ; In the formula, For the first Frequency regulation capacity requirements of a combined wind power and pumped storage system per unit time period For the first Frequency regulation capacity clearing parameters for each unit time period For the first Frequency regulation mileage requirements of a combined wind power and pumped storage system per unit time period For the first Frequency regulation mileage clearing parameters for each unit time period; Some traditional FM market models neglect FM mileage demand, considering only FM capacity demand, thus failing to accurately reflect the actual value of FM services. Furthermore, in some traditional FM market models, the coupling relationship between FM capacity and FM mileage is often simplified, leading to inefficient resource allocation. In addition, the value of dual variables in traditional FM market models is underestimated, failing to provide accurate price signals for market participants.

[0046] This application, within the constraints of the frequency regulation market, simultaneously considers both frequency regulation capacity and mileage requirements, achieving refined modeling of the frequency regulation ancillary services market. It introduces dual variables to reflect the marginal value of frequency regulation capacity and mileage, providing price signals to market participants and guiding decision-making agents in the frequency regulation ancillary services market, thereby improving the accuracy of price signals and leading to optimized decision-making strategies. Furthermore, it correlates the mileage contribution of individual generators with their technical characteristics such as regulation rate and response time, realizing a performance-based market mechanism that rewards generators with better regulation performance with higher returns, incentivizing technological upgrades. Through modeling the constraints of the frequency regulation market, both frequency regulation capacity and mileage can be optimized simultaneously, avoiding resource misallocation, improving system frequency regulation efficiency, and enhancing system incentive performance.

[0047] In some embodiments, risk preference coefficients are obtained when constructing decision models using multi-agent reinforcement learning and deep neural networks based on historical datasets, electricity market models, and joint operation system models. These risk preference coefficients are derived from historical datasets, which can uncover hidden risk preference patterns in the data and better reflect the actual risk preferences of the decision agents. In practical applications, as new data is continuously added, the decision model can continuously update and optimize the risk preference coefficients to adapt to changes in the market environment.

[0048] In some embodiments, a decision-making model is constructed using multi-agent reinforcement learning and deep neural networks based on historical datasets, electricity market models, and joint operation system models, including: Obtain initialization parameters and initial risk coefficients; divide the historical dataset to obtain training and test sets; Training models are constructed based on deep neural networks, initialization parameters, and at least two decision agents. Each training model includes an Actor network set, a Critic network set, and a target network set. The Actor network set includes at least two Actor networks, the Critic network set includes at least two Critic networks, and the target network set includes at least two sets of target networks. Each of the at least two Actor networks, at least two Critic networks, and at least two sets of target networks corresponds one-to-one with at least two decision agents. A first cache set is obtained based on the training set, the electricity market model, the joint operation system model, and the Actor network set of the training model. The first cache set represents the set of experience tuples for each decision agent in each unit time period during the historical target period, obtained based on the training set. Each experience tuple includes the current state, output action, return reward, and next state corresponding to the decision agent. The current state represents the state information of the corresponding decision agent in the current unit time period, obtained based on the training set, the electricity market model, and the joint operation system model. The output action represents the action information obtained by the corresponding Actor network based on the current state. The return reward represents the reward information of the corresponding decision agent obtained based on the simulation clearing result. The simulation clearing result represents the clearing information obtained by the electricity market model based on the output action. The next state represents the state information of the next unit time period obtained by the joint operation system model based on the output action. Update the Actor network set, Critic network set, and target network set of the trained model according to the first cache set; The first average total reward for each decision agent is obtained based on the returned reward. The first average total reward represents the average total reward of the decision agent in each historical target period based on the training set. When the first average total revenue does not converge, repeat the process of obtaining the first cache set based on the Actor network set of the training set, the electricity market model, the joint operation system model, and the training model; When the convergence value of the first average total return is higher than the first preset value, the test model is obtained based on the training model corresponding to the highest first average total return. An initial risk coefficient is introduced into the Actor network set of the test model and the Actor network set of the test model is updated. A second cache set is obtained based on the test set, the electricity market model, the joint operation system model and the Actor network set of the test model. The second cache set represents the set of experience tuples of each decision agent in each unit time period based on the test set for the historical target time period. Update the initial risk coefficients and the Actor network set of the test model based on the second cache set; The second average total reward for each decision agent is obtained based on the returned reward. The second average total reward represents the average total reward of the decision agent in each historical target period based on the test set. When the second average total revenue does not converge, repeat the process of obtaining the second cache set based on the Actor network set of the test set, the electricity market model, the joint operation system model, and the test model. When the second average total return converges to a value higher than the second preset value, a decision model is obtained based on the Actor network set, and a risk preference coefficient is obtained based on the updated initial risk coefficient. Optionally, the second preset value is less than the first preset value to prevent overfitting.

[0049] The system continuously records all data from its online operation, including input status, output actions, actual market clearing results, and actual returns, forming new empirical tuple data. Using this new empirical tuple data, the decision-making model undergoes small-scale online incremental learning, enabling it to quickly adapt to subtle market changes.

[0050] Every so often, such as every quarter or half a year, or when there are significant changes in market rules, all new and old data are collected, and a new generation of decision-making models are repeatedly built based on historical datasets, electricity market models, and joint operation system models using multi-agent reinforcement learning and deep neural networks to ensure the continued advancement and adaptability of decision-making strategies.

[0051] Calculation example A joint operation system based on IEEE-30 nodes was used as an example for verification. This joint operation system consists of 41 lines, 20 load nodes, 6 conventional units, 2 pumped storage power stations, and 2 wind turbines.

[0052] Conventional generating units G1 through G6 are located at nodes 1, 2, 5, 8, 11, and 13, respectively. Pumped storage power stations E1 and E2 are located at nodes 18 and 24, respectively. Pumped storage power station E1 has a capacity of 1200MW / 3600MWh, and pumped storage power station E2 has a capacity of 1000MW / 4000MWh. The charging and discharging efficiency of both pumped storage power stations is 90%, the maximum allowable power generation is 90% of the capacity, and the minimum allowable power generation is 10% of the capacity. Assume that the total frequency regulation demand of the joint operation system equals 5% of the total load. Wind turbine units W1 and W2 with installed capacities of 3200MW and 2000MW, respectively, are added at nodes 5 and 8. The power generation price of wind turbine unit W1 is 180 yuan / MWh, and the power generation price of wind turbine unit W2 is 195 yuan / MWh.

[0053] The daily load curve of the joint operation system is as follows Figure 2 As shown in Tables 1 and 2, the pricing information for various conventional wind turbine units and pumped storage power stations is presented. Assuming that the predicted wind power output follows a normal distribution with a standard deviation of 0.1, a large number of scenarios are generated through Monte Carlo simulation sampling to represent the uncertainty of the predicted wind power output. Then, the K-means clustering algorithm is used to obtain ten typical scenarios. The predicted wind turbine output under these ten scenarios is shown in Tables 1 and 2. Figure 3 , Figure 4 As shown in Table 3, the probability values ​​for the ten scenarios are as follows.

[0054] Table 1. Application Information for Conventional Generating Units

[0055] Table 2 Application Information for Pumped Storage Power Stations

[0056] Table 3 Probabilities of 10 Typical Scenarios

[0057] When the risk appetite coefficient is 0.9, the bidding results for various conventional generating units, wind turbines, and pumped storage power stations in the electricity market are as follows: Figure 5 As shown. By Figure 5 It can be seen that conventional and wind turbine units account for 97.8% of the electricity supply in the power market. For conventional units, conventional unit G2 achieves a 94.63% capacity utilization rate due to its marginal cost advantage (its bid is 18.6% lower than the market average). During periods of high load, such as 17:00-19:00, conventional unit G2 reaches its maximum output, and conventional unit G4, with a lower bid, supplements the supply. Conventional unit G4, acting as a peak-shaving unit, contributes 8.02% of the incremental power supply during peak load periods. For wind turbine units, the winning bid volume of wind turbine unit W1 fluctuates significantly within 24 hours, ranging from 5.24MW to a maximum of 1928.04MW. Wind turbine unit W2, due to its higher bid, has a lower overall output and is discontinuous. During periods of lower wind turbine output, such as 15:00-19:00, the output of conventional turbines G2 and G4 increases; during periods of higher wind turbine output, such as 2:00-12:00, the output of conventional turbines decreases relatively.

[0058] Pumped storage power stations have a participation rate of only 2.92% in the electricity market, but exhibit a significant bi-peak characteristic: during periods of low load, such as 2:00-10:00, pumped storage power stations flexibly adjust their pumping operations based on electricity market signals to replenish the energy consumed during charging, enabling them to fully participate in the frequency regulation ancillary services market and obtain frequency regulation revenue. During periods of high load, pumped storage power stations reach peak electricity prices and generate electricity to obtain profits from the electricity market.

[0059] The frequency modulation capacity and mileage won in the frequency modulation ancillary services market are as follows: Figure 6 , Figure 7 As shown. From Figure 6 , Figure 7The bidding results for frequency regulation ancillary services across different time periods demonstrate that pumped storage power stations exhibit a significant competitive advantage in this market. Two pumped storage power stations account for 98.86% of the frequency regulation capacity demand and 99.92% of the frequency regulation mileage demand of the joint operation system. This phenomenon stems from a triple-driven factor: market mechanism design, resource and technological characteristics, and pricing mechanisms. First, the current frequency regulation ancillary service market generally adopts a "performance-based payment" principle, with the clearing price for frequency regulation mileage essentially reflecting the marginal demand of the joint operation system for high-quality regulation resources. Second, pumped storage power stations, with their instantaneous ramping capability and zero ramping cost, have a significantly higher frequency regulation mileage factor than conventional units, meaning they can contribute more effective mileage while providing the same frequency regulation capacity. Finally, under the market marginal clearing mechanism, the lower-level optimization objective (minimizing procurement costs) will prioritize resources that can reduce the overall cost per unit of frequency regulation mileage. Although pumped storage power stations may not be the cheapest, their superior regulation performance makes them the most economical option for meeting system frequency regulation requirements. Conventional units, limited by technical constraints such as ramp rate and response delay, find it difficult to compete for dominance in a performance-oriented market.

[0060] Compared to pumped storage power station E2, pumped storage power station E1 has a higher charging and discharging power, which means that E1 has a greater instantaneous regulation capability. In the frequency regulation ancillary services market, pumped storage power station E1 will be given priority in frequency regulation. Moreover, pumped storage power station E2 has a lower price and is closer to wind turbine nodes, so frequency regulation tasks will be allocated preferentially. At the same time, due to the participation of wind turbines, the frequency regulation capacity and frequency regulation mileage will also increase by 19%-28% compared to when pumped storage power stations participate in the power energy-frequency regulation ancillary services market alone.

[0061] The above results provide clear insights for consortium operation and system planning. For operational practice, pumped storage power stations should establish a strategic positioning that emphasizes both energy arbitrage and frequency regulation services: charging with low-cost electricity during off-peak hours and reserving sufficient capacity to participate in the frequency regulation ancillary services market; achieving energy arbitrage through discharging during peak hours, while simultaneously leveraging rapid response characteristics to obtain high-value frequency regulation revenue, thereby maximizing the value of flexible regulation.

[0062] Figure 8 To consider the bidding situation of each power generation unit in the electricity market when taking into account the risk preference coefficient, and Figure 9 This is a chart showing the percentage of successful bids for each power generation unit in the electricity market without considering risk appetite coefficients. Figure 8 , Figure 9The comparison shows that, when considering the risks arising from the uncertainty of wind turbine power generation forecasts, compared to not considering the risk preference coefficient, wind turbines (W1) are the main contributors, and their share of the awarded electricity volume decreases, while the share of awarded electricity volume for conventional turbines increases. This is because the uncertainty of wind power generation forecasts leads to forecasting deviations. To avoid high deviation costs, wind turbine manufacturers will proactively reduce their bids. Simultaneously, dispatching agencies or market participants will also reduce the proportion of wind power bids to ensure system reliability. For pumped storage power stations, the uncertainty of wind turbine power generation forecasts increases the frequency regulation pressure on the system, causing pumped storage power stations to suppress fluctuations through charging and discharging levels, i.e., participating more in the frequency regulation ancillary services market. Therefore, their awarded electricity volume in the electricity market will also decrease slightly.

[0063] Table 4 shows the proportion of frequency regulation capacity and mileage won in the frequency regulation ancillary services market under two scenarios: considering the risk preference coefficient and not considering it. Analysis of the data in the table shows that after considering the risk preference coefficient, the proportion of frequency regulation capacity and mileage won by pumped storage power stations in the frequency regulation ancillary services market increases significantly. Pumped storage power station E1, as the main provider in the frequency regulation ancillary services market, sees a significant increase in its proportion of frequency regulation capacity and mileage. Pumped storage power station E2, due to its high bid and low regulation capacity, shows a slight downward trend. The proportion of conventional units G4 decreases significantly. This is because the uncertainty of wind power leads to an increase in the total frequency regulation demand of the joint operation system. Conventional units have slower ramp rates and response speeds, making it difficult to cope with large power fluctuations at high frequencies. Therefore, the frequency regulation proportion of conventional units decreases, while pumped storage power stations, with their flexibility, rapid response capabilities, and zero ramp constraints, increase their proportion in the frequency regulation ancillary services market.

[0064] Table 4. Percentage of frequency regulation capacity and mileage won by different power generation units in the frequency regulation ancillary services market under two scenarios.

[0065] Taking pumped storage power station E1 as an example, based on 10 typical scenarios derived from clustering, the daily revenue of the pumped storage power station under two different decision-making strategies was calculated. The revenue situation of each scenario is as follows: Figure 10 As shown, Figure 10 In this context, risk representation is considered when considering the risk preference coefficient, while risk representation is not considered when not considering the risk preference coefficient.

[0066] from Figure 10As can be seen, the revenue performance of pumped storage power station E1 exhibits significant risk sensitivity. Without considering the risk preference coefficient, the daily revenue range of the power station is [659,400, 725,000] yuan, with a revenue fluctuation of 65,600 yuan. However, after optimizing the scheduling based on the decision-making strategy obtained through the decision-making model using the risk preference coefficient, the revenue range of pumped storage power station E1 is optimized to [678,300, 720,100] yuan, with the fluctuation reduced to 41,800 yuan, a decrease of 36.3%. Particularly noteworthy is that under 10 typical operating scenarios, the risk avoidance strategy cumulatively brings an additional 31,200 yuan in revenue to pumped storage power station E1. This empirical result verifies the dual advantages of the optimized scheduling method for the wind power and pumped storage combined operation system proposed in this application: while ensuring the basic revenue level, it can significantly reduce revenue volatility (by more than 36%), thus providing more robust decision support for pumped storage power stations to participate in electricity market bidding.

[0067] Finally, it should be noted that the above content is only used to illustrate the technical solution of the present invention, and is not intended to limit the scope of protection of the present invention. Simple modifications or equivalent substitutions made by those skilled in the art to the technical solution of the present invention do not depart from the essence and scope of the technical solution of the present invention.

Claims

1. An optimized scheduling method for a combined wind power and pumped storage system, characterized in that, include: Obtain real-time datasets and risk preference coefficients; The real-time dataset represents the set of input data for a real-time target period, including market information data, equipment status data, and forecast data; the risk preference coefficient represents the weighting factor between the returns and risks of the joint operation system. The decision value set is obtained by solving the decision model based on the real-time dataset and the risk preference coefficient; The decision value set includes the winning bid amount in the electricity market, the winning bid amount in the frequency regulation capacity, and the winning bid amount in the frequency regulation mileage for each power generation unit, wherein the power generation unit represents a wind turbine or a pumped storage power station. The joint operation system is scheduled according to the set of decision values; The decision model includes an objective function that aims to minimize the total electricity purchase cost in the market, which includes the electricity market and the frequency regulation ancillary service market. The decision-making model is constructed using multi-agent reinforcement learning and deep neural networks based on historical datasets, electricity market models, and joint operation system models. The historical dataset represents the set of input data for the historical target period. The joint operation system model includes a power grid model, at least two decision agents, and at least one conventional unit model; the at least two decision agents include at least one wind turbine model and at least one pumped storage power station model. The wind turbine model is constructed based on the uncertainty of wind turbine output prediction; The pumped storage power station model is constructed based on the reservoir capacity status of the pumped storage power station; The decision-making model, constructed using multi-agent reinforcement learning and deep neural networks based on historical datasets, electricity market models, and joint operation system models, includes: Obtain initialization parameters and initial risk coefficients; divide the historical dataset to obtain training and test sets; Training models are constructed based on deep neural networks, the initialization parameters, and the at least two decision agents. Each training model includes an Actor network set; the Actor network set includes at least two Actor networks; and each of the at least two Actor networks corresponds one-to-one with the at least two decision agents. A first cache set is obtained based on the training set, the electricity market model, the joint operation system model, and the Actor network set of the training model. This first cache set represents a set of experience tuples for each decision agent in each unit time period of the historical target time period, obtained based on the training set. Each experience tuple includes the current state, output action, return reward, and next state corresponding to the decision agent. The current state represents the state information of the corresponding decision agent in the current unit time period, obtained based on the training set, the electricity market model, and the joint operation system model. The output action represents the action information obtained by the corresponding Actor network based on the current state. The return reward represents the reward information of the corresponding decision agent obtained based on the simulated clearing result, and the simulated clearing result represents the clearing information obtained by the electricity market model based on the output action. The next state represents the state information of the joint operation system model in the next unit time period, obtained based on the output action. Update the training model according to the first cache set; The first average total reward for each decision agent is obtained based on the returned reward, and the first average total reward represents the average total reward of the decision agent in each historical target period based on the training set. When the first average total revenue does not converge, the process of obtaining the first cache set based on the training set, the electricity market model, the joint operation system model, and the Actor network set of the training model is repeated. When the convergence value of the first average total return is higher than the first preset value, a test model is obtained based on the training model corresponding to the highest first average total return. The initial risk coefficient is introduced into the Actor network set of the test model and the Actor network set of the test model is updated. A second cache set is obtained based on the test set, the electricity market model, the joint operation system model and the Actor network set of the test model. The second cache set represents the set of experience tuples of each decision agent in each unit time period of the historical target time period obtained based on the test set. The initial risk coefficient is updated based on the second cache set, and the Actor network set of the test model is updated accordingly. The second average total reward for each decision agent is obtained based on the returned reward, and the second average total reward represents the average of the total reward of the decision agent in each historical target period based on the test set; When the second average total revenue does not converge, the process of obtaining the second cache set based on the test set, the electricity market model, the joint operation system model, and the Actor network set of the test model is repeated. When the second average total return converges to a value higher than the second preset value, the decision model is obtained based on the Actor network set, and the risk preference coefficient is obtained based on the updated initial risk coefficient.

2. The optimized scheduling method for a wind power and pumped storage combined operation system according to claim 1, characterized in that, The total electricity purchase cost represents the linear superposition of the first total electricity purchase cost and the second total electricity purchase cost. The first total electricity purchase cost represents the total electricity purchase cost of the electricity market, and the second total electricity purchase cost represents the total electricity purchase cost of the frequency regulation ancillary service market.

3. The optimized scheduling method for a combined wind power and pumped storage system according to claim 1, characterized in that, The decision-making model also includes constraints, including power balance and line transmission capacity constraints, unit operation constraints, and frequency regulation market constraints. The power balance and line transmission capacity constraints characterize the coupling constraints between the node power load and the inter-node line transmission capacity based on the power grid model. The unit operation constraints include the internal control constraints of the joint operation system and the state constraints of the pumped storage power station. The frequency regulation market constraint characterizes the coupling constraint between demand and clearing based on the frequency regulation ancillary services market, the at least one wind turbine model, and the at least one pumped storage power station model.

4. The optimized scheduling method for a combined wind power and pumped storage system according to claim 3, characterized in that, The internal regulation constraints of the joint operation system are characterized by the regulation coupling constraints based on the electric energy market, the at least one wind turbine model, and the at least one pumped storage power station model.

5. The optimized scheduling method for a combined wind power and pumped storage system according to claim 3, characterized in that, The state constraints of the pumped storage power station include the operating condition constraints of the pumped storage power station, the start-stop state switching constraints of the pumped storage power station, and the minimum start-stop time constraints of the pumped storage power station. The operating condition constraints of the pumped storage power station are characterized by the coupling constraints between the charging and discharging power and the charging and discharging state of the electric energy market and the at least one pumped storage power station model. The start-stop state switching constraint of the pumped storage power station represents the start-stop state constraint of the at least one pumped storage power station model under power generation and pumping conditions, respectively. The minimum start-up and shutdown time constraint of the pumped storage power station represents the coupling constraint between the minimum continuous power generation time, minimum continuous shutdown time, power generation operation state, and pumping operation state of the at least one pumped storage power station model.

6. The optimized scheduling method for a wind power and pumped storage combined operation system according to claim 1, characterized in that, The risk preference coefficient is obtained when constructing the decision model using multi-agent reinforcement learning and deep neural networks based on historical datasets, electricity market models, and joint operation system models.

7. The optimized scheduling method for a wind power and pumped storage combined operation system according to any one of claims 1-6, characterized in that, The equipment status data includes the actual load curve, the actual output of the wind turbine, and the wind speed.

8. The optimized scheduling method for a wind power and pumped storage combined operation system according to any one of claims 1-6, characterized in that, The forecast data includes the forecast load curve and the forecast output of wind turbines.