A method for collaborative scheduling of wind-solar-hydrogen cluster based on multi-agent reinforcement learning

By constructing a scheduling agent for the wind-solar-hydrogen production cluster power supply and distribution system using a multi-agent reinforcement learning method, the system solves the global collaborative optimization scheduling problem of the wind-solar-hydrogen production cluster power supply and distribution system under uncertain environments, and realizes the collaborative power supply and distribution scheduling of the system and the satisfaction of grid operation constraints.

CN122437169APending Publication Date: 2026-07-21STATE GRID JIANGSU ELECTRIC POWER CO LTD RESEARCH INSTITUTE +1
View PDF -1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID JIANGSU ELECTRIC POWER CO LTD RESEARCH INSTITUTE
Filing Date
2026-06-24
Publication Date
2026-07-21

Smart Images

  • Figure CN122437169A_ABST
    Figure CN122437169A_ABST
Patent Text Reader

Abstract

The present application relates to power supply and distribution dispatching control technical field, especially to a kind of wind and light hydrogen production cluster collaborative scheduling method based on multi-agent reinforcement learning, method includes: obtaining the operation state data of wind and light hydrogen production cluster power supply and distribution system;According to operation state data, wind power unit, photovoltaic unit and hydrogen production unit are respectively constructed as dispatching agent;Based on the local operating state of each dispatching agent and the power grid operating constraint of wind and light hydrogen production cluster power supply and distribution system, a multi-agent collaborative scheduling model is constructed;According to the collaborative scheduling strategy corresponding to each dispatching agent generated by the multi-agent collaborative scheduling model;According to collaborative scheduling strategy, the dispatching instruction of wind power unit, photovoltaic unit and hydrogen production unit is generated respectively, so that wind and light hydrogen production cluster power supply and distribution system is under the condition of meeting the power grid operating constraint to carry out collaborative power supply and distribution scheduling.By the present application, the global collaborative optimization scheduling problem of wind and light hydrogen production cluster power supply and distribution system under uncertain operating environment is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power supply and distribution dispatch control technology, and in particular to a collaborative dispatch method for wind and solar hydrogen production clusters based on multi-agent reinforcement learning. Background Technology

[0002] With the continuous expansion of installed capacity of new energy sources such as wind power and photovoltaic power, using renewable energy electricity for water electrolysis to produce hydrogen has become an important technological path to enhance the absorption capacity of new energy, promote the synergistic utilization of electricity and hydrogen, and build a green energy system. The power supply and distribution system of wind-solar-hydrogen production clusters typically consists of wind power units, photovoltaic units, hydrogen production units, grid-connected nodes, and related power supply and distribution equipment. Among them, wind power units and photovoltaic units have obvious randomness, fluctuation, and intermittency, while hydrogen production units participate in system power balance and new energy absorption as adjustable electrical loads. Multiple units form a coupled operation relationship through the power supply and distribution network.

[0003] Existing wind-solar-hydrogen production scheduling technologies typically employ centralized optimization, predictive model-based scheduling, or the aggregation of multiple wind-solar-hydrogen production units into a unified control system. These scheduling processes rely on relatively accurate renewable energy output forecasts, equipment model parameters, and global operational information, using pre-established optimization models to solve for the operational plans of wind power, solar power, and hydrogen production loads. However, in actual operation, renewable energy output and market signals are constantly changing, and wind power, solar power, and hydrogen production units have different operating characteristics and regulatory constraints. Centralized scheduling methods struggle to adapt to complex operating environments in a timely manner, while aggregated control methods fail to fully reflect the local state differences and collaborative relationships between units. This makes it difficult for wind-solar-hydrogen production cluster power supply and distribution systems to achieve globally coordinated and optimized scheduling under uncertain operating environments.

[0004] The information disclosed in this background section is intended only to enhance the understanding of the general background of this disclosure and should not be construed as an admission or in any way implying that the information constitutes prior art known to those skilled in the art. Summary of the Invention

[0005] This invention provides a collaborative scheduling method for wind and solar hydrogen production clusters based on multi-agent reinforcement learning, which can effectively solve the problems in the background technology.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A collaborative scheduling method for wind-solar-hydrogen production clusters based on multi-agent reinforcement learning is applied to a power supply and distribution system for wind-solar-hydrogen production clusters, including wind power units, photovoltaic units, and hydrogen production units. The method includes: Obtain the operating status data of the power supply and distribution system of the wind-solar hydrogen production cluster; Based on the operational status data, the wind power unit, the photovoltaic unit, and the hydrogen production unit are respectively constructed as scheduling intelligent agents; Based on the local operating states of each scheduling agent and the power grid operation constraints of the wind-solar-hydrogen production cluster power supply and distribution system, a multi-agent collaborative scheduling model is constructed. Based on the multi-agent cooperative scheduling model, a cooperative scheduling strategy is generated for each of the scheduling agents. According to the collaborative scheduling strategy, scheduling instructions are generated for the wind power unit, the photovoltaic unit, and the hydrogen production unit, respectively, so that the wind-solar-hydrogen production cluster power supply and distribution system can carry out collaborative power supply and distribution scheduling under the condition of satisfying the grid operation constraints.

[0007] Furthermore, acquiring the operational status data of the power supply and distribution system of the wind-solar hydrogen production cluster includes: Collect new energy output data, grid connection node voltage data, injected active power data, and injected reactive power data of the wind power unit and the photovoltaic unit; Collect hydrogen production power data, hydrogen production quantity data, grid connection node voltage data, real-time electricity price data, and real-time hydrogen price data of the hydrogen production unit; Based on the new energy output data, the grid connection node voltage data, the injected active power data, the injected reactive power data, the hydrogen production power data, the hydrogen production volume data, the real-time electricity price data, and the real-time hydrogen price data, the operating status data of the wind-solar-hydrogen production cluster power supply and distribution system is formed.

[0008] Furthermore, based on the operational status data, the wind power unit, the photovoltaic unit, and the hydrogen production unit are respectively constructed as scheduling agents, including: Based on the operational status data, each wind power unit is constructed as a wind power dispatching intelligent agent for outputting wind power dispatching actions; Based on the operational status data, each photovoltaic unit is constructed as a photovoltaic scheduling agent for outputting photovoltaic scheduling actions; Based on the operational status data, each hydrogen production unit is constructed as a hydrogen production scheduling agent for outputting hydrogen production scheduling actions; The scheduling intelligence is composed of the wind power scheduling intelligence, the photovoltaic scheduling intelligence, and the hydrogen production scheduling intelligence, and a global operating state is formed based on the local operating state of each scheduling intelligence.

[0009] Furthermore, a multi-agent cooperative scheduling model is constructed, including: Based on the local operating state of each scheduling agent, the scheduling actions of each scheduling agent, and the global operating state, a Markov game model is established to describe the collaborative scheduling process of the wind-solar-hydrogen production cluster power supply and distribution system. The Markov game model is described by state space, action space, state transition function, and reward function. Based on the Markov game model, the local operating states of the wind power dispatching agent and the photovoltaic dispatching agent are determined, including local renewable energy output, grid-connected node voltage amplitude, and injected power. Based on the Markov game model, the local operating state of the hydrogen production scheduling agent is determined to include local hydrogen production power, hydrogen production volume, grid-connected node voltage amplitude, real-time electricity price, and real-time hydrogen price. Based on the Markov game model, the set of local operating states of all the scheduling agents is determined as the global operating state of the wind-solar-hydrogen production cluster power supply and distribution system.

[0010] Furthermore, constructing a multi-agent cooperative scheduling model also includes: According to the Markov game model, the scheduling actions of the wind power dispatching agent and the photovoltaic dispatching agent are determined as active power output adjustment and reactive power output adjustment. Based on the active power output adjustment, the reactive power output adjustment, and the corresponding predicted base output, the active power output command and the reactive power output command of the wind power unit and the photovoltaic unit are determined. Based on the Markov game model, the scheduling action of the hydrogen production scheduling agent is determined as the hydrogen production power setpoint. Based on the active power output command, the reactive power output command, and the hydrogen production power setting value, the power grid operation constraints are established, including upper and lower limit constraints on output, hydrogen production power limit constraints, hydrogen production power ramp rate constraints, power balance constraints, voltage safety constraints, and line power flow constraints.

[0011] Further, generating the cooperative scheduling strategy corresponding to each of the aforementioned scheduling agents includes: Based on the multi-agent cooperative scheduling model, a multi-objective cooperative evaluation function is established to evaluate the scheduling actions of each of the scheduling agents. Based on the on-grid electricity revenue of the wind power unit and the photovoltaic unit, determine the economic benefit evaluation value of the wind power dispatching intelligent agent and the photovoltaic dispatching intelligent agent; The economic benefit evaluation value of the hydrogen production scheduling agent is determined based on the hydrogen production revenue and electricity purchase cost of the hydrogen production unit. Based on the deviation between the voltage amplitude of the grid-connected node and the reference voltage, a grid security evaluation value is determined to constrain the voltage safety of the power supply and distribution system of the wind-solar-hydrogen production cluster. Based on the exceedance of the line power flow and line power flow limit of the relevant lines, a collaborative penalty evaluation value is determined to constrain the overall operating boundary of the power supply and distribution system of the wind-solar-hydrogen production cluster; wherein the lines associated with each agent include at least the power supply and distribution lines between the node where the corresponding agent is located and the grid-connected node; Based on the weighted result of the economic benefit evaluation value, the power grid security evaluation value, and the collaborative penalty evaluation value, the instantaneous evaluation value corresponding to the multi-objective collaborative evaluation function is generated.

[0012] Furthermore, generating the collaborative scheduling strategy corresponding to each of the aforementioned scheduling agents also includes: Based on the multi-agent cooperative scheduling model, a multi-agent deep deterministic policy gradient model including an actor network and a central critic network is constructed. During the training phase, the global running state and the scheduling actions of each scheduling agent are input into the central critic network to obtain the action value for evaluating the joint scheduling actions; During the training phase, the central critic network is updated based on the instantaneous evaluation value, the global running state at the next scheduling time, and the scheduling action output by the target policy network. During the training phase, the actor network corresponding to each scheduling agent is updated based on the action value output by the central critic network. During the execution phase, the local operating state of each scheduling agent is input into the corresponding actor network to generate a collaborative scheduling strategy for each scheduling agent.

[0013] Furthermore, coordinated power supply and distribution dispatching is carried out according to dispatching instructions, including: The multi-agent deep deterministic policy gradient model was pre-trained offline using historical wind and solar power output data, historical electricity price data, and historical hydrogen price data. The actor network obtained through offline pre-training is deployed to the corresponding wind power unit, photovoltaic unit, and hydrogen production unit; At the current scheduling moment, each of the scheduling agents collects its own local operating state and uses the corresponding actor network to generate the original scheduling action; The original scheduling action is constrained to obtain a scheduling instruction that satisfies the upper and lower limits of output, the limit of hydrogen production power, and the ramp rate of hydrogen production power. The active and reactive power outputs of the wind power unit, the active and reactive power outputs of the photovoltaic unit, and the hydrogen production power of the hydrogen production unit are controlled according to the dispatch instructions. The system collects real-time operating data of the wind-solar hydrogen production cluster power supply and distribution system after executing the scheduling command, and uses the real-time operating data to fine-tune the multi-agent deep deterministic strategy gradient model online to adapt to changes in long-term operating conditions.

[0014] The technical solution of this invention can achieve the following technical effects: By collecting the operating status of the wind-solar-hydrogen production cluster, the wind power, photovoltaic and hydrogen production units are constructed as scheduling agents. Based on the multi-agent collaborative scheduling model, scheduling strategies are generated to achieve power supply and distribution collaborative optimization scheduling that meets the grid operation constraints. This effectively solves the problem of global collaborative optimization scheduling of the power supply and distribution system of the wind-solar-hydrogen production cluster under uncertain operating conditions.

[0015] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart illustrating a collaborative scheduling method for wind-solar hydrogen production clusters based on multi-agent reinforcement learning. Figure 2 This is a schematic diagram of the structure of a wind-solar hydrogen production cluster system. Detailed Implementation

[0018] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0020] Example 1; like Figure 1As shown, this application provides a collaborative scheduling method for wind-solar hydrogen production clusters based on multi-agent reinforcement learning. The method includes: S10: Obtain the operating status data of the power supply and distribution system of the wind-solar hydrogen production cluster; S20: Based on the operational status data, the wind power unit, photovoltaic unit, and hydrogen production unit are respectively constructed as scheduling intelligent agents; S30: Based on the local operating states of each scheduling agent and the grid operation constraints of the wind-solar-hydrogen production cluster power supply and distribution system, a multi-agent collaborative scheduling model is constructed. S40: Generate the cooperative scheduling strategy for each scheduling agent based on the multi-agent cooperative scheduling model; S50: Generate scheduling instructions for wind power units, photovoltaic units and hydrogen production units respectively according to the coordinated scheduling strategy, so that the wind-solar-hydrogen production cluster power supply and distribution system can carry out coordinated power supply and distribution scheduling under the condition of meeting the grid operation constraints.

[0021] Specifically, in one embodiment, the wind-solar-hydrogen production cluster power supply and distribution system includes a wind power unit, a photovoltaic unit, a hydrogen production unit, a grid-connected node, power supply and distribution lines, and a dispatch control device. The wind power unit, photovoltaic unit, and hydrogen production unit are electrically connected to the grid-connected node through the power supply and distribution lines. The dispatch control device is used to obtain the operating status of the wind-solar-hydrogen production cluster power supply and distribution system and generate corresponding dispatch instructions. First, the dispatch control device acquires the operating status data of the wind-solar-hydrogen production cluster power supply and distribution system. The operating status data includes the new energy output status, grid connection node voltage status, and injection power status of the wind power unit and photovoltaic unit, as well as the hydrogen production power status, hydrogen production operation status, and operation-related information such as electricity price and hydrogen price of the hydrogen production unit, thereby providing status input for subsequent coordinated dispatch. Then, based on the operating status data, the dispatch control device constructs the wind power unit, photovoltaic unit and hydrogen production unit as dispatch intelligent agents respectively, so that the dispatch intelligent agent corresponding to the wind power unit is used to characterize the wind power output regulation capability, the dispatch intelligent agent corresponding to the photovoltaic unit is used to characterize the photovoltaic output regulation capability, and the dispatch intelligent agent corresponding to the hydrogen production unit is used to characterize the hydrogen production load regulation capability. Next, the dispatch control device constructs a multi-agent collaborative dispatch model based on the local operating status of each dispatch agent and the power grid operation constraints of the wind-solar-hydrogen production cluster power supply and distribution system. The local operating status is used to characterize the operating status of the corresponding unit at the current dispatch time. The power grid operation constraints include power balance constraints, voltage safety constraints, line power flow constraints, and equipment operation constraints, so that the multi-agent collaborative dispatch model can reflect the power supply and distribution coupling relationship between the wind power unit, photovoltaic unit, and hydrogen production unit. Subsequently, the dispatch control device generates a collaborative dispatch strategy for each dispatch agent based on the multi-agent collaborative dispatch model, enabling each dispatch agent to meet the overall operational requirements of the wind-solar-hydrogen production cluster power supply and distribution system while considering its own local operating status. Finally, the dispatch control device generates dispatch instructions for wind power units, photovoltaic units, and hydrogen production units respectively according to the collaborative dispatch strategy, and sends the dispatch instructions to the corresponding wind power units, photovoltaic units, and hydrogen production units respectively, so that the wind power units and photovoltaic units can perform active or reactive power output adjustment, and the hydrogen production units can perform hydrogen production power adjustment. Thus, the wind-solar-hydrogen production cluster power supply and distribution system completes collaborative power supply and distribution dispatch under the condition of meeting the grid operation constraints.

[0022] The technical solution of this invention collects the operating status of wind-solar-hydrogen production clusters, constructs wind power, photovoltaic and hydrogen production units as scheduling agents, and generates scheduling strategies based on a multi-agent collaborative scheduling model to achieve power supply and distribution collaborative optimization scheduling that meets grid operation constraints. This effectively solves the problem of global collaborative optimization scheduling of power supply and distribution systems for wind-solar-hydrogen production clusters under uncertain operating environments.

[0023] Furthermore, acquiring operational status data of the power supply and distribution system for the wind-solar-hydrogen production cluster includes: Collect new energy output data, grid connection node voltage data, injected active power data, and injected reactive power data for wind power units and photovoltaic units; Collect hydrogen production power data, hydrogen production volume data, grid connection node voltage data, real-time electricity price data, and real-time hydrogen price data of the hydrogen production unit; Based on the power output data of new energy sources, grid connection node voltage data, injected active power data, injected reactive power data, hydrogen production power data, hydrogen production volume data, real-time electricity price data, and real-time hydrogen price data, the operating status data of the wind-solar-hydrogen production cluster power supply and distribution system is formed.

[0024] Furthermore, based on operational status data, the wind power unit, photovoltaic unit, and hydrogen production unit are each constructed as a scheduling intelligent agent, including: Based on the operational status data, each wind power unit is constructed as a wind power dispatching intelligent agent for outputting wind power dispatching actions; Based on the operational status data, each photovoltaic unit is constructed as a photovoltaic scheduling agent for outputting photovoltaic scheduling actions; Based on the operational status data, each hydrogen production unit is constructed as a hydrogen production scheduling agent for outputting hydrogen production scheduling actions; The scheduling intelligence is composed of wind power scheduling intelligence, photovoltaic scheduling intelligence, and hydrogen production scheduling intelligence, and the global operating status is formed based on the local operating status of each scheduling intelligence.

[0025] Furthermore, constructing a multi-agent cooperative scheduling model includes: Based on the local operating state, scheduling actions, and global operating state of each scheduling agent, a Markov game model is established to describe the collaborative scheduling process of the power supply and distribution system of the wind-solar-hydrogen production cluster. The Markov game model describes the system through state space, action space, state transition function, and reward function. Based on the Markov game model, the local operating states of the wind power dispatching agent and the photovoltaic dispatching agent are determined, including local renewable energy output, grid-connected node voltage amplitude, and injected power. Based on the Markov game model, the local operating state of the hydrogen production scheduling agent is determined, including local hydrogen production power, hydrogen production quantity, grid-connected node voltage amplitude, real-time electricity price, and real-time hydrogen price. Based on the Markov game model, the set of local operating states of all scheduling agents is determined as the global operating state of the wind-solar-hydrogen production cluster power supply and distribution system.

[0026] Furthermore, constructing a multi-agent cooperative scheduling model also includes: Based on the Markov game model, the scheduling actions of the wind power dispatching agent and the photovoltaic dispatching agent are defined as active power output adjustment and reactive power output adjustment. Based on the active power output adjustment, reactive power output adjustment and the corresponding predicted base output, determine the active power output command and reactive power output command for the wind power unit and the photovoltaic unit. Based on the Markov game model, the scheduling action of the hydrogen production scheduling agent is determined as the hydrogen production power setpoint. Based on the active power output command, reactive power output command, and hydrogen production power setpoint, power grid operation constraints are established, including upper and lower limits of output, hydrogen production power limit constraints, hydrogen production power ramp rate constraints, power balance constraints, voltage safety constraints, and line power flow constraints.

[0027] Furthermore, generating the collaborative scheduling strategy corresponding to each scheduling agent includes: Based on the multi-agent cooperative scheduling model, a multi-objective cooperative evaluation function is established to evaluate the scheduling actions of each scheduling agent. Based on the on-grid electricity revenue of wind power units and photovoltaic units, determine the economic benefit evaluation values ​​of wind power dispatching intelligent agents and photovoltaic dispatching intelligent agents; The economic benefit evaluation value of the hydrogen production scheduling agent is determined based on the hydrogen production revenue and electricity purchase cost of the hydrogen production unit. Based on the deviation between the voltage amplitude of the grid-connected node and the reference voltage, determine the grid security evaluation value used to constrain the voltage safety of the power supply and distribution system of the wind-solar-hydrogen production cluster; Based on the exceedance of the line power flow and line power flow limit of the relevant lines, determine the collaborative penalty evaluation value used to constrain the overall operation boundary of the power supply and distribution system of the wind-solar-hydrogen production cluster; wherein the lines associated with each agent include at least the power supply and distribution lines between the node where the corresponding agent is located and the grid-connected node; Based on the weighted results of the economic benefit evaluation value, power grid safety evaluation value, and collaborative penalty evaluation value, the instantaneous evaluation value corresponding to the multi-objective collaborative evaluation function is generated.

[0028] Furthermore, generating the collaborative scheduling strategy for each scheduling agent also includes: Based on the multi-agent cooperative scheduling model, a multi-agent deep deterministic policy gradient model including an actor network and a central critic network is constructed. During the training phase, the global running state and the scheduling actions of each scheduling agent are input into the central critic network to obtain the action value for evaluating the joint scheduling actions; During the training phase, the central critic network is updated based on the immediate evaluation value, the global running state at the next scheduling time, and the scheduling action output by the target policy network. During the training phase, the actor network corresponding to each scheduling agent is updated based on the action value output by the central critic network. During the execution phase, the local operating state of each scheduling agent is input into the corresponding actor network to generate the collaborative scheduling strategy for each scheduling agent.

[0029] Furthermore, coordinated power supply and distribution dispatching according to dispatching instructions includes: Offline pre-training of a multi-agent deep deterministic policy gradient model was performed using historical wind and solar power output data, historical electricity price data, and historical hydrogen price data. The actor network obtained through offline pre-training will be deployed to the corresponding wind power unit, photovoltaic unit, and hydrogen production unit; At the current scheduling moment, each scheduling agent collects its own local operating state and uses the corresponding actor network to generate the original scheduling action; The original scheduling actions are constrained to obtain scheduling instructions that satisfy the upper and lower limits of output, the limit of hydrogen production power, and the ramp rate of hydrogen production power. According to the dispatch instructions, the active and reactive power outputs of the wind power unit, the active and reactive power outputs of the photovoltaic unit, and the hydrogen production power of the hydrogen production unit are controlled respectively. Real-time operational data of the power supply and distribution system of the wind-solar hydrogen production cluster is collected after the execution of scheduling instructions, and the real-time operational data is used to fine-tune the multi-agent deep deterministic strategy gradient model online to adapt to changes in long-term operating conditions.

[0030] As a preferred embodiment of the above embodiments, in one specific embodiment: Set of intelligent agents: Definition Individual agents, including A smart entity for a wind farm A photovoltaic power station intelligent agent and A smart entity for a hydrogen production station, namely ; State space: for any intelligent agent Its time Local observation status include: Local renewable energy output: wind farms or photovoltaic power station Actual output; Local hydrogen production station status: Current total operating power of the hydrogen production station Hydrogen production capacity ; Common coupling point electrical quantities: node voltage magnitude Injected active power reactive power ; Market Signals: Real-time Electricity Prices Real-time hydrogen price ; Global state The set of all local observation states ; Action Space: Intelligent Agent action This determines its scheduling instructions; For intelligent agents in wind farms and intelligent agents in photovoltaic power plants, actions This is expressed as an adjustment in active power output. and reactive power output adjustment And there are: ; ; in, Represents intelligent agents During scheduling The injected active power; Represents intelligent agents During scheduling reactive power; Represents intelligent agents During scheduling The basis for prediction is meritorious and effective; Represents intelligent agents During scheduling The prediction basis is reactive power output; The action must meet the upper and lower limits of output constraints: ; ; in, Represents intelligent agents The lower limit of the allowed active power output; Represents intelligent agents The maximum allowed active power output; Represents intelligent agents The lower limit of the allowed reactive power output; Represents intelligent agents The maximum allowable reactive power output; For the intelligent agent of the hydrogen production station, actions This is represented as the hydrogen production power setpoint. Limited by the power limits and ramp-up constraints of hydrogen production stations: ; ; in, , These are the lower and upper limits of hydrogen production capacity, respectively. Maximum gradeability; intelligent agent Instant rewards gained through joint actions It consists of three parts: ; in, These are weighting coefficients used to balance the importance of each optimization objective; Rewards for economic benefits; Rewards for power grid safety; To coordinate punishment and reward; Economic benefit reward This reflects the operational benefits of the intelligent agent, for intelligent agents in wind farms and photovoltaic power plants: ; in, For real-time electricity prices; To inject active power; The scheduling time interval; For intelligent agents at hydrogen production stations: ; in, Real-time hydrogen price; For hydrogen production quality, compared with the hydrogen production power setpoint The relationship is expressed through the efficiency curve. Sure: ; in, This is the higher calorific value of hydrogen. Power Grid Safety Rewards To incentivize agents to maintain node voltage safety, a voltage deviation penalty function is used: ; in, This refers to the node voltage amplitude. Reference voltage; , The upper and lower limits for safe voltage operation; Collaborative punishment and reward Used to constrain the overall operational boundaries of the cluster and prevent global problems caused by local decisions, mainly including penalties for exceeding line power flow limits: ; in, To work with intelligent agents Related route collection; For the line exist The trends of the moment; For line power flow limits; This is the penalty coefficient; Furthermore, the system must also satisfy global power balance constraints: ; in, The number of intelligent agents in the wind farm; The number of intelligent agents in the photovoltaic power station; The number of intelligent agents in the hydrogen production station; For the actual power output of the wind farm; For the actual output of photovoltaic power plants; Set the hydrogen production power value; The power transmitted to the power grid (can be positive or negative); Furthermore, the core of the centralized training and distributed execution architecture lies in utilizing a centralized critic network to acquire the states of all agents during training. and actions To assess the value of the action This effectively addresses the issue of environmental non-stationarity; however, during the execution phase, the actor network of each agent relies only on local observations. Actions can be output. This enables fully distributed collaboration.

[0031] Specifically, for intelligent agents Its actor network Parameterization The input is the local observation state. The output is an action. Central Commentator Network Parameterization The input is the global state. and the actions of all intelligent agents The output is an estimate of the action value; During training, the loss function of the commentator network is: ; in, Indicates from the experience replay pool Sample training samples and calculate the expectation. For intelligent agents During scheduling Instant rewards gained; value of the target action For the target value, For the target policy network, Discount factor; The policy gradient of the actor network is: ; in, To represent intelligent agents Actor network objective function Regarding actor network parameters The gradient; Indicates from the experience replay pool Mid-sampling scheduling time global state And calculate the expectation of the sampling results; For intelligent agents Action variables; Furthermore, offline pre-training uses historical wind and solar power output data, electricity price data, and hydrogen price data as the training set; in the online fine-tuning stage, real-time data collected during actual operation is used to continuously optimize the network parameters using a small-batch gradient update method to adapt to changes in system operating characteristics. Furthermore, in actual operation, at each scheduling moment Each agent performs the following operations in parallel: (i) Acquiring local observation status ; (ii) will Input actor network To obtain the original action ; (iii) Post-process the original actions to ensure that equipment operating constraints (such as power limits and ramp rate constraints) are met. (iv) Perform the final action Update the device's operating status; (v) Waiting for the next scheduling time.

[0032] Example 2; In a preferred embodiment, a real-world integrated wind-solar-hydrogen production energy base is used as an example. This base includes two wind farms, one photovoltaic power station, and two hydrogen production stations. All units are connected to the main grid through a 220kV substation. The system structure is as follows: Figure 2 As shown; Step S1 involves modeling the multi-agent system, defining two wind farms (W1, W2), one photovoltaic power station (S1), and two hydrogen production stations (H1, H2) as independent agents. Each agent can collect electrical quantities such as voltage and power at its local node as local observations. The communication network between agents allows for the exchange of necessary global information during the training phase. Step S2: Define a Markov game mathematical model, taking wind farm agent W1 as an example, its local observation state Including: Current contributions and efforts Currently, no effort is being put into production. Grid connection point voltage Real-time electricity price The action space is the active power adjustment amount at the next moment. and reactive power adjustment It must meet the upper and lower limits of output constraints; Taking the intelligent agent H1 of the hydrogen production station as an example, its local observation status Including: current hydrogen production capacity Cumulative hydrogen production, grid connection voltage Real-time electricity price Real-time hydrogen price The action space is the hydrogen production power setpoint for the next moment. It must meet power limits and ramp rate constraints; Reward function design: Taking the intelligent agent H1 of the hydrogen production station as an example, assuming the current hydrogen price Yuan / Nm 3 Electricity price Yuan / kWh; if H1 operates at 10MW with an efficiency of 55 kWh / kg, then the mass of hydrogen produced is... kg, revenue from hydrogen sales is Yuan, electricity purchase cost is Yuan, then economic benefit reward (A negative value indicates that the cost exceeds the benefit, and the dispatch strategy will tend to produce more hydrogen when the electricity price is low and less hydrogen when the electricity price is high); If the grid connection point voltage is 1.05 pu, the reference voltage is 1.0 pu, and the voltage safety upper and lower limits are 0.95~1.05 pu, then the grid safety bonus... If the power transmitted by the main transformer at the aggregation station does not exceed the limit, then a coordinated penalty and reward will be applied. ; Step S3: Construct the MADDPG training environment. Based on a deep learning framework (such as PyTorch or TensorFlow), build a simulation environment that encapsulates a simplified model of power system flow calculation and hydrogen production equipment. Initialize the Actor network and Critic network for five agents, as well as the corresponding target network. Each Actor network is a 3-layer fully connected neural network (input layer, 128-dimensional hidden layer, output layer), using ReLU as the activation function. Each Critic network is a 4-layer fully connected neural network, with the input being the concatenation of the global state and all actions. Step S4: Use historical one year's worth of wind and solar power output data and time-of-use electricity price data as the training set to conduct offline pre-training; Training parameter settings: discount factor Learning rate Experience replay pool capacity Batch size Target network soft update coefficient ; The training process is as follows: (i) At the beginning of each episode, the environment state is initialized by randomly selecting one day's time-series data from the historical data; (ii) at each time step The Actor network of each agent is based on local observations. Adding exploration noise (Ornstein-Uhlenbeck noise) to obtain the action ; (iii) Performing joint actions Environmental renewal trends, calculating hydrogen content and instant rewards for various intelligent systems. To obtain the state at the next moment. ; (iv) Experience Store the experience in the experience replay pool. The experience replay pool is a fixed-capacity experience storage structure. When the number of stored experiences exceeds the capacity limit, old experiences are removed according to the first-in-first-out rule. (v) Randomly sample a small batch of data from the replay pool and update the central Critic network; (vi) Update the parameters of each Actor network using policy gradients; (vii) After training for 5000 episodes, the target network parameters are softly updated, the policy networks of each agent converge, and the cumulative reward tends to stabilize; Step S5: Online application – The trained Actor network parameters are deployed to the control systems of various wind and solar power plants and hydrogen production stations; in actual operation, each agent executes in parallel at each scheduling moment (e.g., every 5 minutes): Wind farm intelligent agent W1: Collects current active power output, reactive power output, and grid connection point voltage, inputs them into the Actor network, and obtains active power adjustment. and reactive power adjustment After being limited in amplitude, the signal is sent to the wind turbine controller. Hydrogen production station intelligent agent H1: Collects current hydrogen production power, grid connection voltage, real-time electricity price, and hydrogen price, inputs them into the Actor network, and obtains the hydrogen production power setpoint for the next time step. After the ramp rate constraint check, the data is sent to the electrolytic cell controller. Meanwhile, the system continuously collects actual operating data and regularly (e.g., daily) fine-tunes the model using new data to adapt to long-term changes such as seasonal variations and equipment aging.

[0033] Although this application has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made thereto without departing from the spirit and scope of this application. Accordingly, this specification and drawings are merely exemplary illustrations of the application as defined herein, and are to be considered as covering any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Thus, if such modifications and modifications fall within the scope of this application and its equivalents, this application intends to include such modifications and modifications.

Claims

1. A collaborative scheduling method for wind-solar hydrogen production clusters based on multi-agent reinforcement learning, characterized in that, The method, applied to a wind-solar-hydrogen production cluster power supply and distribution system including wind power units, photovoltaic units, and hydrogen production units, comprises: Obtain the operating status data of the power supply and distribution system of the wind-solar hydrogen production cluster; Based on the operational status data, the wind power unit, the photovoltaic unit, and the hydrogen production unit are respectively constructed as scheduling intelligent agents; Based on the local operating states of each scheduling agent and the power grid operation constraints of the wind-solar-hydrogen production cluster power supply and distribution system, a multi-agent collaborative scheduling model is constructed. Based on the multi-agent cooperative scheduling model, a cooperative scheduling strategy is generated for each of the scheduling agents. According to the collaborative scheduling strategy, scheduling instructions are generated for the wind power unit, the photovoltaic unit, and the hydrogen production unit, respectively, so that the wind-solar-hydrogen production cluster power supply and distribution system can carry out collaborative power supply and distribution scheduling under the condition of satisfying the grid operation constraints.

2. The wind-solar hydrogen production cluster collaborative scheduling method based on multi-agent reinforcement learning according to claim 1, characterized in that, Obtaining the operating status data of the power supply and distribution system of the wind-solar hydrogen production cluster includes: Collect new energy output data, grid connection node voltage data, injected active power data, and injected reactive power data of the wind power unit and the photovoltaic unit; Collect hydrogen production power data, hydrogen production quantity data, grid connection node voltage data, real-time electricity price data, and real-time hydrogen price data of the hydrogen production unit; Based on the new energy output data, the grid connection node voltage data, the injected active power data, the injected reactive power data, the hydrogen production power data, the hydrogen production volume data, the real-time electricity price data, and the real-time hydrogen price data, the operating status data of the wind-solar-hydrogen production cluster power supply and distribution system is formed.

3. The wind-solar hydrogen production cluster collaborative scheduling method based on multi-agent reinforcement learning according to claim 1, characterized in that, Based on the operational status data, the wind power unit, the photovoltaic unit, and the hydrogen production unit are respectively constructed as scheduling agents, including: Based on the operational status data, each wind power unit is constructed as a wind power dispatching intelligent agent for outputting wind power dispatching actions; Based on the operational status data, each photovoltaic unit is constructed as a photovoltaic scheduling agent for outputting photovoltaic scheduling actions; Based on the operational status data, each hydrogen production unit is constructed as a hydrogen production scheduling agent for outputting hydrogen production scheduling actions; The scheduling intelligence is composed of the wind power scheduling intelligence, the photovoltaic scheduling intelligence, and the hydrogen production scheduling intelligence, and a global operating state is formed based on the local operating state of each scheduling intelligence.

4. The wind-solar hydrogen production cluster collaborative scheduling method based on multi-agent reinforcement learning according to claim 1, characterized in that, Constructing a multi-agent cooperative scheduling model, including: Based on the local operating state of each scheduling agent, the scheduling actions of each scheduling agent, and the global operating state, a Markov game model is established to describe the collaborative scheduling process of the wind-solar-hydrogen production cluster power supply and distribution system. The Markov game model is described by state space, action space, state transition function, and reward function. Based on the Markov game model, the local operating states of the wind power dispatching agent and the photovoltaic dispatching agent are determined, including local renewable energy output, grid-connected node voltage amplitude, and injected power. Based on the Markov game model, the local operating state of the hydrogen production scheduling agent is determined to include local hydrogen production power, hydrogen production volume, grid-connected node voltage amplitude, real-time electricity price, and real-time hydrogen price. Based on the Markov game model, the set of local operating states of all the scheduling agents is determined as the global operating state of the wind-solar-hydrogen production cluster power supply and distribution system.

5. The wind-solar hydrogen production cluster collaborative scheduling method based on multi-agent reinforcement learning according to claim 4, characterized in that, The construction of a multi-agent cooperative scheduling model also includes: According to the Markov game model, the scheduling actions of the wind power dispatching agent and the photovoltaic dispatching agent are determined as active power output adjustment and reactive power output adjustment. Based on the active power output adjustment, the reactive power output adjustment, and the corresponding predicted base output, the active power output command and the reactive power output command of the wind power unit and the photovoltaic unit are determined. Based on the Markov game model, the scheduling action of the hydrogen production scheduling agent is determined as the hydrogen production power setpoint. Based on the active power output command, the reactive power output command, and the hydrogen production power setting value, the power grid operation constraints are established, including upper and lower limit constraints on output, hydrogen production power limit constraints, hydrogen production power ramp rate constraints, power balance constraints, voltage safety constraints, and line power flow constraints.

6. The wind-solar hydrogen production cluster collaborative scheduling method based on multi-agent reinforcement learning according to claim 1, characterized in that, Generate the collaborative scheduling strategy corresponding to each of the aforementioned scheduling agents, including: Based on the multi-agent cooperative scheduling model, a multi-objective cooperative evaluation function is established to evaluate the scheduling actions of each of the scheduling agents. Based on the on-grid electricity revenue of the wind power unit and the photovoltaic unit, determine the economic benefit evaluation value of the wind power dispatching intelligent agent and the photovoltaic dispatching intelligent agent; The economic benefit evaluation value of the hydrogen production scheduling agent is determined based on the hydrogen production revenue and electricity purchase cost of the hydrogen production unit. Based on the deviation between the voltage amplitude of the grid-connected node and the reference voltage, a grid security evaluation value is determined to constrain the voltage safety of the power supply and distribution system of the wind-solar-hydrogen production cluster. Based on the exceedance of the line power flow and line power flow limit of the relevant lines, a collaborative penalty evaluation value is determined to constrain the overall operating boundary of the power supply and distribution system of the wind-solar-hydrogen production cluster; wherein the lines associated with each agent include at least the power supply and distribution lines between the node where the corresponding agent is located and the grid-connected node; Based on the weighted result of the economic benefit evaluation value, the power grid security evaluation value, and the collaborative penalty evaluation value, the instantaneous evaluation value corresponding to the multi-objective collaborative evaluation function is generated.

7. The wind-solar hydrogen production cluster collaborative scheduling method based on multi-agent reinforcement learning according to claim 6, characterized in that, Generating the collaborative scheduling strategy corresponding to each of the aforementioned scheduling agents further includes: Based on the multi-agent cooperative scheduling model, a multi-agent deep deterministic policy gradient model including an actor network and a central critic network is constructed. During the training phase, the global running state and the scheduling actions of each scheduling agent are input into the central critic network to obtain the action value for evaluating the joint scheduling actions; During the training phase, the central critic network is updated based on the instantaneous evaluation value, the global running state at the next scheduling time, and the scheduling action output by the target policy network. During the training phase, the actor network corresponding to each scheduling agent is updated based on the action value output by the central critic network. During the execution phase, the local operating state of each scheduling agent is input into the corresponding actor network to generate a collaborative scheduling strategy for each scheduling agent.

8. The wind-solar hydrogen production cluster collaborative scheduling method based on multi-agent reinforcement learning according to claim 7, characterized in that, Coordinated power supply and distribution dispatching according to dispatching instructions includes: The multi-agent deep deterministic policy gradient model was pre-trained offline using historical wind and solar power output data, historical electricity price data, and historical hydrogen price data. The actor network obtained through offline pre-training is deployed to the corresponding wind power unit, photovoltaic unit, and hydrogen production unit; At the current scheduling moment, each of the scheduling agents collects its own local operating state and uses the corresponding actor network to generate the original scheduling action; The original scheduling action is constrained to obtain a scheduling instruction that satisfies the upper and lower limits of output, the limit of hydrogen production power, and the ramp rate of hydrogen production power. The active and reactive power outputs of the wind power unit, the active and reactive power outputs of the photovoltaic unit, and the hydrogen production power of the hydrogen production unit are controlled according to the dispatch instructions. The system collects real-time operating data of the wind-solar hydrogen production cluster power supply and distribution system after executing the scheduling command, and uses the real-time operating data to fine-tune the multi-agent deep deterministic strategy gradient model online to adapt to changes in long-term operating conditions.