Coordinated control method and device for output voltage of proton exchange membrane fuel cell power generation system
Through the multi-agent depth deterministic strategy gradient algorithm, the stability of the output voltage of the proton exchange membrane fuel cell is achieved, and the coordination and control problem between oxygen flow rate and DC/DC converter duty cycle is solved, which improves the tracking accuracy of the output voltage and the robustness of the system.
Patent Information
- Application Number
- CN202510294608.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art is difficult to achieve the stability of the output voltage of the proton exchange membrane fuel cell (PEMFC), especially under different load conditions, the coordinated control between the oxygen flow rate and the duty cycle of the DC/DC converter is not easy to achieve, resulting in fluctuations and instability of the output voltage.
Using the multi-agent depth deterministic strategy gradient algorithm, through leader selection training and leadership-following training, a first agent for controlling the blower and a second agent for controlling the DC/DC converter are established, and the action space, state space and reward functions are defined respectively to realize coordinated control between the blower air flow rate and the duty cycle of the DC/DC converter.
It improves the tracking accuracy and stability of the PEMFC output voltage, enhances the robustness of the system, improves the overall operating efficiency, and solves the problems of output voltage fluctuations and instability.
Smart Images

Figure CN120184288A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of output voltage control of proton exchange membrane fuel cells, and particularly relates to a method and device for coordinated control of the output voltage of a proton exchange membrane fuel cell power generation system. Background Art
[0002] In recent years, hydrogen energy has developed vigorously due to its advantages of cleanness and high mass / energy density. As one of the main scenarios for current hydrogen energy utilization, hydrogen fuel cells can provide a more stable power output compared with renewable energy sources such as wind and solar, and can relieve the grid pressure during peak power demand. Among hydrogen fuel cells, proton exchange membrane fuel cells (PEMFCs) have the advantages of rapid startup and strong ability to respond to dynamic loads, and are widely used in the fields of power generation and transportation. However, a PEMFC is a non-linear time-varying system with complex internal parameters and involving multiple energy flows. Under different load conditions, the battery voltage will change violently and fluctuate, thus reducing the stability of the output voltage. Therefore, in order to stably maintain the working performance of the proton exchange membrane fuel cell, the proton exchange membrane fuel cell needs to be equipped with corresponding auxiliary subsystems. Considering the different dynamic characteristics of each subsystem, it is necessary to coordinate multiple subsystems to stabilize the output voltage. According to relevant research, the air flow rate of the blower and the duty ratio of the DC / DC converter are two important factors affecting the output voltage of the PEMFC. However, current existing technologies usually model and analyze these two factors separately, and only improve the output voltage by improving the performance of the blower controller or the DC / DC converter controller.
[0003] For example, some scholars have proposed a sliding mode control strategy that can simultaneously adjust the flow rates of hydrogen and oxygen. However, since this strategy cannot adjust the working state of the DC / DC converter, its influence on the output voltage of the PEMFC is relatively small; some scholars have constructed an optimal voltage controller for PEMFCs based on distributed deep reinforcement learning and applied the integrated gradient method to adjust the output voltage. The experimental results show that this strategy significantly enhances the robustness and adaptability of the PEMFC system. However, since the model of this method is often based on an assumed environment during the training stage, it may show poor adaptability when facing a real and uncertain operating environment. At the same time, its convergence and stability are difficult to guarantee, resulting in its not being easy to apply; some researchers have proposed a method for adjusting the output characteristics of PEMFCs based on active disturbance rejection control to solve the problems of slow response and overshoot in the process of tracking the oxygen excess ratio in the PEMFC system by traditional PID control, thus achieving rapid and accurate adjustment of the oxygen excess ratio. However, the robustness of this algorithm is low and it is not easy to apply in practice.
[0004] Furthermore, the above research has paid little attention to the coupling effect between the two, and failed to achieve coordinated control between them. Generally speaking, the relationship between the oxygen flow rate and the output voltage of the fuel cell is highly non-linear, and there is a close coupling relationship between the oxygen flow rate and the duty ratio of the DC / DC converter at the output voltage level. Simple linear models and methods are difficult to achieve coordinated control of the oxygen flow rate and the duty ratio. At the same time, the adjustment of the oxygen flow rate is usually realized by a blower, and its response speed is relatively slow, while the electronic control response speed of the DC / DC converter is fast, which is likely to cause mismatch with the oxygen supply system, resulting in overshoot or steady-state error. It can be seen that the coordinated control of the oxygen flow rate and the duty ratio of the DC / DC converter involves multiple physical processes, and the system dynamic behavior is complex and highly non-linear. These problems make it difficult to achieve coordinated control between the oxygen flow rate and the duty ratio.
[0005] Therefore, efforts are still needed to find effective ways to improve the stability of the output voltage of the PEMFC power generation system. Summary of the Invention
[0006] Aiming at the deficiencies in the prior art, the present invention provides a method and device for coordinated control of the output voltage of a proton exchange membrane fuel cell power generation system, which simultaneously considers the influence of the air flow rate of the blower and the duty ratio of the DC / DC converter on the output voltage of the PEMFC, and can achieve coordinated control between the air flow rate in the blower and the duty ratio of the DC / DC converter in a non-linear and adaptive manner. The training process has good convergence, so as to enhance the robustness of the system, improve the tracking accuracy of the output voltage of the PEMFC, and further improve the voltage stability and overall operation efficiency.
[0007] To achieve the above object, the present invention provides the following technical solutions:
[0008] According to the first aspect of the present invention, a method for coordinated control of the output voltage of a proton exchange membrane fuel cell power generation system is provided. The method is applied to a proton exchange membrane fuel cell power generation system, and the proton exchange membrane fuel cell power generation system includes a DC / DC converter. The method includes:
[0009] Respectively establish a first intelligent agent for controlling the blower and a second intelligent agent for controlling the DC / DC converter, and respectively define the corresponding action space, state space and reward function;
[0010] Establish a leader selection training framework including a plurality of parallel training units, and each training unit includes a first intelligent agent and a second intelligent agent;
[0011] Train through the established leader selection training framework, so that the actions output by the plurality of training units interact with the environment, determine the leader training unit and the follower training unit, and then construct a leader-follower training framework;
[0012] Through the leader-follower training framework, the trained first agent and second agent are used to coordinately control the air flow rate of the blower and the duty cycle of the DC / DC converter respectively to obtain a stable output voltage.
[0013] Preferably, the first agent is used to control the air flow rate by controlling the opening degree of the air supply valve and the rotational speed of the blower;
[0014] The second agent controls the duty cycle of the DC / DC converter by controlling the conduction time of the switching device in the DC / DC converter.
[0015] Preferably, the action space of the first agent is defined as:
[0016]
[0017] where a oxy is the action space of the first agent, b is the opening degree of the air supply valve, and n is the rotational speed of the blower.
[0018] Preferably, the state space of the first agent is defined as:
[0019]
[0020] where s oxy is the state space of the first agent, v o (t - 1) is the air flow rate at the previous moment t - 1, v o (t) is the air flow rate at the current moment t, b(t - 1) is the opening degree of the air supply valve at the previous moment t - 1, n(t - 1) is the rotational speed of the blower at the previous moment t - 1, e s (t - 1) is the stack voltage error at the previous moment t - 1, e s (t) is the stack voltage error at the current moment t.
[0021] Preferably, the reward function of the first agent is defined as:
[0022]
[0023] where r oxy (t) is the reward function of the first agent, e o (t) is the output voltage error at the current moment t, v o (t) is the air flow rate at the current moment t, v o (t - 1) is the air flow rate at the previous moment t - 1, e s(t) is the stack voltage error at the current moment t, and w1, w2, and w3 are the weight coefficients of the first agent related to the output voltage error, air flow rate difference, and stack voltage error, respectively.
[0024] Preferably, the action space of the second agent is defined as:
[0025]
[0026] In the formula, a dc is the action space of the second agent, t on is the conduction time of the switching device in the DC / DC converter, and T is the switching period.
[0027] Preferably, the state space of the second agent is defined as:
[0028]
[0029] In the formula, s dc is the state space of the second agent, D dc (t - 1) is the duty cycle at the previous moment t - 1, D dc (t) is the duty cycle at the current moment t, t on (t - 1) is the conduction time of the switching device at the previous moment t - 1, e o (t - 1) is the output voltage error at the previous moment t - 1, e o (t) is the output voltage error at the current moment t.
[0030] Preferably, the reward function of the second agent is defined as:
[0031]
[0032] In the formula, r dc (t) is the reward function of the second agent, D dc (t - 1) is the duty cycle at the previous moment t - 1, D dc (t) is the duty cycle at the current moment t, e o (t) is the output voltage error at the current moment t, e s (t) is the stack voltage error at the current moment t, and μ1, μ2, and μ3 are the weight coefficients of the second agent related to the output voltage error, duty cycle difference, and stack voltage error, respectively.
[0033] Preferably, through the established leader selection training framework for training, the actions output by the multiple training units interact with the environment to determine the leader training unit and the follower training unit, and then a leader-follower training framework is constructed, including: performing a first data acquisition process and a first training process;
[0034] The first data acquisition process includes the following steps:
[0035] S31: Initialize the agent parameters of each agent in the leader selection training framework and obtain the initial state;
[0036] S32: Each agent selects an action according to the current state to interact with the environment;
[0037] S34: Store the data set generated by the interaction into the experience replay pool and obtain the next state;
[0038] S35: Determine whether the number of data in the experience replay pool reaches a preset threshold. If not, set the current state equal to the next state, and re-execute steps S32 to S34; if so, execute the first training process;
[0039] The first training process includes the following steps:
[0040] Each agent extracts a mini-batch data set from the experience replay pool;
[0041] Calculate the gradient using the extracted mini-batch data set, and update the parameters of each agent according to the gradient;
[0042] Based on the reward functions of the first agent and the second agent, determine the average reward values of the first agent and the second agent in each training unit, and select the training unit with the largest average reward index as the leader training unit. Use the training units in the leader selection training framework other than the leader training unit as follower training units.
[0043] Preferably, training is performed through a leader-follower training framework, and the trained first agent and second agent are used to coordinately control the air flow rate of the blower and the duty cycle of the DC / DC converter respectively to obtain a stable output voltage, including: performing a second data acquisition process and a second training process;
[0044] The second data acquisition process includes the following steps:
[0045] S41: Initialize the agent parameters of each agent in the leader-follower training framework and obtain the initial state;
[0046] S42: Select the load of each follower training unit;
[0047] S43: Each agent in each follower training unit selects an action according to the current state to interact with the environment;
[0048] S45: Store the data set generated by the interaction into the experience replay pool and obtain the next state;
[0049] S46: Determine whether the number of data in the experience replay pool reaches a preset threshold. If not, set the current state equal to the next state and restart from step S43 to step S46; if so, execute the second training process.
[0050] The second training process includes the following steps:
[0051] The leader training unit extracts a mini-batch data set from the experience replay pool;
[0052] The leader agent calculates the target value of each data set in the mini-batch data set and determines the gradient;
[0053] Update the agent parameters of each agent in the leader training unit according to the gradient, and send the agent parameters in the leader training unit to each agent in all follower training units for update to obtain the trained first agent and second agent;
[0054] Use the trained first agent and second agent to coordinately control the blower and the DC / DC converter to obtain a stable output voltage.
[0055] According to the second aspect of the present invention, there is provided a proton exchange membrane fuel cell power generation system output voltage coordination control device using the method described in the first aspect of the present invention. The device includes:
[0056] An agent establishment unit for respectively establishing a first agent for controlling the blower and a second agent for controlling the DC / DC converter, and respectively defining corresponding action spaces, state spaces and reward functions;
[0057] A first framework establishment unit for establishing a leader selection training framework including a plurality of parallel training units, each training unit including a first agent and a second agent;
[0058] A leader training unit for training through the established leader selection training framework, enabling the actions output by the plurality of training units to interact with the environment, determining the leader training unit and the follower training unit, and further constructing a leader-follower training framework;
[0059] An actual training unit for training through the leader-follower training framework, and coordinately controlling the air flow rate of the blower and the duty cycle of the DC / DC converter with the trained first agent and second agent to obtain a stable output voltage.
[0060] According to the third aspect of the present invention, there is provided an electronic device. The electronic device includes a processor and a storage medium;
[0061] The storage medium is used to store instructions;
[0062] The processor is configured to operate according to the instructions to perform the steps of the method according to any one of claims 1-7.
[0063] According to a fourth aspect of the present invention, there is provided a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and when the program is executed by a processor, the steps of the method according to the first aspect of the present invention are implemented.
[0064] Compared with the prior art, the beneficial effects of the present application are as follows:
[0065] By adopting the leader multi-agent deep deterministic policy gradient output voltage coordination control strategy, through leader selection training and leader-follower training, trained first and second agents can be obtained, which coordinate with each other to simultaneously coordinate the control of the air flow rate of the blower and the duty cycle of the DC / DC converter. The training process has good convergence, so as to improve the PEMFC output voltage control performance, enhance the robustness of the system, improve the tracking accuracy of the PEMFC output voltage, and further improve the voltage stability and overall operation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] In order to make the system structure and method flow of the present invention clearer, the system and method of the present invention are described in detail in conjunction with the following drawings, where:
[0067] Figure 1 It is a structural diagram of a proton exchange membrane fuel cell power generation system of the present invention.
[0068] Figure 2 It is a structural diagram of the anode hydrogen supply system of the proton exchange membrane fuel cell power generation system of the present invention.
[0069] Figure 3 It is a structural diagram of the cathode air supply system of the proton exchange membrane fuel cell power generation system of the present invention.
[0070] Figure 4 It is a structural diagram of the thermal management system of the proton exchange membrane fuel cell power generation system of the present invention.
[0071] Figure 5 It is an output voltage model of the proton exchange membrane fuel cell power generation system of the present invention.
[0072] Figure 6 It is a circuit structure of the DC / DC converter of the proton exchange membrane fuel cell power generation system of the present invention.
[0073] Figure 7 It is a flowchart of the output voltage coordination control method of the proton exchange membrane fuel cell power generation system of the present invention.
[0074] Figure 8This is a coupling relationship diagram between the air flow rate of the blower and the duty ratio of the DC / DC converter in the output voltage coordinated control method of the proton exchange membrane fuel cell power generation system of the present invention.
[0075] Figure 9 This is a diagram of agent definition in the output voltage coordinated control method of the proton exchange membrane fuel cell power generation system of the present invention.
[0076] Figure 10 This is a leader selection training framework diagram in the output voltage coordinated control method of the proton exchange membrane fuel cell power generation system of the present invention.
[0077] Figure 11 This is a leader selection training flow chart in the output voltage coordinated control method of the proton exchange membrane fuel cell power generation system of the present invention.
[0078] Figure 12 This is an actual training framework diagram in the output voltage coordinated control method of the proton exchange membrane fuel cell power generation system of the present invention.
[0079] Figure 13 This is an actual training flow chart in the output voltage coordinated control method of the proton exchange membrane fuel cell power generation system of the present invention. Detailed implementation manners
[0080] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. The embodiments described in this application are only a part of the embodiments of the present invention, rather than all embodiments. Based on the spirit of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0081] According to the first aspect of the present invention, there is provided a method for coordinating and controlling the output voltage of a proton exchange membrane fuel cell power generation system.
[0082] As Figure 1 shown, in one embodiment, the proton exchange membrane fuel cell power generation system based on this method includes: a hydrogen supply system, an air supply system, a thermal management system, a fuel cell stack, and a DC / DC converter. Among them, the hydrogen supply system, the air supply system, and the thermal management system are connected to the fuel cell stack as auxiliary subsystems, and the stack voltage is converted by the DC / DC converter to obtain the output voltage.
[0083] As Figure 2 shown, the hydrogen supply system consists of a high-pressure hydrogen storage tank, a hydrogen valve, a filter, and a humidifier. Hydrogen is released from the high-pressure hydrogen tank through the hydrogen valve, and after passing through the filter and the humidifier, it is sent to the anode to participate in the reaction.
[0084] AsFigure 3 As shown in the figure, the air supply system includes a blower, a filter, a flow meter, an air compressor, an air supply pipeline, a humidifier, a cooler, and an air return pipeline. The air in the atmosphere enters the supply pipeline through the blower, filter, flow meter, and air compressor. The high-temperature and dry air in the supply pipeline enters the cathode of the fuel cell stack to participate in the reaction after passing through the humidifier and the cooler. The remaining air after the reaction will flow out from the cathode and enter the return pipeline, and finally be discharged into the atmosphere.
[0085] As Figure 4 shown in the figure, the thermal management system consists of a water pump, a thermostat, a heater, a radiator, a filter, a deionized water tank, a compensation water tank, an intercooler, and a hydrogen heat exchange plate. Each component cooperates with each other to jointly control the temperature of the fuel cell stack. Heating process: When the coolant temperature is relatively low, the thermostat closes the radiator channel (meanwhile the heater works simultaneously), opens the heater channel. The coolant pressurized by the water pump is heated in the heater, and after passing through the thermostat, it is divided into two paths. One path enters the flow field of the fuel cell stack, and after heating the fuel cell stack, it returns to the water inlet of the water pump; the other path enters the series water circuit of the intercooler and the hydrogen heat exchange plate, and then returns to the water inlet of the water pump. Cooling process: When the coolant temperature is relatively high and heat dissipation is required, the thermostat closes the heater channel and opens the radiator channel. The coolant pressurized by the water pump is cooled in the radiator, and after passing through the thermostat, it is divided into two paths. One path enters the flow field of the fuel cell stack, and after cooling the fuel cell stack, it returns to the water inlet of the water pump; the other path enters the series water circuit of the intercooler and the hydrogen heat exchange plate, and then returns to the water inlet of the water pump.
[0086] The fuel cell stack is composed of multiple single fuel cells stacked in series. During the operation of a single fuel cell, a series of physical and chemical changes will occur on the surface of the battery electrodes. Each change has a certain resistance, and energy must be consumed to overcome this resistance, resulting in the actual output voltage of the proton exchange membrane fuel cell being less than the ideal output voltage.
[0087] As Figure 5 shown in the figure, the calculation steps for the actual output voltage of a proton exchange membrane fuel cell monomer are as follows:
[0088] V cell = E N - V a - V oh - V c
[0089] In the formula, V cell is the actual output voltage of the fuel cell monomer, E N is the ideal output voltage that a single fuel cell theoretically reaches, V a is the activation overvoltage, representing the voltage loss caused by the slow reaction on the electrode surface, V oh is the ohmic overvoltage, representing the voltage loss generated by the ohmic resistance, Vc is the concentration overvoltage, representing the voltage loss caused by the change in the concentration of reactants on the electrode surface.
[0090] For a fuel cell stack composed of N fuel cell monomers connected in series, its voltage is:
[0091] V st = N·V cell
[0092] Wherein, V st is the stack voltage of the fuel cell stack, V cell is the actual output voltage of the fuel cell monomer, and N is the number of fuel cell monomers included in the fuel cell stack.
[0093] The DC / DC converter adopts a Boost circuit, and the circuit structure is as Figure 6 shown. The Boost boost circuit is a DC / DC conversion circuit, that is, a switching DC boost circuit, which converts a lower DC voltage into a higher DC voltage. The Boost boost circuit can not only increase the voltage, but also achieve the voltage stabilization function of the output voltage within a certain range.
[0094] In this embodiment, the multi-agent deep deterministic policy gradient algorithm is adopted, and through leader selection training and leader-follower training, the air flow rate of the blower in the air supply system and the duty cycle of the DC / DC converter are coordinately controlled to obtain a stable output voltage.
[0095] Specifically, as Figure 7 shown, the method in this embodiment includes the following steps:
[0096] S1: Respectively establish a first agent for controlling the blower and a second agent for controlling the DC / DC converter, and respectively define the corresponding action space, state space and reward function.
[0097] As Figure 8 shown, there is a coupling relationship between the blower and the DC / DC converter at the output voltage level. The Nernst voltage in the stack voltage is related to the oxygen partial pressure, the activation overvoltage in the stack voltage is related to the oxygen concentration, and the oxygen partial pressure and oxygen concentration are related to the oxygen flow rate. Therefore, the stack voltage is affected by the oxygen flow rate in the blower. The stack voltage is converted by the DC / DC converter to obtain the output voltage. When the oxygen flow rate in the blower causes the stack voltage to fluctuate, the output voltage fluctuates accordingly. By reasonably adjusting the duty cycle in real time through the DC / DC converter control strategy, the coordination between the blower and the DC / DC converter is realized to obtain a stable output voltage.
[0098] Considering the coordination between the blower controller and the DC / DC converter controller, as Figure 9As shown, the blower controller and the DC / DC converter controller are regarded as the first intelligent agent oxy and the second intelligent agent dc , and the two coordinate with each other through a centralized training and distributed execution strategy.
[0099] In this embodiment, the first intelligent agent is used to control the air flow rate by controlling the opening degree of the air supply valve and the rotational speed of the blower, and the second intelligent agent is used to control the duty cycle of the DC / DC converter by controlling the conduction time of the switching device.
[0100] Next, define the action space, state space, and reward function of the intelligent agent (agent oxy and agent dc ).
[0101] The action space of the first intelligent agent oxy is defined as:
[0102]
[0103] In the formula, a oxy is the action space of the first intelligent agent, b is the opening degree of the air supply valve, and n is the rotational speed of the blower.
[0104] The state space of the first intelligent agent oxy is defined as:
[0105]
[0106] In the formula, s oxy is the state space of the first intelligent agent, v o (t - 1) is the air flow rate at the previous moment t - 1, v o (t) is the air flow rate at the current moment t, b(t - 1) is the opening degree of the air supply valve at the previous moment t - 1, n(t - 1) is the rotational speed of the blower at the previous moment t - 1, e s (t - 1) is the stack voltage error at the previous moment t - 1, and e s (t) is the stack voltage error at the current moment t.
[0107] The reward function of the first intelligent agent oxy is defined as:
[0108]
[0109] In the formula, r oxy (t) is the reward function of the first intelligent agent, e o (t) is the output voltage error at the current moment t, and v oThe air flow velocity at the current moment \(t\) is \(v\). o The air flow velocity at the previous moment \(t - 1\) is \(e\). s The stack voltage error at the current moment \(t\) is \(w_1\), \(w_2\), and \(w_3\) are the weight coefficients of the first agent related to the output voltage error, air flow velocity difference, and stack voltage error respectively.
[0110] The second agent dc The action space is defined as:
[0111]
[0112] where \(a\ dc is the action space of the second agent, \(t\ on is the conduction time of the switching device in the DC / DC converter, and \(T\) is the switching period.
[0113] The second agent dc The state space is defined as:
[0114]
[0115] where \(s\ dc is the state space of the second agent, \(D\ dc \((t - 1)\) is the duty cycle at the previous moment \(t - 1\), \(D\ dc \((t)\) is the duty cycle at the current moment \(t\), \(t\ on \((t - 1)\) is the conduction time of the switching device at the previous moment \(t - 1\), \(e\ o \((t - 1)\) is the output voltage error at the previous moment \(t - 1\), \(e\ o \((t)\) is the output voltage error at the current moment \(t\).
[0116] The second agent dc The reward function is defined as:
[0117]
[0118] where \(r\ dc \((t)\) is the reward function of the second agent, \(D\ dc \((t - 1)\) is the duty cycle at the previous moment \(t - 1\), \(D\ dc \((t)\) is the duty cycle at the current moment \(t\), \(e\ o \((t)\) is the output voltage error at the current moment \(t\), \(e\ s \((t)\) is the stack voltage error at the current moment \(t\), and \(\mu_1\), \(\mu_2\), \(\mu_3\) are the weight coefficients of the second agent related to the output voltage error, duty cycle difference, and stack voltage error respectively.
[0119] S2. Establish a leader selection training framework including multiple parallel training units, where each training unit includes a first agent and a second agent respectively.
[0120] As an example, as Figure 10 shown, in step S4, the leader selection training framework of this embodiment may include 40 parallel systems, and each parallel system contains two agents (i.e., the first agent oxy and the second agent dc ).
[0121] S3. Train through the established leader selection training framework, so that the actions output by the multiple training units interact with the environment, determine the leader training unit and the follower training unit, and then construct a leader-follower training framework.
[0122] This step generally includes a first data collection process and a first training process.
[0123] As Figure 11 shown, in step S3:
[0124] The first data collection process includes the following steps:
[0125] S31: Initialize the agent parameters of each agent in the leader selection training framework and obtain the initial state as the current state.
[0126] The agent parameters may include the learning rate, discount factor, exploration rate, and soft update coefficient.
[0127] The learning rate refers to the degree of controlling the agent to adjust its policy or value function at each update.
[0128] The discount factor refers to the current value used to calculate the long-term reward, which determines the importance of future rewards for the current decision.
[0129] The exploration rate refers to determining the balance between the agent exploring new strategies and exploiting known information.
[0130] The soft update coefficient refers to being used for smooth transition in the target network update to avoid drastic fluctuations during the training process.
[0131] In reinforcement learning, the state is a description of the current situation of the environment, and the agent selects actions based on the state.
[0132] S32: Each agent selects an action according to the current state to interact with the environment.
[0133] An action is the influence exerted by the agent on the environment, aiming to change the state of the environment to obtain a reward.
[0134] After the agent executes an action, the environment will give a new state and a reward based on the action and the current state. The reward is the immediate feedback of the environment to the agent's action, and is usually used to indicate the quality of the action.
[0135] S33: Store the dataset generated by the interaction in the experience replay pool and obtain the next state.
[0136] The dataset generated by the interaction can include a collection of multiple data samples, where each data sample includes the current state, the current action, the reward corresponding to the current state-action, and the next state.
[0137] The experience replay pool is a data structure used to store the agent's experiences, which helps to break the correlation between samples and improve the learning efficiency.
[0138] The agent obtains the next state of the environment as the basis for the next decision.
[0139] S34: Determine whether the number of data in the experience replay pool reaches a preset threshold. If not, set the current state equal to the next state and re-execute steps S32 to S34; if so, execute the first training process.
[0140] In this step, the threshold is usually set to ensure that there are enough samples for training to avoid overfitting or underfitting.
[0141] The first training process includes:
[0142] S35: Each agent extracts a small batch of datasets from the experience replay pool.
[0143] In this step, the small batch of datasets refers to multiple data samples randomly extracted from the experience replay pool. The number of these multiple data samples is usually less than the total number of data samples in the experience replay pool. This small batch of datasets can be used for subsequent gradient calculation and parameter update.
[0144] S36: Each agent calculates the target value of each dataset in the small batch of datasets to determine the gradient.
[0145] In this step, the agent can use its current policy or value function to calculate the target value of each data sample in the small batch of datasets. The target value is usually calculated based on the optimal estimated value of the next state. For example, it can usually be calculated using methods such as the Bellman equation or temporal difference learning.
[0146] By comparing the predicted value of the agent (i.e., the reward obtained after executing the action plus the estimated value of the next state) and the target value, the agent can calculate the loss function and determine the gradient accordingly.
[0147] S37: Each agent updates its respective agent parameters according to the gradient.
[0148] In this step, the agent can use methods such as gradient descent (or its variants, such as stochastic gradient descent, Adam optimizer, etc.) to update the parameters of its policy or value function according to the calculated gradient. This update process aims to minimize the loss function, thereby improving the prediction accuracy or policy quality of the agent.
[0149] S38: Based on the reward functions of the first agent and the second agent, determine the average reward values of the first agent and the second agent in each training unit, select the training unit with the largest average reward value as the leader training unit, and use the training units other than the leader training unit in the leader selection training framework as the follower training units.
[0150] In this step, the agent in the leader training unit is the first leader agent Leader oxy and the second leader agent Leader DC .
[0151] Such as Figure 12 shown, in step S6 of this embodiment, the actual training framework includes a leader training unit (Leader oxy and Leader DC ) and 39 parallel follower training units. The leader agent is the agent in the training unit with the highest average reward value obtained through leader selection training. In actual training, the agent in the leader training unit does not participate in the data collection process, but is responsible for: randomly extracting a data set from the experience replay pool for learning; updating the parameters of its own actor; and sending the parameters to the agents in all follower training units for updating.
[0152] In the 39 parallel training units, each parallel training unit contains two agents (agent oxy and agent DC ).
[0153] Through the parallel leader training unit and follower training units, a leader-follower training framework can be constructed.
[0154] S4: Train through the leader-follower training framework, and use the trained first agent and second agent to coordinately control the air flow rate of the blower and the duty cycle of the DC / DC converter respectively to obtain a stable output voltage.
[0155] Step S4 generally includes a second data collection process and a second training process.
[0156] Such as Figure 13 shown, in step S4 of this embodiment, the second data collection process includes:
[0157] S41: Initialize the agent parameters of each agent in the leader-follower training framework and obtain the initial state.
[0158] Similar to step S31, the agent parameters here may include the learning rate, discount factor, exploration rate, and soft update coefficient.
[0159] S42: Select the load for each follower training unit.
[0160] In this step, the load refers to the specific tasks or workloads assigned to each follower training unit in the leader-follower training framework, which can be understood as the actual problems or challenges that the agent needs to handle and solve. It may involve a series of actions that the agent needs to take, goals to be achieved, or performance metrics to be optimized, etc.
[0161] The selection of the load may affect the training effect and performance of the agent.
[0162] S43: Each agent in each follower training unit selects an action according to the current state to interact with the environment.
[0163] In this step, each follower agent will select an action according to the current environmental state and its own policy. This action will cause the agent to interact with the environment and generate corresponding results (such as a new state, reward, etc.).
[0164] S45: Store the dataset generated by the interaction in the experience replay pool and obtain the next state.
[0165] The dataset here is similar to the dataset in step S33 and also includes a collection of multiple data samples, where each data sample includes the current state, current action, reward corresponding to the current state-action, and the next state.
[0166] S46: Determine whether the number of data in the experience replay pool reaches a preset threshold. If not, set the current state equal to the next state and re-execute steps S43 to S46; if so, execute the second training process.
[0167] The second training process includes:
[0168] S47: The leader training unit extracts a mini-batch dataset from the experience replay pool.
[0169] S48: The leader agent calculates the target value of each dataset in the mini-batch dataset and determines the gradient.
[0170] S49: Update the agent parameters of each agent in the leader training unit according to the gradient, and send the agent parameters in the leader training unit to each agent in all follower training units for update, so as to obtain the trained first agent and second agent.
[0171] Through this step, the leader and follower agents can work together to improve the performance of the entire system. After multiple iterations of training, the trained first agent and second agent can be obtained.
[0172] S410: Use the trained first agent and second agent to coordinately control the blower and the DC / DC converter to obtain a stable output voltage.
[0173] Through this step, the agent can dynamically adjust the control parameters of the blower and the DC / DC converter according to the current state of the system (such as input voltage, load change, etc.) to achieve a stable output voltage. This coordinated control can improve the efficiency and stability of the system, while reducing the risk of energy waste and equipment damage.
[0174] According to the second aspect of the present invention, there is provided a device for coordinating the output voltage of a proton exchange membrane fuel cell power generation system using the method described in the first aspect of the present invention, characterized by comprising:
[0175] An agent establishment unit, configured to respectively establish a first agent for controlling a blower and a second agent for controlling a DC / DC converter, and respectively define corresponding action spaces, state spaces and reward functions;
[0176] A first framework establishment unit, configured to establish a leader selection training framework including a plurality of parallel training units, and each training unit includes a first agent and a second agent respectively;
[0177] A leader training unit, configured to perform training through the established leader selection training framework, enable the actions output by the plurality of training units to interact with the environment, determine the leader training unit and the follower training unit, and further construct a leader-follower training framework;
[0178] An actual training unit, configured to perform training through the leader-follower training framework, and use the trained first agent and second agent to respectively coordinately control the air flow rate of the blower and the duty ratio of the DC / DC converter to obtain a stable output voltage.
[0179] According to a third aspect of the present invention, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, it implements the output voltage coordinated control method of the proton exchange membrane fuel cell power generation system according to the first aspect of the present invention.
[0180] According to a fourth aspect of the present invention, there is provided a computer-readable storage medium storing a computer program, which when executed by a processor, implements the output voltage coordinated control method of the proton exchange membrane fuel cell power generation system according to the first aspect of the present invention.
[0181] The beneficial effects of the present invention are as follows. Compared with the prior art, by adopting the leader multi-agent deep deterministic policy gradient output voltage coordinated control strategy, through leader selection training and leader-follower training, well-trained first and second agents can be obtained, which coordinate with each other to simultaneously coordinate the control of the air flow rate of the blower and the duty cycle of the DC / DC converter. The training process has good convergence, so as to improve the PEMFC output voltage control performance, enhance the robustness of the system, improve the tracking accuracy of the PEMFC output voltage, and further improve the voltage stability and overall operation efficiency.
[0182] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0183] The computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punched card or raised structures in grooves storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as an instantaneous signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., optical pulses through an optical fiber cable), or electrical signals transmitted through wires.
[0184] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0185] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.
[0186] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: still modifications or equivalent substitutions can be made to the specific embodiments of the present invention, and any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A method for coordinated control of output voltage of a proton exchange membrane fuel cell power generation system, characterized in that: The method is applied to a proton exchange membrane fuel cell power generation system, wherein the proton exchange membrane fuel cell power generation system includes a DC / DC converter; the method includes: Establishing a first agent for controlling the blower and a second agent for controlling the DC / DC converter, respectively, and defining corresponding action spaces, state spaces and reward functions, respectively; Establishing a leader selection training framework including a plurality of training units in parallel, each training unit including a first agent and a second agent; Training is performed by establishing a leader selection training framework, so that the actions output by the multiple training units interact with the environment, a leader training unit and a follower training unit are determined, and then a leader-follower training framework is constructed; The training is performed through a leader-follower training framework, and the trained first agent and the second agent are used to coordinately control the air flow rate of the blower and the duty cycle of the DC / DC converter respectively to obtain a stable output voltage.
2. The method for coordinated control of output voltage of a proton exchange membrane fuel cell power generation system according to claim 1, characterized in that: The first intelligent agent is used to control the air flow rate by controlling the opening and closing degree of the air supply valve and the speed of the blower; The second intelligent agent controls the duty cycle of the DC / DC converter by controlling the conduction time of the switch device in the DC / DC converter.
3. The method for coordinated control of output voltage of a proton exchange membrane fuel cell power generation system according to claim 1, characterized in that: The action space of the first agent is defined as: In the formula, a oxy is the action space of the first agent, b is the opening and closing degree of the air supply valve, and n is the blower speed.
4. The method for coordinated control of output voltage of a proton exchange membrane fuel cell power generation system according to claim 1, characterized in that: The state space of the first agent is defined as: In the formula, s oxy is the state space of the first agent, v o (t-1) is the air velocity at the previous moment t-1, v o (t) is the air flow rate at the current time t, b(t-1) is the opening and closing degree of the air supply valve at the previous time t-1, n(t-1) is the blower speed at the previous time t-1, e s (t-1) is the stack voltage error at the previous moment t-1, e s (t) is the stack voltage error at the current time t.
5. The method for coordinated control of output voltage of a proton exchange membrane fuel cell power generation system according to claim 1, characterized in that: The reward function of the first agent is defined as: In the formula, r oxy (t) is the reward function of the first agent, e o (t) is the output voltage error at the current time t, v o (t) is the air velocity at the current time t, v o (t-1) is the air velocity at the previous moment t-1, e s (t) is the stack voltage error at the current time t, w1, w2, w3 are the weight coefficients of the first intelligent agent related to the output voltage error, air flow rate difference and stack voltage error respectively.
6. The method for coordinated control of output voltage of a proton exchange membrane fuel cell power generation system according to claim 1, characterized in that: The action space of the second agent is defined as: In the formula, a dc is the action space of the second agent, t on is the conduction time of the switching device in the DC / DC converter, and T is the switching period.
7. The method for coordinated control of output voltage of a proton exchange membrane fuel cell power generation system according to claim 1, characterized in that: The state space of the second agent is defined as: In the formula, s dc is the state space of the second agent, D dc (t-1) is the duty cycle at the previous moment t-1, D dc (t) is the duty cycle at the current time t, t on (t-1) is the conduction time of the switch device at the previous moment t-1, e o (t-1) is the output voltage error at the previous moment t-1, e o (t) is the output voltage error at the current time t.
8. The method for coordinated control of output voltage of a proton exchange membrane fuel cell power generation system according to claim 1, characterized in that: The reward function of the second agent is defined as: In the formula, r dc (t) is the reward function of the second agent, D dc (t-1) is the duty cycle at the previous moment t-1, D dc (t) Duty cycle at the current time t, e o (t) is the output voltage error at the current time t, e s (t) is the stack voltage error at the current time t, μ1, μ2, and μ3 are the weight coefficients of the second intelligent agent related to the output voltage error, duty cycle difference, and stack voltage error, respectively.
9. The method for coordinated control of output voltage of a proton exchange membrane fuel cell power generation system according to claim 1, characterized in that: The method includes: performing training by establishing a leader selection training framework, making the actions output by the plurality of training units interact with the environment, determining a leader training unit and a follower training unit, and then constructing a leader-follower training framework, including: executing a first data collection process and a first training process; The first data collection process comprises the following steps: S31: Initialize the leader to select agent parameters of each agent in the training framework and obtain the initial state as the current state; S32: Each agent selects an action based on its current state to interact with the environment; S33: Store the data set generated by the interaction into the experience replay pool and obtain the next state; S34: Determine whether the number of data in the experience replay pool reaches a preset threshold. If not, set the current state equal to the next state and re-execute steps S32 to S34. If reached, execute the first training process. The first training process comprises the following steps: Each agent draws a mini-batch of data from the experience replay pool and calculates the target value; Each agent calculates the target value for each dataset in the mini-batch to determine the gradient; Each agent updates its own agent parameters based on the gradient; Based on the reward function of the first agent and the second agent, the average reward value of the first agent and the second agent in each training unit is determined, and the training unit with the largest average reward value is selected as the leader training unit, and the training units except the leader training unit in the leader selection training framework are used as follower training units.
10. The method for coordinated control of output voltage of a proton exchange membrane fuel cell power generation system according to claim 1, characterized in that: The method comprises the steps of: performing training through a leader-follower training framework, using the trained first agent and the second agent to coordinately control the air flow rate of the blower and the duty cycle of the DC / DC converter respectively, and obtaining a stable output voltage, including: performing a second data acquisition process and a second training process; The second data collection process comprises the following steps: S41: Initialize the agent parameters of each agent in the leader-follower training framework and obtain the initial state; S42: Selecting the load of each follower training unit; S43: Each agent in each follower training unit selects an action based on the current state to interact with the environment; S45: storing the data set generated by the interaction into the experience replay pool and obtaining the next state; S46: Determine whether the number of data in the experience replay pool reaches a preset threshold. If not, set the current state equal to the next state and re-execute steps S43 to S46; if reached, execute the second training process; The second training process comprises the following steps: The leader training unit draws a mini-batch of data from the experience replay pool; The leader agent calculates the target value for each dataset in the mini-batch and determines the gradient; Update the agent parameters of each agent in the leader training unit according to the gradient, and send the agent parameters in the leader training unit to the agents in all follower training units for updating, so as to obtain the trained first agent and the second agent; The blower and the DC / DC converter are coordinated and controlled by the trained first agent and the second agent to obtain a stable output voltage.
11. A device for coordinating and controlling the output voltage of a proton exchange membrane fuel cell power generation system using the method according to any one of claims 1 to 10, characterized in that: include: An agent establishment unit, used to respectively establish a first agent for controlling the blower and a second agent for controlling the DC / DC converter, and define corresponding action spaces, state spaces and reward functions respectively; A first framework establishment unit, for establishing a leader selection training framework including a plurality of training units in parallel, each training unit including a first agent and a second agent; A leader training unit, used to perform training through an established leader selection training framework, so that the actions output by the plurality of training units interact with the environment, determine a leader training unit and a follower training unit, and then construct a leader-follower training framework; The actual training unit is used for training through a leader-follower training framework, and uses the trained first agent and the second agent to coordinate and control the air flow rate of the blower and the duty cycle of the DC / DC converter respectively to obtain a stable output voltage.
12. An electronic device, comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the method according to any one of claims 1-10.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.