Fuel cell hybrid propulsion system and its power fluctuation-consumption co-optimization method

By constructing an intelligent agent for a fuel cell hybrid propulsion system and utilizing a near-end strategy optimization algorithm to allocate battery power in real time, the problem of coordinating power fluctuations and energy consumption during land-to-air transitions in hybrid propulsion systems was solved, achieving efficient and smooth system operation and extended battery life.

CN121697510BActive Publication Date: 2026-06-30NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
Filing Date
2026-02-06
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

During the dynamic transition from land to air in existing hybrid propulsion systems, traditional energy management methods struggle to coordinate and optimize battery power fluctuations and overall system energy consumption, leading to power surges and energy waste.

Method used

A fuel cell hybrid propulsion system is adopted, which combines a fuel cell power generation unit, a lithium-ion battery unit, an electric drive unit, a sensing and data acquisition unit, and an energy management controller. An intelligent agent is constructed through a near-end strategy optimization algorithm to allocate battery power in real time to optimize power fluctuations and energy consumption.

Benefits of technology

It achieves efficient and smooth operation of fuel cell hybrid propulsion system in dynamic environment, reduces component fatigue wear, extends battery life and improves energy conversion efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121697510B_ABST
    Figure CN121697510B_ABST
Patent Text Reader

Abstract

This invention discloses a fuel cell hybrid propulsion system and its power fluctuation-consumption synergistic optimization method. The system includes a fuel cell power generation unit, a lithium-ion battery unit, an electric drive unit, a sensing and data acquisition unit, and an energy management controller. Based on near-end strategy optimization, an intelligent agent is constructed with system state and operating condition information as input and power source power allocation commands as output. A reward function integrating total energy consumption penalty and power fluctuation penalty is designed to drive the intelligent agent to learn the optimal energy management strategy in a continuous action space. This invention can suppress frequent and drastic fluctuations in lithium-ion battery power, extend battery life, ensure system safety, and minimize overall hydrogen fuel consumption while meeting dynamic power demand and safety constraints. It achieves synergistic optimization of long-term operational economy and system output stability, enabling efficient, smooth, and long-life operation of the hybrid propulsion system in dynamic environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of energy management of hybrid propulsion systems, and specifically relates to a fuel cell hybrid propulsion system and its power fluctuation-consumption co-optimization method. Background Technology

[0002] Hybrid propulsion systems face severe challenges during dynamic transitions between land and air transport. An inherent contradiction exists between instantaneous high-power demands and sustained cruise efficiency. Traditional rule-based or threshold-based energy management methods struggle to simultaneously meet power response requirements while maintaining overall energy efficiency and output stability. This leads to power surges during mode switching, increasing component stress and reducing overall economic efficiency. Therefore, researching methods for synergistic optimization of power fluctuations and energy consumption has become a core technology for improving the reliability and economy of hybrid propulsion systems, and is of significant practical importance for promoting the commercialization of urban air mobility.

[0003] Research on energy management for hybrid propulsion systems remains an area requiring further exploration. Existing technical solutions provide valuable references and foundations for further research in this field. For example, Chinese invention patent application CN202411745253.1, entitled "A Power Matching Optimization Method for Land-Air Conversion in a Hybrid Propulsion System," achieves a smooth transition in power demand during the land-air conversion of a flying car by dynamically adjusting the power ratio of the motor and engine. Its core is solving the instantaneous power matching problem during mode switching. However, this solution primarily focuses on short-term power balance and fails to coordinate the optimization of total energy consumption and system output volatility from a global perspective. Furthermore, Chinese invention patent CN202010043003.9, entitled "A Fuel Cell Hybrid Propulsion System and Control Method for Underwater Vehicles," proposes managing fuel cell operating mode switching through hysteresis control based on lithium battery state of charge, achieving efficiency optimization of the underwater propulsion system. However, its control strategy is essentially a segmented management based on fixed thresholds. In flying car scenarios, the land-to-air transition mode may cause frequent jumps in fuel cell output power near the threshold, resulting in system power fluctuations. In summary, current energy management methods for hybrid propulsion systems generally suffer from the following common problems: current energy management methods for hybrid propulsion systems, especially those for dynamic land-to-air transition applications, mostly rely on rule-based or fixed-threshold control strategies. These methods lack adaptability under dynamic and continuous complex operating conditions and struggle to achieve integrated, adaptive, and real-time optimization of the system's overall total energy consumption and power output volatility. Summary of the Invention

[0004] In view of the shortcomings of the prior art, and in order to solve the problem that the prior art cannot coordinately optimize battery power fluctuation and overall system energy consumption, the purpose of this invention is to provide a fuel cell hybrid propulsion system and its power fluctuation-consumption co-optimization method, so as to achieve efficient, smooth and long-life operation of the hybrid propulsion system in dynamic environments.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0006] A fuel cell hybrid propulsion system is characterized by comprising: a fuel cell power generation unit, a lithium-ion battery unit, an electric drive unit, a sensing and data acquisition unit, and an energy management controller. The sensing and data acquisition unit is used to acquire real-time state data required by the energy management controller, which is configured to execute the aforementioned power fluctuation-consumption co-optimization method based on near-end strategy optimization to achieve optimal real-time allocation of output power between the fuel cell power generation unit and the lithium-ion battery unit.

[0007] The fuel cell power generation unit includes: a fuel cell stack and its associated hydrogen supply system, air supply system, thermal management system and DC-DC converter, used to provide the main driving range power.

[0008] The lithium-ion battery unit includes: a high-power-density lithium-ion battery pack and its management system, used to provide peak power, recover braking energy and smooth load fluctuations.

[0009] The electric drive unit includes a drive motor, a motor controller, and a propeller, used to convert electrical energy into propulsion power.

[0010] The sensing and data acquisition unit includes: a sensor computing module, a battery management system, current and voltage sensors, and a flight attitude sensor.

[0011] The sensor computing module is used to monitor the overall power requirements of the machine.

[0012] Battery management system is used to monitor the state of charge of lithium-ion battery cells;

[0013] Current and voltage sensors are used to monitor the real-time output power of fuel cell power generation units and lithium-ion battery units;

[0014] Flight attitude sensors are used to sense flight attitude, airspeed, and ambient wind speed, providing input for simulation environment models.

[0015] The energy management controller includes: an environment simulation and interaction module, a near-end strategy optimization algorithm module, and a power allocation module.

[0016] The environment simulation and interaction module is responsible for building a digital twin simulation environment within the controller to simulate the operation of a hybrid propulsion system in a dynamic wind field.

[0017] The near-end policy optimization algorithm module is responsible for implementing the near-end policy optimization algorithm and driving the controller to perform self-learning and optimization.

[0018] The power allocation module is used to apply the trained model to real-time control.

[0019] A method for co-optimizing power fluctuation and energy consumption in a fuel cell hybrid propulsion system includes the following steps:

[0020] Step 1): Real-time acquisition of system status data, including overall power demand, battery state of charge, and current power of fuel cells and lithium-ion batteries. Inputting the agent's action commands (defined as changes in battery power) into the environment, a simulation environment for the hybrid propulsion system in a dynamic wind field is constructed. The dynamic behavior of the hybrid propulsion system is simulated through environmental simulation and interaction modules. Based on this, a reward function reflecting energy consumption and fluctuation levels is designed, and all interaction information is stored in an experience pool to provide data support for subsequent optimization learning.

[0021] Step 2): Based on the interaction data collected in Step 1), construct and train the policy network and the value network. The policy network generates a probability distribution of battery power adjustment actions according to the current state, while the value network evaluates the long-term value of the state. The merits of a single step action are quantified by calculating the advantage function, and the objective function is optimized using the pruning mechanism unique to proximal policy optimization. This ensures the stability of the training process and enables the algorithm to effectively optimize both power fluctuation and energy consumption objectives.

[0022] Step 3): Based on the pruning optimization objective function and the loss function of the value network obtained in Step 2), the gradient descent algorithm is used to update the parameters of both the policy network and the value network simultaneously. By repeating Steps 1) to 3), a closed-loop iterative process of "data collection - network optimization - parameter update" is formed until the policy performance converges, ultimately resulting in an optimization decision model that can intelligently balance power fluctuations and energy consumption.

[0023] Step 4): Solidify the parameters of the policy network model trained and converged in Step 3) and deploy them to the actual energy management controller. During real-time flight, the controller outputs the optimal battery power adjustment command directly through this network based on the current system state, and coordinates the allocation with the fuel cell power to achieve online real-time optimization control that suppresses power fluctuations and minimizes overall energy consumption under dynamic conditions.

[0024] Further, step 1) specifically includes:

[0025] Based on the current state st and current action a t Calculate the instant reward R t and the next state s t+1 The information is then stored in an experience pool. Optionally, in this invention, the state variable is set as the power demand P. dem Battery capacity (SOC), fuel cell power (P) fc and lithium-ion battery power P b a t Set to ΔP b (That is, the difference between the lithium-ion battery power and the power of the previous second). R provided by the environment. t Calculated in the following way:

[0026]

[0027] Among them, R ecms (t) is the reward function based on the strategy of minimizing equivalent costs; R sm (t) represents the smooth control reward; R c (t) is the penalty for violating the constraint, which is a large negative constant when the constraint is violated.

[0028] R ecms (t) and R sm (t) uses an exponential function and a piecewise defined function, respectively. Optionally, in this invention, R... ecms The design of (t) is as follows:

[0029]

[0030] Where C1 and C2 represent the prices of hydrogen and electricity, respectively; f fc (t) and f b (t) represents the hydrogen and electricity consumption at time t, respectively; λ is the equivalent consumption factor, where a higher λ and a lower λ indicate that the hybrid power system is more inclined to use hydrogen and electricity, respectively.

[0031] Optionally, in this invention, R sm The design of (t) is as follows:

[0032]

[0033] Where K1 is the smoothing control weight coefficient; C3 is the fluctuation range of the lithium-ion battery; the reward function gives positive or negative rewards for ΔPb that is greater than or less than C3.

[0034] R c (t) is calculated by the following formula:

[0035]

[0036] Where K2 is the penalty value for exceeding the performance constraints of the hybrid propulsion system;

[0037] The performance constraints of the hybrid propulsion system are as follows:

[0038]

[0039] Among them, P fc,max This is the maximum power of a hydrogen fuel cell; P b,min and P b,max These are the minimum and maximum power of a lithium-ion battery; SOC. min and SOC max These represent the minimum and maximum states of charge, respectively.

[0040] The smoothing control weight coefficient K1 and the penalty value K2 for exceeding the constraint are key hyperparameters in the reward function, which are obtained through a grid search method.

[0041] Further, step 2) involves constructing and training the policy network and value network based on the interaction data collected in step 1), specifically including:

[0042] 21) Estimate action probabilities and state values;

[0043] 22) Optimize the value network;

[0044] 23) Pruning and optimizing the objective function.

[0045] Further, step 21) specifically estimates the action probability through an action policy network. This network uses the current system state s as an example. t As input, after calculating its parameter θ, the output is an action probability density function. For ΔP b For these continuous actions, the output is specifically represented by the parameters (mean and variance) of a Gaussian distribution, thus defining the probability distribution of the available actions. Simultaneously, the value of a state is evaluated through a value network. This network operates with the same state s. t As input, and through the calculation of its parameter φ, a scalar estimate is output. ( This value represents the value derived from state s based on the current policy. t The mathematical expectation of the cumulative future rewards that can be obtained from starting a state is used to judge the long-term quality of the state.

[0046] Furthermore, the input to step 22) is the current environment state s. t The output is the estimated value. ( During the optimization process, the value network minimizes the mean squared error (MSE) loss function. This allows for the adjustment of its parameters, thereby minimizing the error between the estimated value function and the actual discounted return (target return value). Defined as:

[0047]

[0048] in, It is the mathematical expectation; ( ) is a state The value estimate; γ is the discount factor; It is the algorithm's reward at time t+1; T is the final time;

[0049] Furthermore, step 23) avoids excessively rapid policy updates by limiting the change in the ratio of the new policy's probability to the old policy's probability. The objective function can be expressed as:

[0050]

[0051]

[0052]

[0053] Where θ is the parameter of the policy network. θ is the probability of the new policy; θ′ is the parameter of the old policy network. It is the probability of the old strategy; δ is the dominance function of the action, used as the weights to update the action probability; it is calculated through generalized dominance estimation. t and δ t+1 These are the temporal difference learning errors at times t and t+1, respectively; the function Limiting the ratio The change is kept within the range of [1−ε,1+ε] to prevent excessive policy changes; ε is a hyperparameter used to control the update magnitude.

[0054] Further, in step 3), the parameters of the policy network and the value network are updated based on the objective function calculated in step 2). The optimization objective of the value network is to minimize the error between the estimated value and the actual return. In each iteration, a mini-batch of data is sampled from the experience pool, and the value loss is calculated. Subsequently, gradient descent is used to simultaneously update the policy network parameters θ and the value network, repeating steps 1) to 3), periodically replacing the old policy network with the updated one, and continuing to interact with the environment to collect new data. This offline training process continues until the policy performance converges or the preset maximum number of iterations is reached.

[0055] Further, in step 4), specifically after the offline training in step 3) is completed, the parameters of the converged policy network are solidified and deployed into the actual energy management controller. During flight, the controller collects the current state s in real time at a fixed frequency (e.g., 10Hz). t . s t The input is fed into the loaded policy network. The network performs forward computation, outputs the current optimal action (i.e., the optimal battery power increment), and sends the optimal control commands to the battery management system and fuel cell unit for execution. In this way, the trained intelligent policy that can collaboratively optimize fluctuations and consumption can be executed online in real time, enabling the system to maintain a high-efficiency and smooth operating state in complex and dynamic flight environments.

[0056] The beneficial effects of this invention are:

[0057] (1) Unlike traditional local optimization approaches based on segmented operating conditions or threshold triggering, this invention constructs an agent through a near-end strategy optimization algorithm that can make decisions in a continuous action space based on the real-time state and future expectations of the system. The reward function trained on this agent incorporates both total energy consumption and power fluctuation penalty terms, thereby enabling it to autonomously learn and execute a globally optimal strategy that balances overall energy economy and output stability throughout the complex and dynamic land-air conversion process.

[0058] (2) By directly optimizing the long-term cumulative reward, this invention can effectively smooth the power output curve of the power source while ensuring power demand, and reduce the power impact caused by mode switching or load changes. This not only reduces the fatigue loss of key components and extends the system life, but also improves the average energy conversion efficiency under all operating conditions by avoiding frequent start-stop or drastic fluctuations of equipment such as fuel cells in the inefficient range.

[0059] (3) This invention proposes a reward function architecture that combines an exponential function and a piecewise defined function. This design maintains the differentiability and continuity required for policy gradient optimization, incorporates domain-specific nonlinear responses applicable to practical energy dispatch scenarios, and maintains computational tractability when dealing with high-dimensional state-action spaces. A unified reward function mitigates battery power fluctuations, balancing the collaborative contributions of fuel cells and batteries under dynamic wind disturbances. This approach effectively protects battery health and extends battery life. Attached Figure Description

[0060] Figure 1 This is a structural diagram of the system of the present invention.

[0061] Figure 2 This is a flowchart of the method of the present invention. Detailed Implementation

[0062] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to embodiments and accompanying drawings. The content mentioned in the embodiments is not intended to limit the present invention.

[0063] Reference Figure 1 As shown, a fuel cell hybrid propulsion system includes: a fuel cell power generation unit, a lithium-ion battery unit, an electric drive unit, a sensing and data acquisition unit, and an energy management controller. The sensing and data acquisition unit is used to collect real-time status data required by the energy management controller. The energy management controller is configured to execute the aforementioned power fluctuation-consumption co-optimization method based on near-end strategy optimization to achieve optimal real-time allocation of output power between the fuel cell power generation unit and the lithium-ion battery unit.

[0064] The fuel cell power generation unit includes: a fuel cell stack and its associated hydrogen supply system, air supply system, thermal management system and DC-DC converter, used to provide the main driving range power.

[0065] The lithium-ion battery unit includes: a high-power-density lithium-ion battery pack and its management system, used to provide peak power, recover braking energy and smooth load fluctuations.

[0066] The electric drive unit includes a drive motor, a motor controller, and a propeller, used to convert electrical energy into propulsion power.

[0067] The sensing and data acquisition unit includes: a sensor computing module, a battery management system, current and voltage sensors, and a flight attitude sensor.

[0068] The sensor computing module is used to monitor the overall power requirements of the machine.

[0069] Battery management system is used to monitor the state of charge of lithium-ion battery cells;

[0070] Current and voltage sensors are used to monitor the real-time output power of fuel cell power generation units and lithium-ion battery units;

[0071] Flight attitude sensors are used to sense flight attitude, airspeed, and ambient wind speed, providing input for simulation environment models.

[0072] The energy management controller includes: an environment simulation and interaction module, a near-end strategy optimization algorithm module, and a power allocation module.

[0073] The environment simulation and interaction module is responsible for building a digital twin simulation environment within the controller to simulate the operation of a hybrid propulsion system in a dynamic wind field.

[0074] The near-end policy optimization algorithm module is responsible for implementing the near-end policy optimization algorithm and driving the controller to perform self-learning and optimization.

[0075] The power allocation module is used to apply the trained model to real-time control.

[0076] Reference Figure 2 As shown, a power fluctuation-consumption co-optimization method for a fuel cell hybrid propulsion system includes the following steps:

[0077] 1) Construct a simulation environment for the hybrid propulsion system in a dynamic wind field. This environment module simulates the dynamic behavior of the fuel cell hybrid propulsion system. Real-time system status data is collected, including overall power demand, battery state of charge, and current power of the fuel cell and lithium-ion battery. The action commands of the intelligent agent are input into the environment. Based on this, the instantaneous reward reflecting the energy consumption and fluctuation level is calculated, and all interaction information is stored in the experience pool to provide data support for subsequent optimization learning.

[0078] Step 1) specifically includes:

[0079] The environmental simulation and interaction module simulates the dynamic behavior of a hybrid propulsion system in a dynamic wind field. Based on the current state s... t and current action a t Calculate the instant reward R t and the next state s t+1 The information is then stored in an experience pool. Optionally, in this invention, the state variable is set as the power demand P. dem Battery capacity (SOC), fuel cell power (P) fc and lithium-ion battery power P b a t Set to ΔP b (That is, the difference between the lithium-ion battery power and the power of the previous second), which reduces the frequent fluctuations in lithium-ion battery power, thereby improving the exploration efficiency in the action space. R provided by the environment t It can be calculated in the following ways:

[0080]

[0081] Among them, R ecms (t) is the reward function based on the strategy of minimizing equivalent costs; R sm (t) represents the smooth control reward; R c (t) is the penalty for violating the constraint, which is a large negative constant when the constraint is violated.

[0082] R ecms (t) and R sm (t) uses an exponential function and a piecewise defined function, respectively.

[0083] R ecms (t) quantifies the nonlinear relationship between energy utilization and system efficiency, where the reward decays exponentially with increasing hydrogen and electricity consumption to incentivize minimizing operating costs. Optionally, in this invention, R... ecms The design of (t) is as follows:

[0084]

[0085] Where C1 and C2 represent the prices of hydrogen and electricity, respectively; f fc (t) and f b (t) represents the hydrogen and electricity consumption at time t, respectively; λ is the equivalent consumption factor, where a higher λ and a lower λ indicate that the hybrid power system is more inclined to use hydrogen and electricity, respectively.

[0086] R sm (t) is used to implement a threshold-driven segmented architecture and operates through a dual response mechanism: (1) when power fluctuations remain below a specified change threshold, positive reinforcement is provided through exponential decay rewards, thereby promoting operational stability; (2) for fluctuations exceeding the threshold, exponential penalties are applied to dynamically suppress high-amplitude changes. The segmented definition function R sm (t) Battery power fluctuations are regulated by imposing a penalty only when the deviation exceeds a predefined threshold, tolerating small changes while effectively limiting significant disturbances. Optionally, in this invention, R... sm The design of (t) is as follows:

[0087]

[0088] Where K1 is the smoothing control weight coefficient; C3 is the fluctuation range of the lithium-ion battery; the reward function gives positive or negative rewards for ΔPb that is greater than or less than C3.

[0089] R c (t) is calculated by the following formula:

[0090]

[0091] Where K2 is the penalty value for exceeding the performance constraints of the hybrid propulsion system;

[0092] The performance constraints of the hybrid propulsion system are as follows:

[0093]

[0094] Among them, P fc,max This is the maximum power of a hydrogen fuel cell; P b,min and P b,max These are the minimum and maximum power of a lithium-ion battery; SOC. min and SOCmax These represent the minimum and maximum states of charge, respectively.

[0095] The smoothing control weight coefficient (K1) and the penalty value for exceeding the constraint (K2) are key hyperparameters in the reward function, which are obtained through a grid search method.

[0096] 2) Based on the interaction data collected in step 1), construct and train the policy network and the value network. The policy network generates a probability distribution of battery power adjustment actions according to the current state, while the value network evaluates the long-term value of the state. The merits of a single step are quantified by calculating the advantage function, and the objective function is optimized using the pruning mechanism unique to proximal policy optimization. This ensures the stability of the training process and enables the algorithm to effectively optimize both power fluctuation and energy consumption objectives.

[0097] Step 2) specifically includes:

[0098] 21) Estimate action probabilities and state values;

[0099] 22) Optimize the value network;

[0100] 23) Pruning and optimizing the objective function.

[0101] Step 21) specifically estimates the action probability through an action policy network. This network uses the current system state s. t As input, after calculating its parameter θ, the output is an action probability density function. For ΔP b For these continuous actions, the output is specifically represented by the parameters (mean and variance) of a Gaussian distribution, thus defining the probability distribution of the available actions. Simultaneously, the value of a state is evaluated through a value network. This network operates with the same state s. t As input, and through the calculation of its parameter φ, a scalar estimate is output. ( This value represents the value derived from state s based on the current policy. t The mathematical expectation of the cumulative future rewards that can be obtained from starting a state is used to judge the long-term quality of the state.

[0102] The input to step 22) is the current environment state st, and the output is the estimated value. ( During the optimization process, the value network minimizes the mean squared error (MSE) loss function. This allows for the adjustment of its parameters, thereby minimizing the error between the estimated value function and the actual discounted return (target return value). Defined as:

[0103]

[0104] in, It is the mathematical expectation; ( ) is a state The value estimate; γ is the discount factor; It is the algorithm's reward at time t+1; T is the final time.

[0105] Step 23) avoids excessively rapid policy updates by limiting the change in the ratio of the new policy's probability to the old policy's probability. The objective function can be expressed as:

[0106]

[0107]

[0108]

[0109] Where θ is the parameter of the policy network. θ is the probability of the new policy; θ′ is the parameter of the old policy network. It is the probability of the old strategy; δ is the dominance function of the action, used as the weights to update the action probability; it is calculated through generalized dominance estimation. t and δ t+1 These are the temporal difference learning errors at times t and t+1, respectively; the function Limiting the ratio The change is kept within the range of [1−ε,1+ε] to prevent excessive policy changes; ε is a hyperparameter used to control the update magnitude.

[0110] 3) Update the parameters of the policy network and value network based on the objective function calculated in step 2). The optimization objective of the value network is to minimize the error between the estimated value and the actual return. In each iteration, a mini-batch of data is sampled from the experience pool, and the value loss is calculated. Subsequently, gradient descent is used to simultaneously update the policy network parameters θ and the value network, repeating steps 1) to 3), periodically replacing the old policy network with the updated one, and continuing to interact with the environment to collect new data. This offline training process continues until the policy performance converges or the preset maximum number of iterations is reached.

[0111] 4) After the offline training in step 3) is completed and the policy network reaches convergence, the optimal network parameters obtained from the training are solidified. Subsequently, the complete policy network model containing these solidified parameters is deployed to the hardware storage unit of the energy management controller.

[0112] Before executing the mission, the controller completes initialization, loading the policy network model into its runtime memory to put it in a ready state. Upon entering the flight phase, the controller periodically collects the current state variables of the entire system at a fixed sampling frequency (e.g., 10Hz), forming a real-time state vector s. t .

[0113] The controller will s t The input is fed into the loaded policy network. Based on its fixed parameters and internal structure, the network performs a forward propagation inference calculation and directly outputs the state corresponding to the current state s. t The optimal action command is the optimal power increment setpoint for the battery system within the current control cycle.

[0114] Finally, the energy management controller converts this power increment setpoint into specific control signals and sends them to the battery management system and associated power actuators, thereby achieving real-time closed-loop optimization management of the hybrid power system. The trained intelligent strategies that can collaboratively optimize fluctuations and consumption are executed online in real time, enabling the system to maintain a highly efficient and smooth operating state in complex and dynamic flight environments.

[0115] This invention is based on near-end strategy optimization. It constructs an agent that takes system state and operating condition information as input and power source allocation commands as output. A reward function integrating total energy consumption penalty and power fluctuation penalty is designed to drive the agent to learn the optimal energy management strategy in a continuous action space. This method can suppress frequent and drastic fluctuations in lithium-ion battery power, extend battery life, ensure system safety, and minimize overall hydrogen fuel consumption while meeting dynamic power demands and safety constraints. This achieves a synergistic optimization of long-term operational economy and system output stability.

[0116] This invention has many specific applications. The above description is only a preferred embodiment of this invention. It should be noted that for those skilled in the art, several improvements can be made without departing from the principle of this invention, and these improvements should also be considered within the scope of protection of this invention.

Claims

1. A method for co-optimizing power fluctuation and energy consumption in a fuel cell hybrid propulsion system, characterized in that, The steps are as follows: Step 1): Collect system status data in real time, input the action commands of the intelligent agent and the changes in battery power into the environment, build a simulation environment in the dynamic wind field, and simulate the dynamic behavior of the hybrid propulsion system through the environment simulation and interaction module; Design a reward function that reflects energy consumption and fluctuation levels, and store all interaction information in an experience pool to provide data support for subsequent optimization learning; Step 2): Based on the interaction data collected in Step 1), construct and train the policy network and the value network; the policy network generates the probability distribution of battery power adjustment actions according to the current state, while the value network evaluates the long-term value of the state; the merits of a single step action are quantified by calculating the advantage function, and the objective function is optimized by the pruning mechanism unique to proximal policy optimization, ensuring the stability of the training process and enabling the algorithm to effectively optimize the two objectives of power fluctuation and energy consumption. Step 3): Based on the pruning optimization objective function and the loss function of the value network obtained in Step 2), the gradient descent algorithm is used to update the parameters of the policy network and the value network simultaneously; by repeating Step 1) to Step 3), a closed-loop iterative process of "data collection-network optimization-parameter update" is formed until the policy performance converges, and finally an optimization decision model that intelligently balances power fluctuations and energy consumption is obtained. Step 4): Solidify the parameters of the policy network model trained and converged in Step 3) and deploy it to the actual energy management controller; Step 1) specifically includes: System state data includes, overall power demand, battery state of charge, current power of fuel cell and lithium ion battery; according to current state s t and current action a t , calculate immediate reward R t and next state s t+1 , and store information in experience pool; state variables are set as power demand P dem , battery capacity SOC, fuel cell power P fc and lithium ion battery power P b ; a t is set as ΔP b , ΔP b is the difference between lithium ion battery power and previous second power; R t provided by the environment is calculated in the following way: ; Among them, R ecms (t) is the reward function based on the strategy of minimizing equivalent costs; R sm (t) represents the smooth control reward; R c (t) is the penalty for violating the constraint, which gives a large negative constant when the constraint is violated; R ecms (t) and R sm (t) Using exponential and piecewise defined functions respectively, for R ecms The design of (t) is as follows: ; Where C1 and C2 represent the prices of hydrogen and electricity, respectively; f fc (t) and f b (t) represents the hydrogen and electricity consumption at time t, respectively; λ is the equivalent consumption factor, where a higher λ and a lower λ indicate that the hybrid power system is more inclined to use hydrogen and electricity, respectively. For R sm The design of (t) is as follows: ; Where K1 is the smoothing control weight coefficient; C3 is the fluctuation range of the lithium-ion battery; the reward function gives positive or negative rewards for ΔPb that is greater than or less than C3 respectively; R c (t) is calculated by the following formula: ; Where K2 is the penalty value for exceeding the performance constraints of the hybrid propulsion system; The performance constraints of the hybrid propulsion system are as follows: ; Among them, P fc,max This is the maximum power of a hydrogen fuel cell; P b,min and P b,max These are the minimum and maximum power of a lithium-ion battery; SOC. min and SOC max These are the minimum and maximum states of charge, respectively. In step 4), after the offline training in step 3) is completed, the parameters of the converged policy network are solidified and deployed to the energy management controller, as follows: During flight, the controller collects the current state s in real time at a fixed frequency. t , will s t The input is fed into the loaded policy network, which performs forward calculations and outputs the current optimal action, namely the optimal battery power increment. The optimal control commands are then sent to the battery management system and fuel cell unit for execution. The trained intelligent policy that can collaboratively optimize fluctuations and consumption can be executed online in real time, enabling the system to maintain a high-efficiency and smooth operating state in complex and dynamic flight environments.

2. The power fluctuation-consumption co-optimization method for a fuel cell hybrid propulsion system according to claim 1, characterized in that, In step 2), based on the interaction data collected in step 1), the policy network and value network are constructed and trained, specifically including: 21) Estimate action probabilities and state values; 22) Optimize the value network; 23) Pruning and optimizing the objective function; Step 21) specifically estimates the action probability through an action policy network; this network uses the current system state s t As input, after calculating its parameter θ, the output is an action probability density function. For ΔP b These continuous actions are represented by parameters of a Gaussian distribution, thus defining the probability distribution of the available actions; the value of a state is evaluated through a value network, which evaluates the state s under the same conditions. t As input, and through the calculation of its parameter φ, a scalar estimate is output. ( This value represents the value derived from state s based on the current policy. t The mathematical expectation of the cumulative future rewards that can be obtained from starting is used to judge the long-term quality of the state; The input for step 22) is the current environment state s. t The output is the estimated value. ( During the optimization process, the value network minimizes the mean squared error loss function. To adjust its parameters and minimize the error between the estimated value function and the actual discounted return, Defined as: ; in, It is the mathematical expectation; ( ) is a state The value estimate; γ is the discount factor; It is the algorithm's reward at time t+1; T is the final time; Step 23) avoids excessively rapid policy updates by limiting the change in the ratio of the probabilities of the new policy to the old policy. The objective function is expressed as: ; ; ; Where θ is the parameter of the policy network. θ is the probability of the new policy; θ′ is the parameter of the old policy network. It is the probability of the old strategy; It is the dominance function of the action, used as the weights to update the action probability, and is calculated through generalized dominance estimation; δ t and δ t+1 These are the temporal difference learning errors at times t and t+1, respectively; the function Limiting the ratio The change is kept within the range of [1−ε,1+ε] to prevent excessive policy changes; ε is a hyperparameter used to control the update magnitude.

3. The power fluctuation-consumption co-optimization method for a fuel cell hybrid propulsion system according to claim 2, characterized in that, In step 3), the parameters of the policy network and the value network are updated based on the objective function calculated in step 2); the optimization objective of the value network is to minimize the error between the estimated value and the actual return. In each iteration, a small batch of data is sampled from the experience pool, and the value loss is calculated. Use the gradient descent algorithm to update the policy network parameters θ and the value network simultaneously, repeating steps 1) to 3), periodically replacing the old policy network with the updated policy network, and continue to interact with the environment to collect new data; this offline training process continues until the policy performance converges or the preset maximum number of iterations is reached.

4. A fuel cell hybrid propulsion system for the power fluctuation-consumption co-optimization method of claim 1, characterized in that, include: Fuel cell power generation unit, lithium-ion battery unit, electric drive unit, sensing and data acquisition unit, and energy management controller; The sensing and data acquisition unit is used to acquire the status data required by the energy management controller in real time. The energy management controller is used to execute a power fluctuation-consumption co-optimization method based on near-end strategy optimization to achieve optimal real-time allocation of the output power of the fuel cell power generation unit and the lithium-ion battery unit.

5. The fuel cell hybrid propulsion system according to claim 4, characterized in that, The fuel cell power generation unit includes: a fuel cell stack and its supporting hydrogen supply system, air supply system, thermal management system and DC-DC converter, used to provide driving range power; The lithium-ion battery unit includes: a high-power-density lithium-ion battery pack and its management system, used to provide peak power, recover braking energy and smooth load fluctuations; The electric drive unit includes a drive motor, a motor controller, and a propeller, used to convert electrical energy into propulsion power; The sensing and data acquisition unit includes: a sensor computing module, a battery management system, current and voltage sensors, and a flight attitude sensor; the sensor computing module is used to monitor the overall power demand; the battery management system is used to monitor the state of charge of the lithium-ion battery cells; the current and voltage sensors are used to monitor the real-time output power of the fuel cell power generation unit and the lithium-ion battery cells; and the flight attitude sensor is used to sense flight attitude, airspeed, and ambient wind speed, providing input for the simulation environment model. The energy management controller includes: an environment simulation and interaction module, a near-end strategy optimization algorithm module, and a power allocation module; the environment simulation and interaction module constructs a digital twin simulation environment running in a dynamic wind field within the controller; the near-end strategy optimization algorithm module implements the near-end strategy optimization algorithm, driving the controller to perform self-learning and optimization; the power allocation module is used to apply the trained model to real-time control.

Citation Information

Patent Citations

  • A fuel cell hybrid propulsion system and control method for an underwater vehicle

    CN111204430B

  • Hybrid propulsion system air-ground conversion power matching optimization method

    CN119226680A

  • Energy self-adaptive optimization system and method for hydrogen fuel cell hybrid power unmanned aerial vehicle

    CN119167791A

  • Self-adaptive working condition sensing fuel cell hybrid tramcar hierarchical management method

    CN120422725A