Fuel cell hybrid electric vehicle energy management method and device based on multi-agent reinforcement learning, storage medium and equipment

The energy management method using multi-agent reinforcement learning solves the problem of low power source collaboration efficiency in fuel cell hybrid electric vehicles, achieving efficient power distribution and vehicle energy consumption optimization, and extending the life of the power source.

CN119872282BActive Publication Date: 2025-11-11SHENZHEN INST OF ADVANCED TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411893623.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-11-11
Estimated Expiration
2044-12-20

AI Technical Summary

Technical Problem

Existing energy management strategies for fuel cell hybrid electric vehicles struggle to fully leverage the unique advantages of each power source and achieve efficient collaboration, especially when faced with multiple power sources of different structures and characteristics, making efficient power distribution difficult.

Method used

An energy management method based on multi-agent reinforcement learning is adopted. Through a learning framework of centralized training and distributed control, output power control strategy models for fuel cells and batteries are designed respectively. The simulation model is used to simulate driving conditions and train the energy management strategy model to optimize power allocation to achieve the coordinated operation of multiple power sources.

Benefits of technology

It achieves efficient collaborative power distribution among multiple power sources, reduces reliance on expert experience, fully leverages the advantages of each power source, reduces overall vehicle hydrogen consumption, maintains a stable state of charge in the hybrid energy storage system, and slows down the lifespan degradation of the core power source.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119872282B_ABST
    Figure CN119872282B_ABST
Patent Text Reader

Abstract

This application provides a method, device, storage medium, and equipment for energy management of fuel cell hybrid electric vehicles based on multi-agent reinforcement learning. The method includes constructing a simulation model of the fuel cell hybrid electric vehicle to simulate preset driving conditions and obtain vehicle power demand data; based on the vehicle power demand data, determining the power allocation results of each power source using an energy management strategy model based on multi-agent reinforcement learning; based on the power allocation results, determining the vehicle state information before and after power allocation using the simulation model, and obtaining the reward value of the energy management strategy model; and training the energy management strategy model using the power allocation results, vehicle state information, and reward value. The energy management strategy model includes a fuel cell output power control strategy model and a battery output power control strategy model based on reinforcement learning. This method can achieve efficient collaborative power allocation among multiple power sources, fully leveraging the unique advantages of multiple power sources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of new energy vehicle technology, specifically, it relates to a method, device, storage medium, and equipment for energy management of fuel cell hybrid electric vehicles based on multi-agent reinforcement learning. Background Technology

[0002] In recent years, fuel cell hybrid electric vehicles (FCEVs) have been considered one of the best solutions for the electrification transformation of automobiles due to their integrated multi-power source system and the environmental advantage of each power source achieving zero emissions. During the operation of FCEVs, fuel cells should avoid operating under certain degradation conditions, such as frequent starts and stops, rapid load changes, and idling, as these conditions can adversely affect fuel cell durability. Batteries and supercapacitors should avoid high discharge rates and be protected against overcharging and over-discharging, which is crucial for protecting the health of the power source system and maintaining the stability of the vehicle's power output. Therefore, an energy management strategy that efficiently optimizes the energy distribution among the various power sources is particularly important. This not only ensures smooth vehicle operation and significantly improves its overall performance but also effectively reduces the vehicle's equivalent hydrogen consumption while delaying the degradation of fuel cells and batteries and maintaining the stable state of charge of batteries and supercapacitors.

[0003] To address the energy management problem of fuel cell hybrid electric vehicles (FCEVs), researchers have proposed several energy management strategies, which can be broadly categorized into three types based on the different strategies employed: rule-based, optimization-based, and learning-based strategies. Rule-based energy management strategies primarily rely on expert experience and the overall vehicle powertrain architecture design to formulate energy allocation rules among multiple power sources. This strategy is easy to implement and meets real-time energy allocation requirements, but it requires expert experience to design the power allocation rules, and its control performance is poor under complex operating conditions. Optimization-based energy management strategies rely on designing an energy management cost function, followed by using advanced optimization algorithms to find the output power sequence of each power source that minimizes this cost function. However, the computational complexity is high, making it difficult to meet the real-time requirements of power allocation. Learning-based methods utilize artificial intelligence to build energy management strategy models. These models determine the output power of each power source based on real-time vehicle conditions and power demand. This strategy not only meets real-time control requirements but also achieves excellent control performance, demonstrating the advantages of intelligent strategies. In recent years, thanks to the rapid development of artificial intelligence technology, learning-based methods, especially reinforcement learning methods, have become a research hotspot in the field of energy management, particularly for the core issue of hybrid electric vehicle energy management.

[0004] Currently, for energy management of multi-power-source fuel cell hybrid electric vehicles, deep reinforcement learning-based methods mainly include: hierarchical energy management strategies combining fuzzy logic and deep reinforcement learning, and centralized energy management strategies based on deep reinforcement learning. Specifically, the hierarchical energy management strategy uses an adaptive fuzzy controller at the upper layer to determine the output power of the supercapacitor, while the lower layer uses deep reinforcement learning to determine the output power of the fuel cell and battery, achieving power allocation across multiple power sources. The centralized energy management strategy, on the other hand, uses a deep reinforcement learning algorithm to implement a central controller that centrally allocates the output power of multiple power sources based on the current vehicle state and required power. Although existing energy management strategies have shown some effectiveness in control performance and real-time performance, they inevitably have some limitations or shortcomings that need improvement. For example, the hierarchical energy management strategy still relies heavily on experience to design adaptive fuzzy control rules, limiting the strategy's adaptability. The centralized energy management strategy, when faced with multiple power sources with different structures and characteristics, struggles to fully explore and utilize the advantages of each power source to achieve more efficient collaborative power allocation. Summary of the Invention

[0005] The technical problem addressed in this application is: how to provide an energy management method for fuel cell hybrid electric vehicles based on multi-agent reinforcement learning that can fully leverage the unique advantages of each power source and coordinate the work of each power source.

[0006] This application provides an energy management method for fuel cell hybrid electric vehicles based on multi-agent reinforcement learning, the energy management method for fuel cell hybrid electric vehicles including:

[0007] Data acquisition phase: Construct a simulation model of the fuel cell hybrid electric vehicle, use the simulation model to simulate preset driving conditions, and obtain vehicle power demand data for a predetermined duration;

[0008] Model training phase: Based on vehicle power demand data, the power allocation results of each power source are determined using a pre-built energy management strategy model based on multi-agent reinforcement learning; based on the power allocation results, the vehicle state information before and after power allocation is determined using the simulation model, and the reward value of the energy management strategy model is obtained; the energy management strategy model is trained using the power allocation results, the vehicle state information, and the reward value to complete one round of training; the energy management strategy model is trained repeatedly for multiple rounds until the training termination condition is met;

[0009] Among them, the energy management strategy model includes a fuel cell output power control strategy model based on reinforcement learning and a battery output power control strategy model based on reinforcement learning.

[0010] Optionally, the method for obtaining vehicle power demand data for a predetermined duration by simulating preset driving conditions using the simulation model includes:

[0011] The simulation model is run according to the preset driving conditions to obtain the vehicle driving data at the current moment. The vehicle driving data includes the total mass of the vehicle, gravitational acceleration, rolling resistance coefficient, driving slope, air resistance coefficient, frontal area, vehicle speed, mass coefficient and vehicle acceleration.

[0012] The vehicle's power demand at the current moment is calculated based on the vehicle's driving data.

[0013] Optionally, methods for determining the power allocation results of each power source based on vehicle power demand data and utilizing a pre-built energy management strategy model based on multi-agent reinforcement learning include:

[0014] Based on the vehicle's current power demand and the vehicle's status information from the previous moment, the energy management strategy model of the current training round is used to simultaneously determine the output power of the fuel cell and the battery at the current moment.

[0015] The output power of the supercapacitor at the current moment is determined based on the vehicle's current power demand, the fuel cell's output power, and the battery's output power.

[0016] Optionally, based on the power allocation result, the method for determining the vehicle state information before and after power allocation using the simulation model includes:

[0017] Based on the power allocation results at the current moment, the simulation model is used to determine the vehicle's equivalent hydrogen consumption, the change in the battery's state of charge, the change in the supercapacitor's state of charge, and the degradation of the core power source before and after the power allocation.

[0018] Optionally, methods for simultaneously determining the output power of the fuel cell and the battery at the current moment using the energy management strategy model of the current training round include:

[0019] The fuel cell output power control strategy model of the current training round is used to determine the output power of the fuel cell at the current moment based on the local observation state information of the fuel cell.

[0020] The battery output power control strategy model of the current training round determines the battery output power at the current moment based on the local observation state information of the battery.

[0021] Optionally, the local observation state information of the fuel cell includes the vehicle's power demand at the current moment, the fuel cell's output power at the previous moment, the supercapacitor's state of charge, and the difference between the supercapacitor's state of charge and the target state of charge; the local observation state information of the battery includes the vehicle's power demand at the current moment, the battery's state of charge at the previous moment, the difference between the battery's state of charge and the target state of charge, the supercapacitor's state of charge, and the difference between the supercapacitor's state of charge and the target state of charge.

[0022] Optionally, the method for obtaining the reward value of the energy management strategy model includes:

[0023] Based on the power allocation results and vehicle status information, the first part of the reward value is obtained using the reward function of the fuel cell output power control strategy model;

[0024] Based on the power allocation results and vehicle status information, the second part of the reward value is obtained using the reward function of the battery output power control strategy model.

[0025] This application also provides an energy management device for a fuel cell hybrid electric vehicle based on multi-agent reinforcement learning, the energy management device for the fuel cell hybrid electric vehicle comprising:

[0026] The data acquisition module is used to construct a simulation model of a fuel cell hybrid electric vehicle, and use the simulation model to simulate preset driving conditions to obtain vehicle power demand data for a predetermined duration.

[0027] The model training module is used to determine the power allocation results of each power source based on the vehicle's power demand data and a pre-built energy management strategy model based on multi-agent reinforcement learning; based on the power allocation results, the module uses the simulation model to determine the vehicle state information before and after power allocation and obtains the reward value of the energy management strategy model; the module trains the energy management strategy model using the power allocation results, the vehicle state information, and the reward value to complete one round of training; the module repeats the training of the energy management strategy model for multiple rounds until the training termination condition is met.

[0028] Among them, the energy management strategy model includes a fuel cell output power control strategy model based on reinforcement learning and a battery output power control strategy model based on reinforcement learning.

[0029] This application also provides a computer-readable storage medium storing a fuel cell hybrid electric vehicle energy management program based on multi-agent reinforcement learning. When the fuel cell hybrid electric vehicle energy management program is executed by a processor, it implements the aforementioned fuel cell hybrid electric vehicle energy management method based on multi-agent reinforcement learning.

[0030] This application also provides a computer device, the computer device including a computer-readable storage medium, a processor, and a fuel cell hybrid electric vehicle energy management program based on multi-agent reinforcement learning stored in the computer-readable storage medium, wherein the fuel cell hybrid electric vehicle energy management program, when executed by the processor, implements the aforementioned fuel cell hybrid electric vehicle energy management method based on multi-agent reinforcement learning.

[0031] This application provides a method, device, storage medium, and equipment for energy management of fuel cell hybrid electric vehicles based on multi-agent reinforcement learning, which has the following technical effects:

[0032] Employing a learning framework that combines centralized training and distributed control effectively promotes collaboration among multiple power sources and mitigates the instability of multi-agent interaction environments. Compared to hierarchical and centralized energy management strategies, the energy management strategy based on multi-agent deep reinforcement learning enables efficient collaborative power allocation among multiple power sources, effectively reducing reliance on expert experience and fully leveraging the unique advantages of multiple power sources. It also addresses multiple energy management and control objectives, including reducing overall vehicle hydrogen consumption, maintaining a stable state of charge in the hybrid energy storage system, and delaying the degradation of the core power source's lifespan. Attached Figure Description

[0033] Figure 1 The flowchart shows the main steps of a fuel cell hybrid electric vehicle energy management method based on multi-agent reinforcement learning according to one or more embodiments.

[0034] Figure 2 An architectural diagram of a simulation model of a fuel cell hybrid electric vehicle according to one or more embodiments;

[0035] Figure 3 This is a schematic block diagram of a fuel cell hybrid electric vehicle energy management device based on multi-agent reinforcement learning according to one or more embodiments;

[0036] Figure 4 This is a schematic diagram of a computer device according to one or more embodiments. Detailed Implementation

[0037] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0038] Before describing the various embodiments of this application in detail, the technical concept of this application is first briefly described: Current energy management methods for multi-power-source fuel cell hybrid electric vehicles based on deep reinforcement learning suffer from problems such as difficulty in fully exploring and utilizing the advantages of each power source and difficulty in achieving efficient cooperation among the power sources. To address this, the energy management method for fuel cell hybrid electric vehicles based on multi-agent reinforcement learning provided in this application adopts a learning framework of centralized training and distributed control. It independently designs reinforcement learning-based output power control strategy models for both the fuel cell and the battery, which can effectively promote the cooperation capabilities among multiple power sources and fully leverage the advantages of each power source. The specific principles of the energy management method and control method for fuel cell hybrid electric vehicles based on multi-agent reinforcement learning of this application are described below with reference to more embodiments.

[0039] Specifically, such as Figure 1 As shown, the energy management method for fuel cell hybrid electric vehicles based on multi-agent reinforcement learning in this embodiment includes the following steps:

[0040] Step S10, Data Acquisition Stage: Construct a simulation model of the fuel cell hybrid electric vehicle, use the simulation model to simulate preset driving conditions, and obtain vehicle power demand data for a predetermined duration;

[0041] Step S20, Model Training Phase: Based on the vehicle power demand data, the power allocation results of each power source are determined using a pre-built energy management strategy model based on multi-agent reinforcement learning; based on the power allocation results, the vehicle state information before and after power allocation is determined using a simulation model, and the reward value of the energy management strategy model is obtained; the energy management strategy model is trained using the power allocation results, vehicle state information, and reward value to complete one round of training; the energy management strategy model is repeatedly trained for multiple rounds until the training termination condition is met; wherein, the energy management strategy model includes a fuel cell output power control strategy model based on reinforcement learning and a battery output power control strategy model based on reinforcement learning.

[0042] In one or more embodiments, the overall architecture of the simulation model of the fuel cell hybrid electric vehicle built during the data acquisition phase is as follows: Figure 2As shown, the vehicle mainly consists of a multi-power source system architecture comprised of fuel cells, batteries, and supercapacitors. The simulation models for each power source are built using the fuel cell stack voltage model, the battery internal resistance equivalent circuit model, and the supercapacitor classical equivalent circuit model, respectively. For example, the preset driving condition is a hybrid driving cycle consisting of multiple international standard driving cycles. This hybrid driving cycle effectively simulates the actual operating conditions of the vehicle, with a sampling interval of one second and a total operating time of approximately one hour. Subsequently, based on vehicle speed data such as the current speed and the next speed in the hybrid driving cycle, the vehicle power demand model is used to calculate the vehicle power demand at each moment in the simulation model of the current fuel cell hybrid electric vehicle.

[0043] For example, vehicle driving data includes total vehicle mass, gravitational acceleration, rolling resistance coefficient, driving slope, air resistance coefficient, frontal area, vehicle speed, mass coefficient, and vehicle acceleration. The vehicle power demand model is used to calculate the vehicle power demand of the fuel cell hybrid electric vehicle simulation model under specified operating conditions, representing the power value required to drive the vehicle at the current vehicle speed and acceleration. For example, this embodiment only considers the straight-line driving condition and does not take into account complex factors such as vehicle steering. Therefore, according to the vehicle longitudinal dynamics model, the vehicle power demand p... req The calculation formula is as follows:

[0044]

[0045] Among them, M,g,f,a,ρ,C d A, v, δ, a represent the total mass of the vehicle, gravitational acceleration, rolling resistance coefficient, driving gradient, air resistance coefficient, frontal area, vehicle speed, mass coefficient, and vehicle acceleration, respectively.

[0046] For example, the simulation model of a fuel cell hybrid electric vehicle also includes a vehicle power balance model. The vehicle power balance model is used to describe the output power of each power source. After the DC-DC converter voltage regulation, drive motor conversion losses, and mechanical losses of the transmission system, the sum of its final usable power must meet the current power demand of the vehicle. Its expression is as follows:

[0047] p req =(p fcs η DC / DC +p bat η Bi-DC / DC +p uc η Bi-DC / DC )η motor η f (2)

[0048] Where, p fcs ,pbat ,p uc η represents the output power of the fuel cell, battery, and supercapacitor system, respectively. DC / DC ,η Bi-DC / DC These represent the efficiencies of the unidirectional and bidirectional DC-DC converters used for voltage conversion, respectively. η motor ,η f This is used to describe the transmission efficiency of motors and transmission systems.

[0049] For example, the fuel cell system model will be established using a stack voltage model, where the operating voltage v of each individual cell is... fc The voltage is determined by the open-circuit voltage E of the battery cell, the activation voltage, the ohmic voltage, and the concentration overvoltage, respectively represented by v. act ,v ohm ,v conc The calculation equation is shown below:

[0050] v fc =Ev act -v ohm -v conc (3)

[0051] fuel cell system voltage v st Then, the single-cell operating voltage and the number of fuel cell cells N are used to determine the fuel cell's operating voltage. cell Confirmed, as shown below:

[0052] v st =N cell ·v fc (4)

[0053] For example, the battery system model is established using an equivalent circuit model of battery internal resistance. This model simplifies the lithium battery into a system consisting of an ideal voltage source and an equivalent charge / discharge internal resistance R. int This system is designed to rapidly assess and track changes in the state of charge (SOC) of a battery under different operating power levels, providing a reference for energy management strategies. When the battery power p is known... bat Based on this, the battery current i bat and the battery state of charge (SOC) at time t b t at It can be expressed by the following formula:

[0054]

[0055] Among them, v ocv ,η bat Q bat These represent the battery open-circuit voltage, the battery's internal charging and discharging efficiency, and the battery's total capacity, respectively.

[0056] For example, the supercapacitor system model is established using a classic equivalent circuit model, which simplifies the supercapacitor to consist of equivalent series and parallel resistances and an ideal capacitor. This model aims to rapidly assess and track the changes in the state of charge of the supercapacitor under different operating power levels. When the supercapacitor power p is known... uc Based on this, the supercharged state at time t It can be expressed by the following formula:

[0057]

[0058] Among them, C uc ,R s ,R p ,v ucmax These represent the total capacitance, equivalent series internal resistance, equivalent parallel resistance, and maximum voltage of the supercapacitor, respectively.

[0059] The overall approach to the model training phase is as follows: Based on the current vehicle power demand and vehicle status information, the output power of both the fuel cell and battery is determined using the current energy management strategy. The output power of the supercapacitor is determined by the combined output power of the fuel cell and battery, along with the vehicle's current power demand, according to the vehicle power balance model. Based on the power allocation of multiple power sources and the vehicle simulation model, the equivalent hydrogen consumption, the change in state of charge of the hybrid energy storage system, and the degradation of the core power source are determined. By evaluating the changes in vehicle status before and after power allocation, and combining this with the reward function designed for the energy management control objective, the current energy management strategy is evaluated. A higher reward value indicates a better control effect of the current energy management strategy. Historical experience data, consisting of vehicle status information before and after power allocation, power allocation of multiple power sources, and corresponding reward values, will be comprehensively collected and stored. Once a certain amount of historical experience data has accumulated in the experience replay pool, a portion of this historical experience will be randomly selected for training and optimizing the current energy management control strategy, continuously improving its control performance and adaptability. Ultimately, in the hybrid driving cycle, the energy management strategy, through a precise multi-power source power allocation strategy, achieved a stable and significant increase in the cumulative discount reward value, marking the completion of the energy management strategy training. The trained energy management strategy model will demonstrate excellent control performance in actual vehicle operating conditions, significantly optimizing overall vehicle energy consumption, maintaining the stable state of charge of the hybrid energy storage system, and effectively delaying the lifespan degradation of the core power source.

[0060] Before going into detail about the various parts of the model training phase, we will introduce the basics of multi-agent reinforcement learning.

[0061] Multi-agent deep reinforcement learning refers to algorithms that utilize deep reinforcement learning methods to solve decision-making and cooperation problems in multi-agent systems. Among them, the multi-agent deep deterministic policy gradient algorithm, a classic algorithm in the field of multi-agent deep reinforcement learning, is an extension of the deep deterministic policy gradient algorithm. This embodiment will use this algorithm as an example for detailed explanation. In this interactive environment, there are multiple independent agents, i.e., independent control policies, and each agent has an independent policy network μ. i (o i ;φ i and value networks The policy network of agent i is composed of its local observed states o. i Give control action a i The value network of agent i is based on the joint observation state s = [o1, o2, ..., o] of all agents. n ] and control action a = [a1, a2, ..., a n The expected value of future cumulative rewards is calculated, and this metric is used to evaluate the quality of the current state and control actions. To enhance the stability of the training process, a target network with the same structure as the policy network, namely the policy-target network, is introduced into both the policy network and the value network. and value target network The following will use agent i as an example to elaborate on the process of collecting historical interaction experience and the training process of the policy network and value network: Agent i's policy network is based on the current local observation state o i t Determine the current control action The joint action of all intelligent agents It will act on the interactive environment and obtain the corresponding reward signal. and the next state These historical experiences of interaction (s t ,a t ,r t ,s t +1 The historical interaction experiences are stored in the experience replay pool for training the policy network and value network. When the historical interaction experiences in the experience replay pool accumulate to a certain level, a portion of the historical experiences (s) can be randomly extracted. j ,a j ,r j ,s j+1 Then, update the parameters of the policy network and value network. Specifically, the policy network of agent i needs to optimize parameter φ. i The goal is to maximize the cumulative future discount return, and the formula for this goal is as follows:

[0062]

[0063] The policy network updates its parameters using policy gradient boosting.

[0064]

[0065] The value network of agent i optimizes the parameter θ. i To improve the accuracy of predicting future cumulative discount returns, updates are primarily performed by minimizing time difference error. The specific process is as follows: Based on state s j+1 All agents use their target policy network to give their control actions. The joint action of all agents is represented as Agent i uses its value target network to calculate future cumulative discounted returns. The time difference objective is expressed as Then, the state s is determined using the value network. j and action a j The following future cumulative discount return The time difference error can then be calculated using the following formula:

[0066]

[0067] The parameter updates of the value network need to minimize the time difference error, and the parameter update method is as follows:

[0068]

[0069] The update methods for the policy objective network and value objective network of agent i are as follows:

[0070]

[0071] In one or more embodiments, the specific method for determining the power allocation results of each power source using a pre-built energy management strategy model based on multi-agent reinforcement learning includes: determining the output power of the fuel cell and the output power of the battery at the current moment using the energy management strategy model of the current training round, based on the vehicle's current power demand and the vehicle's state information at the previous moment; and determining the output power of the supercapacitor at the current moment based on the vehicle's current power demand, the fuel cell's output power, and the battery's output power.

[0072] For example, the fuel cell and battery, as the core power sources of a hybrid electric vehicle, are configured with two independent strategies to control their output power. These strategies can determine the output power of the corresponding power source based on the local observation state information of the fuel cell or battery, so as to achieve efficient and coordinated distribution of vehicle power output.

[0073] Specifically, the method of simultaneously determining the output power of the fuel cell and the battery at the current moment using the energy management strategy model of the current training round includes: determining the output power of the fuel cell at the current moment based on the local observation state information of the fuel cell using the fuel cell output power control strategy model of the current training round; and determining the output power of the battery at the current moment based on the local observation state information of the battery using the battery output power control strategy model of the current training round.

[0074] For example, the fuel cell output power control strategy focuses more on the fuel cell lifespan degradation and the supercapacitor's state of charge during vehicle operation. Therefore, the local observation state information of the fuel cell is set as follows:

[0075]

[0076] Therefore These represent the vehicle power demand at the current time t, the fuel cell output power at the previous time t-1, the supercapacitor state of charge, and the difference between the supercapacitor state of charge and the target state of charge, respectively.

[0077] For example, the battery output power control strategy focuses on battery degradation and the state of charge of the hybrid energy storage system; therefore, the settings for the battery's local observed state information are as follows:

[0078]

[0079] in, They represent the current time. t The vehicle's required power, the battery's state of charge at the previous time t-1, the difference between the battery's state of charge and the target state of charge, the supercapacitor's state of charge, and the difference between the supercapacitor's state of charge and the target state of charge.

[0080] The output power control strategy models for fuel cells and batteries independently determine their output power based on their respective local observation state information, while the output power of the supercapacitor is jointly determined by the actual power demand of the vehicle and the output power already allocated to the fuel cell and battery, as shown below:

[0081]

[0082] in, These represent the minimum and maximum output power of fuel cells, batteries, and supercapacitors, respectively.

[0083] In one or more embodiments, the method for determining vehicle state information before and after power allocation using the simulation model based on the power allocation result includes: determining the vehicle's equivalent hydrogen consumption, battery state of charge (SOC) change, supercapacitor SOC change, and core power source attenuation before and after power allocation using the simulation model based on the current power allocation result. The SOC change of the battery and supercapacitor can be calculated using the aforementioned battery system model and supercapacitor system model, respectively. The vehicle's equivalent hydrogen consumption is obtained by adding the fuel cell hydrogen consumption to the equivalent hydrogen consumption of the battery and supercapacitor. The fuel cell hydrogen consumption can be determined based on the fuel cell hydrogen consumption rate curve (power versus hydrogen consumption curve). The equivalent hydrogen consumption of the battery and supercapacitor is determined by their power, the average hydrogen consumption rate of the fuel cell stack during vehicle operation, the average hydrogen consumption rate and average power of the fuel cell during vehicle operation, and the average charge / discharge efficiency of the battery and supercapacitor. Core power source attenuation mainly includes the attenuation of the fuel cell and battery. The fuel cell utilizes a rapid assessment model for the lifespan degradation of automotive fuel cells, primarily based on the duration / number of idling, high-power load changes, heavy loads, and start-stop cycles during operation. The battery, on the other hand, employs a control-oriented semi-empirical degradation assessment model, with influencing factors including charge / discharge rate, depth of charge / discharge, and cumulative charge throughput.

[0084] In one or more embodiments, the method for obtaining the reward value of the energy management strategy model includes: obtaining a first part of the reward value using the reward function of the fuel cell output power control strategy model based on the power allocation result and vehicle state information; and obtaining a second part of the reward value using the reward function of the battery output power control strategy model based on the power allocation result and vehicle state information.

[0085] The quality of the reward function design will directly affect the optimization effect of the control strategy. Its design should focus on control objectives such as reducing the hydrogen consumption of the whole vehicle, maintaining the stable state of charge of the hybrid energy storage system, and delaying the life decay of the core power source. It should also be designed differently based on the local observation state information of different control strategies.

[0086] For example, the control strategy model for fuel cells focuses more on control objectives such as energy consumption, fuel cell lifespan degradation, and supercapacitor state of charge stability; therefore, its reward function is designed as follows:

[0087]

[0088] in, and These represent the hydrogen consumption of a fuel cell and the equivalent hydrogen consumption of a supercapacitor, respectively. and This is used to punish fuel cell systems for operating conditions that shorten their lifespan, such as large load fluctuations and idling. This is used to guide strategies to maintain the stable state of charge of supercapacitors, where, This represents the target state of charge that a supercapacitor needs to maintain.

[0089] For example, the battery control strategy model focuses more on energy consumption, battery state of charge stability, and battery life degradation, and its reward function is designed as follows:

[0090]

[0091] in, This indicates the equivalent hydrogen consumption of the battery. Used to punish batteries that experience high-rate charge and discharge during operation, I C Indicates the battery's charge / discharge rate. This is used to guide strategies to maintain a stable state of charge in the battery. This represents the target state of charge that the battery needs to maintain.

[0092] This embodiment provides an energy management method for fuel cell hybrid electric vehicles based on multi-agent reinforcement learning. It employs a learning framework of centralized training and distributed control, which effectively promotes collaboration among multiple power sources and mitigates the instability of the multi-agent interaction environment. Compared to hierarchical and centralized energy management strategies, the energy management strategy based on multi-agent deep reinforcement learning enables efficient collaborative power allocation among multiple power sources, effectively reducing reliance on expert experience and fully leveraging the unique advantages of multiple power sources. It also addresses multiple energy management and control objectives, including reducing overall vehicle hydrogen consumption, maintaining a stable state of charge in the hybrid energy storage system, and delaying the degradation of the core power source's lifespan.

[0093] like Figure 3As shown, in one or more embodiments, the fuel cell hybrid electric vehicle energy management device based on multi-agent reinforcement learning includes a data acquisition module 100 and a model training module 200. The data acquisition module 100 is used to construct a simulation model of the fuel cell hybrid electric vehicle, using the simulation model to simulate preset driving conditions and obtain vehicle power demand data for a predetermined duration. The model training module 200 is used to determine the power allocation results of each power source based on the vehicle power demand data and a pre-constructed energy management strategy model based on multi-agent reinforcement learning; based on the power allocation results, it uses the simulation model to determine the vehicle state information before and after power allocation and obtains the reward value of the energy management strategy model; it trains the energy management strategy model using the power allocation results, vehicle state information, and reward value, completing one round of training; it repeats the training of the energy management strategy model multiple times until the training termination condition is met; wherein, the energy management strategy model includes a fuel cell output power control strategy model based on reinforcement learning and a battery output power control strategy model based on reinforcement learning. For a more detailed description of the working process of the data acquisition module 100 and the model training module 200, please refer to the relevant description of the embodiment of the energy management method for fuel cell hybrid electric vehicles based on multi-agent reinforcement learning, which will not be repeated here.

[0094] In one or more embodiments, a computer-readable storage medium stores a fuel cell hybrid electric vehicle energy management program based on multi-agent reinforcement learning, which, when executed by a processor, implements the fuel cell hybrid electric vehicle energy management method based on multi-agent reinforcement learning described in the above embodiments.

[0095] In one or more embodiments, the computer device includes a computer-readable storage medium, a processor, and a fuel cell hybrid electric vehicle energy management program based on multi-agent reinforcement learning stored in the computer-readable storage medium. When executed by the processor, the fuel cell hybrid electric vehicle energy management program implements the aforementioned fuel cell hybrid electric vehicle energy management method based on multi-agent reinforcement learning. At the hardware level, such as... Figure 4 As shown, the computer device includes a processor 12, an internal bus 13, a network interface 14, and a computer-readable storage medium 11. The processor 12 reads the corresponding computer program from the computer-readable storage medium and then runs it, forming a request processing device at the logical level. Of course, in addition to the software implementation, one or more embodiments of this specification do not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0096] Exemplary examples show that computer-readable storage media, including both permanent and non-permanent, removable and non-removable media, can be implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer-readable storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0097] The specific embodiments of this application have been described in detail above. Although some embodiments have been shown and described, those skilled in the art should understand that modifications and improvements can be made to these embodiments without departing from the principles and spirit of this application as defined by the claims and their equivalents, and such modifications and improvements should also be within the protection scope of this application.

Claims

1. A method for energy management of fuel cell hybrid electric vehicles based on multi-agent reinforcement learning, characterized in that, The energy management method for the fuel cell hybrid electric vehicle includes: Data acquisition phase: Construct a simulation model of the fuel cell hybrid electric vehicle, use the simulation model to simulate preset driving conditions, and obtain vehicle power demand data for a predetermined duration; Model training phase: Based on vehicle power demand data, the power allocation results of each power source are determined using a pre-built energy management strategy model based on multi-agent reinforcement learning; based on the power allocation results, the vehicle state information before and after power allocation is determined using the simulation model, and the reward value of the energy management strategy model is obtained; the energy management strategy model is trained using the power allocation results, the vehicle state information, and the reward value to complete one round of training; the energy management strategy model is trained repeatedly for multiple rounds until the training termination condition is met; Among them, the energy management strategy model includes a fuel cell output power control strategy model based on reinforcement learning and a battery output power control strategy model based on reinforcement learning. The methods for obtaining the reward value of the energy management strategy model include: Based on the power allocation results and vehicle status information, the first part of the reward value is obtained using the reward function of the fuel cell output power control strategy model; Based on the power allocation results and vehicle status information, the second part of the reward value is obtained using the reward function of the battery output power control strategy model; The reward function for the fuel cell output power control strategy model is: , In the formula, The reward function represents the output power control strategy model for fuel cells. and These represent the hydrogen consumption of a fuel cell and the equivalent hydrogen consumption of a supercapacitor, respectively. and This is used to punish fuel cell systems for operating conditions that shorten their lifespan, such as large load fluctuations and idling. This is used to guide strategies to maintain the stable state of charge of supercapacitors, where, This represents the target state of charge that a supercapacitor needs to maintain. Indicates a supercharged state. This represents the minimum value of the supercharged state. This represents the maximum value of the supercharged state; Indicates the current time The output power of the fuel cell, Indicates the previous moment The output power of the fuel cell, This indicates the maximum output power of the fuel cell; The reward function for the battery output power control strategy model is: , in, This represents the reward function of the battery output power control strategy model. This indicates the equivalent hydrogen consumption of the battery. This is used to punish batteries that experience high-rate charge and discharge during operation. Indicates the battery's charge / discharge rate. This is used to guide strategies to maintain a stable state of charge in the battery. This represents the target state of charge that the battery needs to maintain. Indicates the battery's state of charge. This represents the minimum value of the battery's state of charge. This represents the maximum value of the battery's state of charge.

2. The energy management method for fuel cell hybrid electric vehicles based on multi-agent reinforcement learning according to claim 1, characterized in that, The method for obtaining vehicle power demand data for a predetermined duration by simulating preset driving conditions using the simulation model includes: The simulation model is run according to the preset driving conditions to obtain the vehicle driving data at the current moment. The vehicle driving data includes the total mass of the vehicle, gravitational acceleration, rolling resistance coefficient, driving slope, air resistance coefficient, frontal area, vehicle speed, mass coefficient and vehicle acceleration. The vehicle's power demand at the current moment is calculated based on the vehicle's driving data.

3. The energy management method for fuel cell hybrid electric vehicles based on multi-agent reinforcement learning according to claim 1, characterized in that, The methods for determining the power allocation results of each power source based on vehicle power demand data and using a pre-built energy management strategy model based on multi-agent reinforcement learning include: Based on the vehicle's current power demand and the vehicle's status information from the previous moment, the energy management strategy model of the current training round is used to simultaneously determine the output power of the fuel cell and the battery at the current moment. The output power of the supercapacitor at the current moment is determined based on the vehicle's current power demand, the fuel cell's output power, and the battery's output power.

4. The energy management method for fuel cell hybrid electric vehicles based on multi-agent reinforcement learning according to claim 3, characterized in that, Based on the power allocation results, the method for determining vehicle state information before and after power allocation using the simulation model includes: Based on the power allocation results at the current moment, the simulation model is used to determine the vehicle's equivalent hydrogen consumption, the change in the battery's state of charge, the change in the supercapacitor's state of charge, and the degradation of the core power source before and after the power allocation.

5. The energy management method for fuel cell hybrid electric vehicles based on multi-agent reinforcement learning according to claim 3, characterized in that, Methods for simultaneously determining the output power of the fuel cell and the battery at the current moment using the energy management strategy model of the current training round include: The fuel cell output power control strategy model of the current training round is used to determine the output power of the fuel cell at the current moment based on the local observation state information of the fuel cell. The battery output power control strategy model of the current training round determines the battery output power at the current moment based on the local observation state information of the battery.

6. The energy management method for fuel cell hybrid electric vehicles based on multi-agent reinforcement learning according to claim 5, characterized in that, The local observation state information of the fuel cell includes the current vehicle power demand, the previous fuel cell output power, the supercapacitor state of charge, and the difference between the supercapacitor state of charge and the target state of charge; the local observation state information of the battery includes the current vehicle power demand, the previous battery state of charge, the difference between the battery state of charge and the target state of charge, the supercapacitor state of charge, and the difference between the supercapacitor state of charge and the target state of charge.

7. An energy management device for fuel cell hybrid electric vehicles based on multi-agent reinforcement learning, characterized in that, The energy management device for the fuel cell hybrid electric vehicle includes: The data acquisition module is used to construct a simulation model of a fuel cell hybrid electric vehicle, and to use the simulation model to simulate preset driving conditions to obtain vehicle power demand data for a predetermined duration. The model training module is used to determine the power allocation results of each power source based on the vehicle's power demand data and a pre-built energy management strategy model based on multi-agent reinforcement learning; based on the power allocation results, the module uses the simulation model to determine the vehicle state information before and after power allocation and obtains the reward value of the energy management strategy model; the module trains the energy management strategy model using the power allocation results, the vehicle state information, and the reward value to complete one round of training; the module repeats the training of the energy management strategy model for multiple rounds until the training termination condition is met. Among them, the energy management strategy model includes a fuel cell output power control strategy model based on reinforcement learning and a battery output power control strategy model based on reinforcement learning. The methods for obtaining the reward value of the energy management strategy model include: Based on the power allocation results and vehicle status information, the first part of the reward value is obtained using the reward function of the fuel cell output power control strategy model; Based on the power allocation results and vehicle status information, the second part of the reward value is obtained using the reward function of the battery output power control strategy model; The reward function for the fuel cell output power control strategy model is: , In the formula, The reward function represents the output power control strategy model for fuel cells. and These represent the hydrogen consumption of a fuel cell and the equivalent hydrogen consumption of a supercapacitor, respectively. and This is used to punish fuel cell systems for operating conditions that shorten their lifespan, such as large load fluctuations and idling. This is used to guide strategies to maintain the stable state of charge of supercapacitors, where, This represents the target state of charge that a supercapacitor needs to maintain. Indicates a supercharged state. This represents the minimum value of the supercharged state. This represents the maximum value of the supercharged state; Indicates the current time The output power of the fuel cell, Indicates the previous moment The output power of the fuel cell, This indicates the maximum output power of the fuel cell; The reward function for the battery output power control strategy model is: , in, This represents the reward function of the battery output power control strategy model. This indicates the equivalent hydrogen consumption of the battery. This is used to punish batteries that experience high-rate charge and discharge during operation. Indicates the battery's charge / discharge rate. This is used to guide strategies to maintain a stable state of charge in the battery. This represents the target state of charge that the battery needs to maintain. Indicates the battery's state of charge. This represents the minimum value of the battery's state of charge. This represents the maximum value of the battery's state of charge.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a fuel cell hybrid electric vehicle energy management program based on multi-agent reinforcement learning, which, when executed by a processor, implements the fuel cell hybrid electric vehicle energy management method based on multi-agent reinforcement learning as described in any one of claims 1 to 6.

9. A computer device, characterized in that, The computer device includes a computer-readable storage medium, a processor, and a fuel cell hybrid electric vehicle energy management program based on multi-agent reinforcement learning stored in the computer-readable storage medium. When executed by the processor, the fuel cell hybrid electric vehicle energy management program implements the fuel cell hybrid electric vehicle energy management method based on multi-agent reinforcement learning as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Accelerated training method for multi-objective optimization energy management strategy of fuel cell vehicle

    CN116861790A

  • Fuel cell hybrid electric vehicle energy management method based on deep reinforcement learning

    CN117332677A