Multi-target deep reinforcement learning energy management method for fuel cell hybrid vehicle

Through the multi-objective deep reinforcement learning method, combined with neural network model and deep deterministic strategy gradient algorithm, multi-objective optimization of hybrid vehicle energy management is achieved, solving the problem that existing technology is difficult to achieve multi-objective optimization under complex driving conditions, and improving the performance and strategy adaptability of the vehicle.

CN119928814APending Publication Date: 2025-05-06SHANGHAI JIAOTONG UNIV +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411734985.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-29
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing hybrid vehicle energy management strategies are difficult to achieve multi-target optimization under complex and changing driving conditions, especially in multiple aspects such as fuel economy, battery life, power performance and emissions.

Method used

The multi-objective deep reinforcement learning method is adopted to train neural network models by obtaining typical driving data, identify different driving conditions, and define agents in the simulated environment, design multi-objective optimization reward functions, and train agents using deep deterministic strategy gradients to obtain actual energy management strategies.

Benefits of technology

A comprehensive multi-objective optimization of hybrid vehicle energy has been achieved, the adaptability and robustness of energy management strategies have been improved, the optimal strategy can be learned independently in complex driving environments, and the intelligence level of the system and the performance of the vehicle have been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119928814A_ABST
    Figure CN119928814A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-target deep reinforcement learning energy management method for a fuel cell hybrid vehicle, and the method comprises the following steps: S1, obtaining a typical driving data training neural network model, obtaining a driving condition recognition model, and recognizing different driving conditions based on the driving condition recognition model; s2, establishing a simulation environment, defining an intelligent agent in the simulation environment, and designing a multi-objective optimization reward function; s3, performing gradient training on the intelligent agent by adopting a depth deterministic strategy to obtain a trained intelligent agent; and S4, testing the trained intelligent agent, then obtaining actual driving data, outputting an actual working condition based on the driving working condition identification model, and inputting the actual working condition into the trained intelligent agent to obtain an actual energy management strategy. Compared with the prior art, the method has the advantages that energy management of the multi-target hybrid electric vehicle is achieved, and then the overall performance of the hybrid electric vehicle is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of automobile technology, and in particular to a multi-objective deep reinforcement learning energy management method for a fuel cell hybrid vehicle. Background Art

[0002] As a solution to effectively reduce fuel consumption and emissions, hybrid vehicles have attracted widespread attention. Hybrid vehicles combine the advantages of traditional internal combustion engines and electric motors. Through reasonable energy management strategies, they can achieve efficient energy utilization, thereby improving the economy and power performance of the vehicle.

[0003] At present, the energy management strategies of hybrid vehicles mainly include rule-based methods, optimization methods and learning methods. The rule-based method controls the working state of the engine and motor through preset logical rules. This method is simple and easy to implement, but it is difficult to achieve the optimal effect under complex and changeable driving conditions. Optimization methods such as dynamic programming and Pontryagin minimum principle can find the global optimal solution, but they are limited in practical applications due to the large amount of calculation and the need to know the driving conditions in advance. Learning methods such as neural networks and fuzzy logic predict and optimize energy distribution by learning historical data, but these methods usually rely on a large amount of training data and have poor robustness in dealing with sudden changes.

[0004] Driving conditions have a significant impact on the energy management strategy of hybrid electric vehicles. Different driving conditions (such as urban congestion, highway cruising, etc.) will lead to significantly different vehicle power requirements and energy consumption patterns. Traditional energy management strategies often assume that driving conditions are known and fixed, which is obviously unrealistic in actual driving. Therefore, how to identify and adapt to different driving conditions in real time has become a key issue in improving the performance of hybrid electric vehicles.

[0005] In recent years, deep reinforcement learning has made significant progress in autonomous driving and intelligent control systems. Deep reinforcement learning can perform well in uncertain and dynamic environments by learning optimal strategies through interaction with the environment. Applying deep reinforcement learning to the energy management of hybrid vehicles can achieve adaptive control of different driving conditions through continuous learning and optimization, thereby improving the economy and power performance of the vehicle.

[0006] However, most of the existing energy management strategies based on deep reinforcement learning focus on optimizing a single objective, such as minimizing fuel consumption or maximizing battery life. In practical applications, the energy management of hybrid vehicles needs to consider multiple objectives simultaneously, such as fuel economy, battery life, power performance, and emissions. Summary of the invention

[0007] The purpose of the present invention is to provide a multi-objective deep reinforcement learning energy management method for a fuel cell hybrid vehicle in order to achieve multi-objective energy management of the hybrid vehicle and thereby improve the overall performance of the hybrid vehicle.

[0008] The purpose of the present invention can be achieved by the following technical solutions:

[0009] A multi-objective deep reinforcement learning energy management method for a fuel cell hybrid vehicle, the method comprising the following steps:

[0010] S1, acquiring typical driving data to train a neural network model to obtain a driving condition recognition model, and identifying different driving conditions based on the driving condition recognition model;

[0011] S2. Establish a simulation environment, define the agent in the simulation environment, and design a multi-objective optimization reward function;

[0012] S3, using deep deterministic policy gradient to train the agent and obtain a trained agent;

[0013] S4. Test the trained intelligent agent, then obtain actual driving data, output the actual working condition based on the driving condition recognition model, input the actual working condition into the trained intelligent agent, and obtain the actual energy management strategy.

[0014] Furthermore, the specific steps of S1 are:

[0015] A1. Obtain typical driving data;

[0016] A2. Preprocess the driving data to obtain a data set;

[0017] A3. Divide the data set into training set and test set, select the neural network model, optimizer and loss function;

[0018] A4. Use the training set data to train the neural network model;

[0019] A5. Use the test set data to evaluate the recognition accuracy of the model and obtain the trained driving condition recognition model;

[0020] A6. Deploy the driving condition identification model.

[0021] Furthermore, the state space of the agent is:

[0022] State=(v,SOC,P fc )

[0023] Among them, State is the state space, v is the current vehicle speed, SOC is the remaining battery capacity of the power battery, P fc is the fuel cell power.

[0024] Furthermore, the action space of the agent is:

[0025] Action=P fc

[0026] Among them, Action is the action space, P fc is the fuel cell power, which takes values ​​continuously within its power threshold.

[0027] Furthermore, the multi-objective optimization reward function is:

[0028] Reward=-αC H2 -βSOC-SOC target |-γP fc -P fclast |

[0029] Among them, Reward is the reward function, C H2 is the instantaneous hydrogen consumption corresponding to the fuel cell power at the current moment, SOC is the remaining battery capacity of the power battery at the current moment, SOC target is the remaining battery capacity of the target power battery, P fc is the fuel cell power at the current moment, P fclast is the fuel cell power at the previous moment, and α, β, γ are the weight coefficients corresponding to each reward item.

[0030] Furthermore, the specific steps of S3 are:

[0031] B1. Design strategy network and value network;

[0032] B2. Use the experience replay buffer to store the interaction data between the agent and the environment, including the current state, the action selected at the current moment, the reward function at the current moment, the state at the next moment, and the training end mark;

[0033] B3. Randomly sample data from the experience replay buffer to train the policy network and value network, obtain the actor network and critic network, and obtain the preliminary intelligent agent;

[0034] B4. Run the preliminary agent in the simulation environment, output actions based on the current state of the preliminary agent, and the environment returns the new state and reward value, and soft-updates the parameters of the actor network and the critic network;

[0035] B5. Determine whether the training has converged. If it has converged, the trained agent is obtained. Otherwise, return to B2.

[0036] Furthermore, the experience replay buffer is:

[0037] replay_buffer=[State,Action,Reward,NextState,done]

[0038] Among them, replay_buffer is the experience replay buffer, State is the current state, Action is the action selected at the current moment, Reward is the reward function at the current moment, NextState is the next state, and done is the end of training.

[0039] Furthermore, the parameter updates when training the policy network and the value network are:

[0040] policy(State)=argmax policy (value(State,policy(State)))

[0041] =argmin policy (-value(State,policy(State)))

[0042] value(t)=Reward(t)+κ*target_value(t+1)

[0043] Among them, policy is the policy network, State is the current state, value is the value network, target_value is the target value network, and κ is the discount factor.

[0044] Further, the soft update is:

[0045] target_·(t+1)=τ*·(t)+(1-τ)*target_·(t)

[0046] Among them, target_·(t+1) is the target network at the next moment, ·(t) is the network at the current moment, target_·(t) is the target network at the current moment, τ is the soft update coefficient, and ·(t)=·(State(t),Action(t)) is a simple representation of the function at any time.

[0047] Furthermore, judging whether the training has converged is specifically judging whether the difference between the reward functions of two consecutive trainings is less than a threshold. If so, it has converged, otherwise it has not converged.

[0048] Compared with the prior art, the present invention has the following beneficial effects:

[0049] The reward function of the reinforcement learning method adopted by the present invention takes into account the instantaneous hydrogen consumption corresponding to the current fuel cell power, the current fuel cell power and the current remaining battery capacity of the power battery, and comprehensively considers multiple optimization goals such as emission level and battery life, thereby achieving comprehensive energy management. In addition, the reinforcement learning neural network model identifies different driving conditions, so that the energy management strategy can be flexibly adjusted according to different driving conditions, improves the adaptability and robustness of the strategy, and can autonomously learn the optimal strategy in a complex driving environment, thereby improving the intelligence level of the system and the overall performance of the hybrid vehicle. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 It is a flowchart of the module corresponding to the present invention;

[0051] Figure 2 Flow chart of the model for condition identification;

[0052] Figure 3 Training flow chart for energy management model;

[0053] Figure 4 It is a schematic diagram of energy management strategy;

[0054] Figure 5 This is a diagram of the fuel cell vehicle system. DETAILED DESCRIPTION

[0056] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0057] The present invention aims to propose a hybrid vehicle multi-objective deep reinforcement learning energy management strategy based on driving condition recognition, which realizes efficient management and optimization of hybrid vehicle energy by real-time recognition of driving conditions and combining multi-objective optimization algorithms, thereby improving the economy and power performance of the whole vehicle. The method includes the following steps:

[0058] S1, acquiring typical driving data to train a neural network model to obtain a driving condition recognition model, and identifying different driving conditions based on the driving condition recognition model;

[0059] S2. Establish a simulation environment, define the agent in the simulation environment, and design a multi-objective optimization reward function;

[0060] S3, using deep deterministic policy gradient to train the agent and obtain a trained agent;

[0061] S4. Test the trained intelligent agent, then obtain actual driving data, output the actual working condition based on the driving condition recognition model, input the actual working condition into the trained intelligent agent, and obtain the actual energy management strategy.

[0062] The flowchart of the module corresponding to the method is as follows Figure 1 shown.

[0063] The specific steps are:

[0064] S1. Collect typical driving data and use it to train a neural network model that can identify different driving conditions (such as urban congestion, highway cruising, etc.). The trained model will be used in the development of subsequent energy management strategies to ensure that the strategy can be flexibly adjusted according to current driving conditions;

[0065] S2. Design and build a simulation environment that can simulate a real car model as well as an internal power battery model and a fuel cell model. In this environment, define an agent and design a multi-objective optimization reward function that takes into account multiple optimization objectives, such as fuel economy, emission level, battery life, etc., to ensure that the agent can achieve optimal decisions in multiple dimensions;

[0066] S3. Use the Deep Deterministic Policy Gradient (DDPG) algorithm to train the agent in a constructed simulation environment. Through a large amount of simulation training, the agent can learn how to dynamically adjust the working mode and power distribution of the engine and motor under different driving conditions to achieve the best energy efficiency and power performance. During the training process, the agent's strategy is continuously updated until it converges to the optimal solution;

[0067] S4. Load the trained energy management strategy into the whole vehicle model of the hybrid vehicle, conduct detailed simulation tests and put it into actual use. The test content includes but is not limited to fuel economy, power response time, battery charging and discharging efficiency, etc. By comparing with traditional energy management strategies, the effectiveness and advancement of the new strategy can be verified. In addition, small-scale tests can be carried out on actual roads to further evaluate the actual effect of the strategy.

[0068] The process of training the working condition recognition model in S1 is:

[0069] A1. Collect data covering a variety of typical operating conditions (such as urban roads, highways, rural roads, etc.) through appropriate means. Vehicle-mounted sensors and GPS devices can be used to record the vehicle's speed, acceleration, slope and other data under different driving conditions. Data can also be obtained by reconstructing typical automobile driving cycle test conditions.

[0070] A2. Preprocess the collected data, including data cleaning and normalization, removal of outliers and noise, and extraction of data features by matrix differential changes.

[0071] A3. Divide the data set into training set and test set, and select a suitable time series prediction model (such as long short-term memory network, convolutional neural network, etc.), optimizer and loss function.

[0072] A4. Use the training set data to train the neural network model, introduce the attention mechanism module, and ensure the generalization ability of the model through cross-validation and hyperparameter tuning.

[0073] A5. Use the test set data to evaluate the recognition accuracy of the model to ensure that the model can accurately identify different driving conditions.

[0074] A6. Deploy the trained model to the front-end module of the energy management system, receive vehicle sensor data in real time, and output the current driving condition recognition results.

[0075] The process of constructing the environment, agent, and reward function described in S2 is:

[0076] Construct a virtual driving environment, including the construction of a vehicle longitudinal dynamics model, a power battery model, and a fuel cell model. The specific definitions are as follows: the vehicle longitudinal dynamics model obtains the vehicle demand power from the vehicle's current speed; the power battery model obtains the power battery remaining capacity at the next moment from the power battery output power and the power battery remaining capacity at the current moment; the fuel cell model obtains the instantaneous hydrogen consumption of the fuel cell from the fuel cell output power.

[0077] Define the state space (including current vehicle speed, remaining battery capacity of the power battery, fuel cell power) and action space (fuel cell power) of the agent. The specific definitions are as follows:

[0078] State=(v,SOC,P fc )

[0079] Where State is the state space, v is the current vehicle speed, SOC is the remaining battery capacity of the power battery, P fc is the fuel cell power;

[0080] Action=P fc

[0081] Action is the action space, P fc is the fuel cell power, which takes values ​​continuously within its power threshold.

[0082] Design a multi-objective optimization reward function, including fuel economy, fuel cell life, power battery life, fuel cell dynamic response characteristics, etc. The specific definitions are as follows:

[0083] Reward=-αC H2 -βSOC-SOC target |-γ|P fc -P fclast |

[0084] Where Reward is the reward function, C H2 is the instantaneous hydrogen consumption corresponding to the fuel cell power at the current moment, SOC is the remaining battery capacity of the power battery at the current moment, SOC target is the remaining battery capacity of the target power battery, P fc is the fuel cell power at the current moment, P fclast is the fuel cell power at the previous moment, and α, β, γ are the weight coefficients corresponding to each reward item.

[0085] Set up the interactive interface between the environment and the agent. The agent outputs actions based on the current state, and the environment updates the state based on the actions and returns the new state and reward value. The specific definitions are as follows:

[0086] The agent outputs actions according to the current state using a greedy strategy. Specifically, in the early stages of training, the greedy strategy is used to increase the exploratory nature of the action. As the training progresses, the greedy coefficient is gradually reduced, allowing the agent to gradually converge to the optimal strategy. The implementation formula is as follows:

[0087] Action=policy(State)+ε*random(P fcmin ,P fcmax )

[0088] Action is the output action, policy is the policy network, State is the current state, ε is the greedy coefficient, which gradually decays with the number of evaluation rounds, random(P fcmin ,P fcmax ) is any value within the threshold range of the fuel cell output power;

[0089] The environment is updated according to the action status by the vehicle longitudinal dynamics model, power battery model and fuel cell model constructed by the virtual driving environment.

[0090] The process of training an agent using a deep deterministic policy gradient algorithm as described in S3 is:

[0091] The DDPG algorithm combines the advantages of deep learning and reinforcement learning and is suitable for problems in continuous action space.

[0092] B1. Design a strategy network and a value network to generate actions and evaluate the value of actions, respectively. Design a target strategy network and a target value network to generate the next state action and evaluate the value of the corresponding action selected in the next state. This is simulated by a fully connected neural network. Specifically, the input of the strategy network is a state space composed of three features, and the output is a one-dimensional action space. The output is clipped to the fuel cell power threshold range through the sigmoid function. The value network is also simulated by a fully connected neural network. The input of the neural network is a tensor with four features stacked by the state space and the action space, and the output is a one-dimensional value.

[0093] B2. Use the experience replay buffer to store the interaction data between the agent and the environment, including the current state, the action selected at the current moment, the reward function at the current moment, the state at the next moment, and the training end flag. The specific definitions are as follows:

[0094] replay_buffer=[State,Action,Reward,NextState,done]

[0095] Among them, replay_buffer is the experience replay buffer, State is the current state, Action is the action selected at the current moment, Reward is the reward function at the current moment, NextState is the next state, and done is the end of training.

[0096] B3. Randomly sample data from the experience playback buffer for training to improve data utilization. Use the gradient descent method to train the strategy network and value network, and update the network parameter formula as follows:

[0097] policy(State)=argmax policy (value(State,policy(State)))

[0098] =argmin policy (-value(State,policy(State)))

[0099] value(t)=Reward(t)+κ*target_value(t+1)

[0100] Among them, policy is the policy network, State is the current state, value is the value network, target_value is the target value network, ·(t)=·(State(t),Action(t)) is a simple representation of the function at any time, and κ is the discount factor.

[0101] B4. Run the agent in the simulation environment, output actions based on the current state, and the environment returns the new state and reward value. Use this data to soft-update the parameters of the actor network and critic network (to prevent gradient explosion and speed up convergence), and continuously optimize the agent's strategy. The formula for soft-updating network parameters is as follows:

[0102] target_·(t+1)=τ*·(t)+(1-τ)*target_·(t)

[0103] Among them, target_·(t+1) is the target network at the next moment, ·(t) is the network at the current moment, target_·(t) is the target network at the current moment, and τ is the soft update coefficient.

[0104] B5. Determine whether the training has converged by monitoring the changing trend of the reward value. Stop training when the reward value tends to be stable and reaches the expected goal.

[0105] In the above steps, the energy management model training flow chart is as follows: Figure 3 The working condition identification model flow chart is shown in Figure 2 The energy management strategy diagram is shown in Figure 4 As shown in the figure, the fuel cell vehicle system is as follows Figure 5 shown.

[0106] The energy management strategy proposed in the present invention not only focuses on fuel economy, but also comprehensively considers multiple optimization objectives such as emission levels and battery life, thereby realizing comprehensive energy management; the energy management strategy proposed in the present invention identifies different driving conditions through a neural network model, so that the energy management strategy can be flexibly adjusted according to different driving conditions, thereby improving the adaptability and robustness of the strategy; the energy management strategy proposed in the present invention utilizes a deep reinforcement learning algorithm, and the intelligent agent can autonomously learn the optimal strategy in a complex driving environment, thereby improving the intelligence level of the system.

[0107] The energy management strategy proposed in the present invention not only focuses on fuel economy, but also comprehensively considers multiple optimization objectives such as emission levels and battery life, thereby realizing comprehensive energy management; the energy management strategy proposed in the present invention identifies different driving conditions through a neural network model, so that the energy management strategy can be flexibly adjusted according to different driving conditions, thereby improving the adaptability and robustness of the strategy; the energy management strategy proposed in the present invention utilizes a deep reinforcement learning algorithm, and the intelligent agent can autonomously learn the optimal strategy in a complex driving environment, thereby improving the intelligence level of the system.

[0108] The preferred specific embodiments of the present invention are described in detail above. It should be understood that a person skilled in the art can make many modifications and changes based on the concept of the present invention without creative work. Therefore, any technical solution that can be obtained by a person skilled in the art through logical analysis, reasoning or limited experiments based on the concept of the present invention on the basis of the prior art should be within the scope of protection determined by the claims.

Claims

1. A multi-objective deep reinforcement learning energy management method for a fuel cell hybrid vehicle, characterized in that: The method comprises the following steps: S1, acquiring typical driving data to train a neural network model to obtain a driving condition recognition model, and identifying different driving conditions based on the driving condition recognition model; S2. Establish a simulation environment, define the agent in the simulation environment, and design a multi-objective optimization reward function; S3, using deep deterministic policy gradient to train the agent and obtain a trained agent; S4. Test the trained intelligent agent, then obtain actual driving data, output the actual working condition based on the driving condition recognition model, input the actual working condition into the trained intelligent agent, and obtain the actual energy management strategy.

2. A fuel cell hybrid vehicle multi-objective deep reinforcement learning energy management method according to claim 1, characterized in that: The specific steps of S1 are: A1. Obtain typical driving data; A2. Preprocess the driving data to obtain a data set; A3. Divide the data set into training set and test set, select the neural network model, optimizer and loss function; A4. Use the training set data to train the neural network model; A5. Use the test set data to evaluate the recognition accuracy of the model and obtain the trained driving condition recognition model; A6. Deploy the driving condition identification model.

3. The multi-objective deep reinforcement learning energy management method for a fuel cell hybrid vehicle according to claim 1, characterized in that: The state space of the agent is: State=(v,SOC,P fc ) Among them, State is the state space, v is the current vehicle speed, SOC is the remaining battery capacity of the power battery, P fc is the fuel cell power.

4. A fuel cell hybrid vehicle multi-objective deep reinforcement learning energy management method according to claim 3, characterized in that: The action space of the agent is: Action=P fc Among them, Action is the action space, P fc is the fuel cell power, which takes values ​​continuously within its power threshold.

5. A fuel cell hybrid vehicle multi-objective deep reinforcement learning energy management method according to claim 4, characterized in that: The multi-objective optimization reward function is: Among them, Reward is the reward function, is the instantaneous hydrogen consumption corresponding to the fuel cell power at the current moment, SOC is the remaining battery capacity of the power battery at the current moment, SOC target is the remaining battery capacity of the target power battery, P fc is the fuel cell power at the current moment, P fclast is the fuel cell power at the previous moment, and α, β, γ are the weight coefficients corresponding to each reward item.

6. A fuel cell hybrid vehicle multi-objective deep reinforcement learning energy management method according to claim 1, characterized in that: The specific steps of S3 are: B1. Design strategy network and value network; B2. Use the experience replay buffer to store the interaction data between the agent and the environment, including the current state, the action selected at the current moment, the reward function at the current moment, the state at the next moment, and the training end mark; B3. Randomly sample data from the experience replay buffer to train the policy network and value network, obtain the actor network and critic network, and obtain the preliminary intelligent agent; B4. Run the preliminary agent in the simulation environment, output actions based on the current state of the preliminary agent, and the environment returns the new state and reward value, and soft-updates the parameters of the actor network and the critic network; B5. Determine whether the training has converged. If it has converged, the trained agent is obtained. Otherwise, return to B2.

7. A fuel cell hybrid vehicle multi-objective deep reinforcement learning energy management method according to claim 6, characterized in that: The experience replay buffer is: replay_buffer=[State,Action,Reward,NextState,done] Among them, replay_buffer is the experience replay buffer, State is the current state, Action is the action selected at the current moment, Reward is the reward function at the current moment, NextState is the next state, and done is the end of training.

8. The multi-objective deep reinforcement learning energy management method for a fuel cell hybrid vehicle according to claim 7, characterized in that: The parameter updates when training the policy network and value network are: policy(State)=argmax policy (value(State,policy(State))) =argmin policy (-value(State,policy(State))) value(t)=Reward(t)+κ*target_value(t+1) Among them, policy is the policy network, State is the current state, value is the value network, target_value is the target value network, and κ is the discount factor.

9. A fuel cell hybrid vehicle multi-objective deep reinforcement learning energy management method according to claim 8, characterized in that: Soft updates are: target_·(t+1)=τ*·(t)+(1-τ)*target_·(t) Among them, target_·(t+1) is the target network at the next moment, ·(t) is the network at the current moment, target_·(t) is the target network at the current moment, τ is the soft update coefficient, and ·(t)=·(State(t),Action(t)) is a simple representation of the function at any time.

10. The multi-objective deep reinforcement learning energy management method for a fuel cell hybrid vehicle according to claim 8, characterized in that: To determine whether the training has converged, we need to determine whether the difference between the reward functions of two consecutive trainings is less than a threshold. If so, it has converged, otherwise it has not converged.

Citation Information

Cited By

  • Energy management method based on navigation and reinforcement learning

    CN121133668A