Hybrid electric vehicle eco-driving method based on multi-objective deep reinforcement learning

By using multi-objective deep reinforcement learning algorithms and conditional networks, the energy management strategy of hybrid electric vehicles is optimized, solving the dynamic trade-off between fuel efficiency and battery health, and achieving efficient and stable operation of the eco-driving strategy.

CN116257937BActive Publication Date: 2026-03-31SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-05
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies in hybrid vehicles struggle to effectively balance fuel efficiency and battery health through single-objective deep reinforcement learning, and setting reward weights requires additional time and effort, resulting in limited benefits for eco-driving strategies.

Method used

A multi-objective deep reinforcement learning algorithm is adopted to construct an ecological driving strategy based on conditional networks. Combined with a reward weight sampling mechanism, the energy management strategy is optimized, the weights of each objective are adjusted in real time, and the adaptive cruise and power system are optimized in a coordinated manner. Taking into account the cost of battery degradation, the vehicle power distribution is optimized.

Benefits of technology

It improves the overall efficiency of the hybrid electric vehicle's eco-driving strategy, exhibits excellent fuel economy and battery health stability under various complex driving conditions, has a faster computational convergence speed than traditional methods, and possesses strong adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116257937B_ABST
    Figure CN116257937B_ABST
Patent Text Reader

Abstract

The application discloses a hybrid electric vehicle ecological driving method based on multi-objective deep reinforcement learning, and belongs to the technical field of deep reinforcement learning. The method comprises the following steps: constructing a model of an adaptive cruise control system (ACC) and a power system of a hybrid electric vehicle; using a MODRL algorithm to establish an energy consumption optimization method of the hybrid electric vehicle under a following scene based on MODRL; further using a conditional network (CN) to establish an input network of a weight corresponding to each optimization target, so that the MODRL algorithm is first applied to multi-system collaborative optimization of the hybrid electric vehicle, and adaptive selection of multi-target weights is realized in combination with a reward weight sampling mechanism. The method can solve the multi-target trade-off problem involved in ecological driving of the hybrid electric vehicle, so that the development cycle is shortened while driving and power performance are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of deep reinforcement learning technology, and more specifically, relates to a hybrid electric vehicle ecological driving method based on multi-objective deep reinforcement learning. Background Technology

[0002] Hybrid electric vehicles (HEVs) use clean, economical, and environmentally friendly electricity as one of their driving energy sources, and the promotion of HEVs is expected to alleviate excessive emissions of pollutants. However, the investment cycle for HEV configuration improvements is typically measured in years, with limited benefits. Eco-driving technologies, on the other hand, have a shorter deployment time and lower initial investment, yet they can increase fuel efficiency by up to 45%. Therefore, promoting the application of eco-driving technologies and improving vehicle driving strategies can effectively enhance vehicle energy management performance. The synergistic optimization of Energy Management Systems (EMS) and Adaptive Cruise Control (ACC) is currently a research hotspot in the field of eco-driving. EMS improves fuel economy through the coordination of various components in the powertrain system, while ACC enhances driving safety and comfort by helping the driver adjust vehicle speed or following distance in real time.

[0003] Currently, autonomous learning methods are gradually becoming the preferred approach for solving optimization problems. Deep reinforcement learning (DRL) is a method that combines deep learning (DL) and reinforcement learning (RL), possessing both the powerful representational capabilities of deep learning and the strong reasoning capabilities of reinforcement learning. Deep neural networks (DNNs) significantly reduce dependence on domain knowledge. With the increasing adoption of DRL in EMS and ACC fields, eco-driving strategies based on DRL have been proposed accordingly. However, solving problems through single-objective DRL requires additional time and effort to manually determine reward weights; furthermore, the experience buffer gained from the optimal strategy of a certain weight vector may adversely affect other weight vectors. In recent years, research on multi-objective deep reinforcement learning (MODRL) has made some progress. Deep Q-learning networks (DQNs), i.e., conditional networks (CNs), which use the relative weights of objectives as input conditions, can effectively solve multi-objective trade-offs and high-dimensional input problems, which is very beneficial for realizing eco-driving. Therefore, using an eco-driving strategy based on MODRL to realize multi-objective dynamic trade-offs in hybrid electric vehicle eco-driving can control the vehicle's power distribution online in real time, effectively improving the overall efficiency of the eco-driving strategy. Summary of the Invention

[0004] To address the aforementioned technical problems in this field, this invention provides a hybrid electric vehicle eco-driving method based on multi-objective deep reinforcement learning. An eco-driving strategy based on EMS and ACC collaborative optimization is constructed, incorporating battery degradation costs into the optimization objective. Through a CN-based deep learning model combined with a reward weight sampling mechanism, the weights of each objective are adjusted online in real time to optimize the vehicle's energy management strategy and appropriately switch to a car-following model, thereby maximizing the benefits of eco-driving.

[0005] To address at least one of the aforementioned technical problems, according to one aspect of the present invention, a hybrid electric vehicle eco-driving method based on multi-objective deep reinforcement learning is provided, comprising the following steps:

[0006] S1. Construct a model of the adaptive cruise control system and the powertrain system for hybrid electric vehicles;

[0007] S2. Using a multi-objective deep reinforcement learning algorithm, establish an energy consumption optimization method for hybrid electric vehicles in a car-following scenario.

[0008] S3. Based on neural networks, construct a conditional network based on the target relative weight input;

[0009] S4. Apply multi-objective deep reinforcement learning algorithms to the collaborative optimization of adaptive cruise and energy management. Combined with a reward weight sampling mechanism, establish a hybrid electric vehicle ecological driving strategy based on multi-objective deep reinforcement learning to improve the optimization performance of energy management in vehicle following scenarios.

[0010] According to an embodiment of the present invention, the hybrid electric vehicle eco-driving method based on multi-objective deep reinforcement learning may optionally include the following step S1:

[0011] A hybrid electric vehicle adaptive cruise control system model and a powertrain model are constructed. The adaptive cruise control system mainly realizes the reorganization of the vehicle following model. Based on the dynamic changes in the driving conditions of the vehicle in front, the appropriate following model is selected, including the Krauss vehicle following model and the Intelligent Driving Model (IDM). The powertrain mainly realizes the energy coordination between the engine-generator set (EGS) and the battery pack, and considers the aging problem of the battery pack, and establishes a power battery electrothermal aging system model.

[0012] According to an embodiment of the present invention, the hybrid electric vehicle eco-driving method based on multi-objective deep reinforcement learning may optionally include the following step S2:

[0013] The energy consumption optimization method for hybrid electric vehicles in a car-following scenario based on a multi-objective deep reinforcement learning algorithm can be viewed as a Markov decision process, including the following steps:

[0014] S21. Define the state, action, multi-objective reward function, optimal action-value function, and optimal control policy in deep reinforcement learning;

[0015] S22. The deep reinforcement learning agent receives environmental observations and performs an action based on the current control policy.

[0016] S23. The environment responds to this action, enters a new state, and returns the new state and the reward brought by this action to the deep reinforcement learning agent.

[0017] S24. In the new state, the agent will continue to perform actions, and so on. The deep reinforcement learning agent interacts with the environment continuously until the optimal action-value function (multi-objective Q-value vector) and the optimal control policy are obtained.

[0018] According to an embodiment of the present invention, the hybrid electric vehicle eco-driving method based on multi-objective deep reinforcement learning may optionally include step S21 as follows:

[0019] In deep reinforcement learning, the states and actions, multi-objective reward function, optimal action-value function, and optimal control strategy are determined. Specifically, the states include: the current speed of the master vehicle, the current acceleration of the master vehicle, the current distance traveled by the master vehicle, the current speed of the vehicle ahead, the current acceleration of the vehicle ahead, the current distance traveled by the vehicle ahead, the current following distance, the current engine power, the state of charge (SoC) of the power battery, the state of health (SoH) of the power battery, the average internal temperature of the battery, and the battery capacity decay rate (c). The actions are the following behavior mode and engine power. The reward function is defined, comprising three parts: low fuel consumption, SoC stability, and SoH stability. The specific calculation formula for the reward function is as follows:

[0020]

[0021] In the above formula, R(s,a) is the reward function vector for choosing action a in state s, with each objective assigned a corresponding weight; R1(s,a) is reward function reward 1; R2(s,a) is reward function reward 2; R3(s,a) is reward function reward 3; C f C represents the engine's instantaneous fuel consumption. b Cost of charging batteries; C ag The cost is the battery aging cost; M and V are standardization coefficients.

[0022] The specific formula for calculating the optimal action-value function is as follows:

[0023] Q * (s,a)=Q π(s,a)=max E[R t+1 +λQ * (s t+1 ,a t+1 )|s t ,a t (2)

[0024] In the above formula, Q π (s,a) is the action-value function for choosing action a under policy π state s; t ,a t Let s be the state and action at time t; t+1 ,a t+1 ,R t+1 Let be the state, action, and reward function at time t+1; λ∈[0,1] is the discount factor;

[0025] Optimal control strategy π * satisfy The specific calculation formula is as follows:

[0026]

[0027] According to an embodiment of the present invention, the hybrid electric vehicle eco-driving method based on multi-objective deep reinforcement learning may optionally include step S3 as follows:

[0028] S31. Establish a conditional network based on the target relative weight input;

[0029] S32. Define the control action selection strategy based on the output results of the conditional network;

[0030] S33. Employ Diverse Experience Replay (DER) to perform empirical sampling on the weight vectors of recently unexecuted strategies.

[0031] According to the multi-objective deep reinforcement learning-based hybrid electric vehicle ecological driving method of the present invention, optionally, the conditional network based on the relative weight input of the targets in step S31 is essentially a value-based neural network, specifically including an estimation neural network and a target neural network. The target network is used to update the estimation network. The two have the same internal structure, taking the relative importance of the targets as the input condition, and combining the reward weight sampling mechanism to achieve the trade-off of multiple objectives. When establishing the conditional network, there are three inputs: one is the state observation value, the second is the control quantity, and the third is the target weight. The state observation value includes the current speed of the master vehicle, the current acceleration of the master vehicle, the current distance traveled by the master vehicle, the current speed of the preceding vehicle, the current acceleration of the preceding vehicle, the current distance traveled by the preceding vehicle, the current following distance of the vehicle, the current engine power, the state of charge (SoC) of the power battery, the state of health (SoH) of the power battery, the average internal temperature of the battery, and the battery capacity decay rate (c). The output is a multi-objective Q-value vector.

[0032] According to the multi-objective deep reinforcement learning-based hybrid electric vehicle eco-driving method of the present invention, optionally, the definition of the control action selection strategy in step S32 specifically includes: obtaining the execution action from this random process based on the current strategy and a certain action selection probability ε, which can be expressed as:

[0033] P(a = random choice) = ε

[0034]

[0035] In the above formula, P(·) represents the probability of choosing an action.

[0036] According to an embodiment of the present invention, the hybrid electric vehicle eco-driving method based on multi-objective deep reinforcement learning may optionally include the following steps in step S4:

[0037] S41. Offline training: The model is trained using a multi-objective deep reinforcement learning algorithm to learn the control policy, i.e., the mapping relationship between the input state and the action parameters.

[0038] S42. Read out the parameters of each trained conditional network and download the control strategy to the vehicle controller (VCU).

[0039] S43. Online learning: Obtain relevant information about the vehicle and battery status at the current moment, and apply it to the trained conditional network. Through online real-time adjustment, update the car-following model selection and power allocation decisions.

[0040] According to an embodiment of the present invention, the hybrid electric vehicle eco-driving method based on multi-objective deep reinforcement learning may optionally include the following steps in step S41: training the model using a multi-objective deep reinforcement learning algorithm to learn the mapping relationship between input states and action parameters.

[0041] S411. Initialize the diverse experience replay buffer D; deduplicat weight vector experience pool W.

[0042] S412. Initialize the conditional network value function; obtain the weight vector w at the current time step. t And add it to W;

[0043] S413, the agent randomly selects an action a with probability ε. t The order was given to the environment to execute the a t Otherwise, perform the action.

[0044] S414, environment executes this a t Return rewardr t and new state t+1 ;

[0045] S415, the agent will perform this state transition process: (s t ,a t ,r t ,s t+1 Saved to the diverse experience playback buffer D;

[0046] S416. Randomly select a portion of samples from the diverse experience replay buffer D, using (s j ,a j ,r j ,s j+1 ) indicates that w is randomly selected from the experience pool W of the weight vector. j Then, the target neural network is trained and updated. The learning process is as follows:

[0047]

[0048] In the above formula, y j and y′ j For tags; r j λ represents the reward during the learning process; λ is the discount factor.

[0049] S417. Define TD error for network updates. TD error is:

[0050]

[0051] S418, every time N passes - In each round, the parameters of the online network are copied to the target network;

[0052] S419. Once the training steps are completed, the conditional network training is complete.

[0053] According to another aspect of the present invention, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the hybrid electric vehicle eco-driving method based on multi-objective deep reinforcement learning of the present invention.

[0054] According to another aspect of the present invention, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the hybrid electric vehicle eco-driving method based on multi-objective deep reinforcement learning of the present invention.

[0055] Compared with the prior art, the present invention has at least the following beneficial effects:

[0056] This invention applies a multi-objective deep reinforcement learning algorithm to the collaborative optimization of multiple systems in hybrid electric vehicles (HEVs) for the first time, aiming to improve the overall efficiency of HEV eco-driving strategies. Based on traditional energy indicators, this method incorporates battery degradation costs into the optimization objective, improving overall energy management performance. A conditional network based on multi-objective dynamic weight settings is employed to control the vehicle to learn an economy-oriented reconfiguration car-following model, achieving optimized power allocation. Simulation results show that under 100km of real-world urban and interstate highway driving conditions, the CN-based eco-driving strategy converges more than three times faster than the DDQN-based strategy. The former achieves a dynamic trade-off between fuel economy, SoC stability, and SoH stability, while the latter exhibits significantly lower fuel economy and near-depletion of SoC. The CN-based eco-driving strategy demonstrates strong adaptability under various complex driving conditions. Attached Figure Description

[0057] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings of the embodiments will be briefly described below. Obviously, the drawings described below only relate to some embodiments of the present invention and are not intended to limit the present invention.

[0058] Figure 1 This is a schematic diagram of the car-following model algorithm of the present invention;

[0059] Figure 2 This is a schematic diagram of the structure of the present invention. Detailed Implementation

[0060] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention.

[0061] Unless otherwise defined, the technical or scientific terms used herein shall have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains.

[0062] like Figure 1-2 As shown,

[0063] Example 1:

[0064] A hybrid electric vehicle eco-driving strategy based on multi-objective deep reinforcement learning includes the following steps:

[0065] S1. Construct a model of the adaptive cruise control system and the powertrain system for hybrid electric vehicles;

[0066] S2. Using a multi-objective deep reinforcement learning algorithm, establish an energy consumption optimization method for hybrid electric vehicles in a car-following scenario.

[0067] S3. Based on neural networks, construct a conditional network based on the target relative weight input;

[0068] S4. Apply multi-objective deep reinforcement learning algorithms to the collaborative optimization of adaptive cruise and energy management. Combined with a reward weight sampling mechanism, establish a hybrid electric vehicle ecological driving strategy based on multi-objective deep reinforcement learning to improve the optimization performance of energy management in vehicle following scenarios.

[0069] Figure 1 This is a schematic diagram of the theoretical carousel model algorithm provided in this embodiment of the invention. Please refer to [link / reference]. Figure 1 .

[0070] A hybrid electric vehicle adaptive cruise control system model and a powertrain model are constructed. The adaptive cruise control system mainly realizes the reorganization of the vehicle following model. Based on the dynamic changes in the driving conditions of the vehicle in front, the appropriate following model is selected, including the Krauss vehicle following model and the Intelligent Driving Model (IDM). The powertrain mainly realizes the energy coordination between the two power sources, the engine-generator set (EGS) and the battery pack. Considering the aging problem of the battery pack, a power battery electrothermal aging system model is established, including a second-order RC model and a two-state thermal model.

[0071] The energy consumption optimization method for hybrid electric vehicles in car-following scenarios based on multi-objective deep reinforcement learning algorithms can be viewed as a Markov decision process, specifically including the following steps:

[0072] S21. Define the state, action, multi-objective reward function, optimal action-value function, and optimal control policy in deep reinforcement learning;

[0073] S22. The deep reinforcement learning agent receives environmental observations and performs an action based on the current control policy.

[0074] S23. The environment responds to this action, enters a new state, and returns the new state and the reward brought by this action to the deep reinforcement learning agent.

[0075] S24. In the new state, the agent will continue to perform actions, and so on. The deep reinforcement learning agent interacts with the environment continuously until the optimal action-value function (multi-objective Q-value vector) and the optimal control policy are obtained.

[0076] In step S21 above, the states and actions in deep reinforcement learning, the multi-objective reward function, the optimal action-value function, and the optimal control strategy are determined. Specifically, the states include: the current speed of the master vehicle, the current acceleration of the master vehicle, the current distance traveled by the master vehicle, the current speed of the vehicle in front, the current acceleration of the vehicle in front, the current distance traveled by the vehicle in front, the current following distance, the current engine power, the state of charge (SoC) of the power battery, the state of health (SoH) of the power battery, the average internal temperature of the battery, and the battery capacity decay rate (c). The actions are the following behavior mode and the engine power. The reward function is defined, which includes three parts: low fuel consumption, SoC stability, and SoH stability. The specific calculation formula for the reward function is as follows:

[0077]

[0078] In the above formula, R(s,a) is the reward function vector for choosing action a in state s, with each objective assigned a corresponding weight; R1(s,a) is reward function reward 1; R2(s,a) is reward function reward 2; R3(s,a) is reward function reward 3; C f C represents the engine's instantaneous fuel consumption. b Cost of charging batteries; C ag The cost is the battery aging cost; M and V are standardization coefficients.

[0079] Battery charging cost C b The specific calculation formula is as follows:

[0080]

[0081] In the above formula, SoC ref At a distance of x h The spatial domain index of the SoC is λ, where λ is the decay rate of the spatial domain index SoC. init SoC sust These are the initial SoC value (0.5) and the sustained SoC value (0.2), respectively; x total This is the expected driving range (100km) after the battery is fully charged;

[0082] Battery aging cost C ag The specific calculation formula is as follows:

[0083] C ag =ΔSoH (8)

[0084] In the above formula, ΔSoH is the degradation value of the power battery's health state;

[0085] The specific formula for calculating the optimal action-value function is as follows:

[0086] Q * (s,a)=Q π (s,a)=max E[R t+1 +λQ * (s t+1 ,a t+1 )|s t ,a t (2)

[0087] In the above formula, Q π (s,a) is the action-value function for choosing action a under policy π state s; t ,a t Let s be the state and action at time t; t+1 ,a t+1 ,R t+1 Let be the state, action, and reward function at time t+1; λ∈[0,1] is the discount factor;

[0088] Optimal control strategy π * satisfy The specific calculation formula is as follows:

[0089]

[0090] The construction of a conditional network based on the target relative weight input includes the following steps:

[0091] S31. Establish a conditional network based on the target relative weight input;

[0092] S32. Define the control action selection strategy based on the output results of the conditional network;

[0093] S33. Employ Diverse Experience Replay (DER) to perform empirical sampling on the weight vectors of recently unexecuted strategies.

[0094] The conditional network based on the relative weight of the target input described in step S31 above is essentially a value-based neural network, specifically including an estimation neural network and a target neural network. The target network is used to update the estimation network. The two have the same internal structure, taking the relative importance of the target as the input condition and combining a reward weight sampling mechanism to achieve a trade-off between multiple objectives. When establishing the conditional network, there are three inputs: one is the state observation value, the second is the control quantity, and the third is the target weight. The state observation value includes the current speed of the main vehicle, the current acceleration of the main vehicle, the current distance traveled by the main vehicle, the current speed of the vehicle in front, the current acceleration of the vehicle in front, the current distance traveled by the vehicle in front, the current following distance of the vehicle, the current engine power, the state of charge (SoC) of the power battery, the state of health (SoH) of the power battery, the average internal temperature of the battery, and the battery capacity decay rate (c). The output is a multi-objective Q-value vector.

[0095] The definition of the control action selection strategy in step S32 above specifically includes: obtaining the action to be executed from this random process based on the current strategy and a certain action selection probability ε. This process can be represented as:

[0096] P(a = random choice) = ε

[0097]

[0098] In the above formula, P(·) represents the probability of choosing an action.

[0099] Figure 2 This is a schematic diagram of the hybrid electric vehicle ecological driving strategy structure based on multi-objective deep reinforcement learning provided in an embodiment of the present invention. Please refer to [link / reference]. Figure 2 .

[0100] Establishing a hybrid vehicle eco-driving strategy based on multi-objective deep reinforcement learning includes the following steps:

[0101] S41. Offline training: The model is trained using a multi-objective deep reinforcement learning algorithm to learn the control policy, i.e., the mapping relationship between the input state and the action parameters.

[0102] S42. Read out the parameters of each trained conditional network and download the control strategy to the vehicle controller (VCU).

[0103] S43. Online learning: Obtain relevant information about the vehicle and battery status at the current moment, and apply it to the trained conditional network. Through online real-time adjustment, update the car-following model selection and power allocation decisions.

[0104] In step S41 above, the model is trained using a multi-objective deep reinforcement learning algorithm to learn the mapping relationship between input states and action parameters. Specifically, this includes the following steps:

[0105] S411. Initialize the diverse experience replay buffer D; deduplicat weight vector experience pool W.

[0106] S412. Initialize the conditional network value function; obtain the weight vector w at the current time step. t And add it to W;

[0107] S413, the agent randomly selects an action a with probability ε. t The order was given to the environment to execute the a t Otherwise, perform the action.

[0108] S414, environment executes this at Return rewardr t and new state t+1 ;

[0109] S415, the agent will perform this state transition process: (s t ,a t ,r t ,s t+1 Saved to the diverse experience playback buffer D;

[0110] S416. Randomly select a portion of samples from the diverse experience replay buffer D, using (s j ,a j ,r j ,s j+1 ) indicates that w is randomly selected from the experience pool W of the weight vector. j Then, the target neural network is trained and updated. The learning process is as follows:

[0111]

[0112] In the above formula, y j and y′ j For tags; r j λ represents the reward during the learning process; λ is the discount factor.

[0113] S417. Define TD error for network updates. TD error is:

[0114]

[0115] S418, every time N passes - In each round, the parameters of the online network are copied to the target network;

[0116] S419. Once the training steps are completed, the conditional network training is complete.

[0117] Example 2:

[0118] The computer-readable storage medium of this embodiment stores a computer program that, when executed by a processor, implements the steps in the hybrid electric vehicle eco-driving method based on multi-objective deep reinforcement learning of Embodiment 1.

[0119] The computer-readable storage medium in this embodiment can be an internal storage unit of the terminal, such as the terminal's hard disk or memory; the computer-readable storage medium in this embodiment can also be an external storage device of the terminal, such as a plug-in hard disk, smart memory card, secure digital card, flash memory card, etc. equipped on the terminal; furthermore, the computer-readable storage medium can include both the terminal's internal storage unit and external storage devices.

[0120] The computer-readable storage medium of this embodiment is used to store computer programs and other programs and data required by the terminal. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.

[0121] Example 3:

[0122] The computer device of this embodiment includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps in the hybrid electric vehicle eco-driving method based on multi-objective deep reinforcement learning of Embodiment 1.

[0123] In this embodiment, the processor can be a central processing unit, or other general-purpose processors, digital signal processors, application-specific integrated circuits, off-the-shelf programmable gate arrays or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. The memory can include read-only memory and random access memory, and provides instructions and data to the processor. A portion of the memory can also include non-volatile random access memory. For example, the memory can also store device type information.

[0124] Those skilled in the art will understand that the content disclosed in the embodiments can be provided as a method, system, or computer program product. Therefore, this solution can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this solution can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage) containing computer-usable program code.

[0125] This solution is described with reference to flowchart illustrations and / or block diagrams of methods and computer program products according to embodiments of this solution. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0126] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0127] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0128] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0129] The examples described herein are merely preferred embodiments of the invention and are not intended to limit the concept and scope of the invention. Any modifications and improvements made by those skilled in the art to the technical solutions of the invention without departing from the design concept of the invention should fall within the protection scope of the invention.

Claims

1. A hybrid electric vehicle eco-driving method based on multi-objective deep reinforcement learning, characterized in that, It comprises the following steps: S1, constructing a hybrid electric vehicle adaptive cruise control system model and a power system model; S2, using a multi-objective deep reinforcement learning algorithm, establishing an energy consumption optimization method for a hybrid electric vehicle following scenario based on a multi-objective deep reinforcement learning algorithm; S3, based on a neural network, constructing a conditional network based on target relative weight input; S4, applying the multi-objective deep reinforcement learning algorithm to adaptive cruise control and energy management collaborative optimization, combining a reward weight sampling mechanism, establishing a hybrid electric vehicle eco-driving strategy based on multi-objective deep reinforcement learning, and improving the optimization performance of energy management in a vehicle following scenario; In step S4, the hybrid electric vehicle eco-driving strategy based on multi-objective deep reinforcement learning comprises the following steps: S41, offline training; the model is trained by the multi-objective deep reinforcement learning algorithm to learn the mapping relationship between the input state and the action parameter; S42, the parameters of each trained conditional network are read out and the control strategy is downloaded to the vehicle control unit VCU; S43, online learning; the relevant information about the vehicle state and the battery state at the current time is obtained, which jointly acts on the trained conditional network, and the update of the following model selection and power distribution decision is completed through online real-time adjustment; In step S41, the mapping relationship between the input state and the action parameter is learned by training the model by the multi-objective deep reinforcement learning algorithm, which comprises the following steps: S411, initialize the multi-sample experience replay buffer D; remove the weight vector experience pool W; S412, initialize the condition network value function; obtain the weight vector w of the current time t and add it to W; S413、agent with probability ε randomly selects an action a t , to the environment to perform the a t ; otherwise, perform action S414, the environment executes the a t , returns the reward r t and the new state s t+1 ; S415, the agent saves this state transition: (s t , a t , r t , s t+1 ) to the diverse experience replay buffer D; S416, randomly select part of samples from the diverse experience replay buffer D, use (s j , a j , r j , s j+1 ) to represent; randomly select w j from the weight vector experience pool W; then train and update the target neural network, and the learning process is as follows: In the above equation, y j and y′ j are labels; r j is the reward during learning; and λ is a discount factor. S417, define TD error for network update, TD error is: S418, every N - round, copy the online network parameters to the target network; S419, when the training step is completed, the conditional network training is completed.

2. The method of claim 1, wherein, Step S1 is as follows: constructing a hybrid electric vehicle adaptive cruise control system model and a power system model, wherein the adaptive cruise control system mainly realizes vehicle following model reconstruction, selects a suitable following model according to the dynamic change of the front vehicle driving condition, including Krauss vehicle following model and intelligent driving model; the power system mainly realizes energy coordination between engine-generator set and battery pack, and considers the aging problem of the battery pack, and establishes a power battery electro-thermal aging system model.

3. The method of claim 1, wherein, Step S2 is as follows: The energy consumption optimization method for a hybrid electric vehicle following scenario based on a multi-objective deep reinforcement learning algorithm can be regarded as a Markov decision process, comprising the following steps: S21, define the state, action, multi-objective reward function, optimal action-value function and optimal control strategy in deep reinforcement learning; S22, the deep reinforcement learning agent receives the environment observation value, and performs an action according to the current control strategy; S23, the environment responds to this action, enters a new state, and returns the new state and the reward brought by this action to the deep reinforcement learning agent; S24, in the new state, the agent will continue to perform actions, and so on, the deep reinforcement learning agent and the environment interact continuously until the optimal action-value function and the optimal control strategy are obtained.

4. The hybrid electric vehicle eco-driving method based on multi-objective deep reinforcement learning according to claim 3, characterized in that, Step S21 is as follows: Determine the state and action in deep reinforcement learning, multi-objective reward function, optimal action-value function and optimal control strategy; Specifically including: the state is the current speed of the host vehicle, the current acceleration of the host vehicle, the current driving distance of the host vehicle, the current speed of the front vehicle, the current acceleration of the front vehicle, the current driving distance of the front vehicle, the current vehicle following distance, the current engine power, the state of charge SoC of the power battery, the state of health SoH of the power battery, the average temperature inside the battery and the battery capacity decay rate c; the action is the following behavior mode and the engine power; define the reward function, including low fuel consumption, SoC stability and SoH stability, the specific calculation formula of the reward function reward is: In the above formula, R(s, a) is a reward function vector for selecting action a in state s, each target giving a corresponding weight; R1(s, a) is a reward function reward 1; R2(s, a) is a reward function reward 2; R3(s, a) is a reward function reward 3; C f is the engine instantaneous fuel consumption; C b is the battery charging cost; C ag is the battery aging cost; M, V are standardization coefficients; The specific calculation formula of the optimal action-value function is: Q * (s,a) = Q π (s,a) = max E[R t+1 + λQ * (s t+1 ,a t+1 )|s t ,a t ] (2) In the above equation, Q π (s, a) is the action-value function of choosing action a in policy π state s; s t ,a t is the state and action at time t; s t+1 ,a t+1 R t+1 is the state, action and reward function at time t + 1; and λ ∈ [0, 1] is the discount factor. Optimal control policy π * satisfies The specific calculation formula is 5. The method of claim 1, wherein, Step S3 is specifically: S31, establish a conditional network based on target relative weight input; S32, define a control action selection strategy based on the output result of the conditional network; S33, use multi-experience replay to sample experience for the weight vector of the recent unexecuted strategy.

6. The method of claim 5, wherein, The conditional network based on target relative weight input in step S31 is essentially a value-based neural network, specifically including an estimation neural network and a target neural network, the target network is used to update the estimation network, and the internal structures of the two are the same, the target relative importance is input as a condition, and a reward weight sampling mechanism is combined to realize the trade-off of multiple objectives; When establishing the conditional network, there are three inputs: one is the state observation value, the second is the control amount, and the third is the target weight, wherein the state observation value includes the current speed of the host vehicle, the current acceleration of the host vehicle, the current driving distance of the host vehicle, the current speed of the front vehicle, the current acceleration of the front vehicle, the current driving distance of the front vehicle, the current vehicle following distance, the current engine power, the state of charge SoC of the power battery, the state of health SoH of the power battery, the average temperature inside the battery and the battery capacity decay rate c, and the output is a multi-objective Q value vector.

7. The method of claim 6, wherein, The definition of the control action selection strategy in step S32 specifically includes: according to the current strategy and a certain action selection probability epsilon, an execution action is obtained from this random process, which can be represented as: P(a=random choice)=epsilon In the formula, P(·) is the action selection probability.

8. A computer readable storage medium having stored thereon a computer program, characterized in that: The program is executed by the processor to realize the steps in the hybrid electric vehicle ecological driving method based on multi-objective deep reinforcement learning in any one of claims 1-7.