Driving control method, system and equipment for multi-agent deep reinforcement learning based on expert priority and storage medium

By adopting a driving control method based on expert-first multi-agent deep reinforcement learning in plug-in hybrid vehicles, the shortcomings of energy management and driving performance optimization in the prior art in complex driving scenarios are solved, and more efficient and safer driving performance is achieved.

CN120096595APending Publication Date: 2025-06-06CHANGAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510249848.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The energy management strategies and driving performance optimization of existing plug-in hybrid vehicles are underperforming in complex and dynamic driving scenarios, and traditional control methods tend to view adaptive cruise control and energy management as independent subsystems, resulting in suboptimal solutions.

Method used

The driving control method based on expert-first multi-agent deep reinforcement learning is adopted. By constructing a vehicle follow-up model and power system model, the ACC agent and EMS agent are constructed using the MADDPG algorithm, and a reward function is set for them, and offline training is combined with expert knowledge to obtain driving strategies after expert intervention.

Benefits of technology

It significantly improves the energy efficiency and driving safety and comfort of plug-in hybrid cars, enhances adaptability to different driving cycles, and achieves a significant improvement in driving performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120096595A_ABST
    Figure CN120096595A_ABST
Patent Text Reader

Abstract

The invention relates to the field of automobile ecological driving, in particular to a multi-agent deep reinforcement learning driving control method, system and device based on expert priority and a storage medium. Comprising the following steps: constructing a simulation environment, and loading training data; the method comprises the following steps: constructing two agents ACC and EMS, and constructing an Actor network, a Critic network and a target network; and training an ecological driving strategy, introducing expert knowledge and experience playback during training, generating a driving strategy, and loading network parameters of the driving strategy to a vehicle control unit to realize online application. The invention provides a training method based on expert priority, expert knowledge can be fused into a multi-agent deep reinforcement learning process, and key experience including expert intelligence can be ensured to be utilized more frequently and more effectively by constructing a priority experience playback mechanism, so that the agents can converge to a better strategy more quickly, and the learning efficiency of the multi-agent deep reinforcement learning is improved. And the driving performance is obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of automobile ecological driving, and in particular to a driving control method, system, device and storage medium based on expert-priority multi-agent deep reinforcement learning. Background Art

[0002] As a means of transportation with the dual advantages of fuel and electricity, plug-in hybrid electric vehicles (PHEVs) have become the most promising alternative to traditional fuel vehicles. However, the energy management strategy (EMS) and driving performance optimization of PHEVs still face many challenges. Traditional EMS includes rule-based and optimization-based strategies. Rule-based methods are simple and easy to use, but performance is difficult to guarantee; optimization-based strategies such as dynamic programming (DP), Pontryagin minimum principle, and equivalent consumption minimization strategy perform well in terms of fuel economy, but are computationally limited and rely on predefined driving cycles. Model predictive control combines real-time performance and robustness, but performance is affected when prediction accuracy is limited.

[0003] Currently, most EMSs are optimized for fixed driving cycles, with dynamic parameters related to the powertrain (such as vehicle speed and total required power) remaining unchanged. Therefore, the optimization potential of these EMSs is limited. Complex and dynamic driving scenarios bring new challenges to strengthening EMSs, and eco-driving provides a powerful mechanism to overcome the limitations of the driving environment on the powertrain, contributing to the development of smarter and more efficient transportation systems in the future. Among them, adaptive cruise control (ACC), as the technical basis for achieving this goal, is essential for maintaining a safe distance between vehicles and ensuring effective following. In vehicle following scenarios, eco-driving can be regarded as a combination of energy management and speed planning through hierarchical or holistic control. The current hierarchical control framework usually regards ACC and EMS as independent subsystems. This tandem control approach often leads to suboptimal solutions. ACC usually focuses on dynamic following performance, while EMS determines torque distribution based on the desired vehicle acceleration. Both systems are optimized independently according to their respective objective functions, without fully utilizing the benefits of coordinated control. Although adding specific optimization items to ACC can solve problems such as unnecessary rapid acceleration, the optimal output of ACC may not always match the acceleration required by EMS. Furthermore, due to the nonlinear relationship between the operating state and efficiency of each power source, the optimal speed cannot be determined in the driving cycle alone. This highlights the need for a more comprehensive control strategy to optimize vehicle speed and energy management simultaneously.

[0004] Eco-driving is understood as a collaboration between a pair of intelligent agents, ACC and EMS. The development of deep reinforcement learning (DRL) has brought new hope for eco-driving strategy optimization. Algorithms such as deep Q-learning, dual deep Q-learning, and deep deterministic policy gradient (DDPG) are widely used. However, deep Q-learning has limited adaptability and lacks interpretability when facing different driving cycles. Multi-agent deep reinforcement learning (MADRL) techniques such as multi-agent deep deterministic policy gradient (MADDPG) have certain potential in eco-driving, but still face problems such as slow convergence and difficulty in exploration. Summary of the invention

[0005] In response to the problems mentioned in the prior art, the present invention proposes a driving control method, system, device and storage medium based on expert-priority multi-agent deep reinforcement learning to solve the above-mentioned problems existing in the prior art, improve the energy efficiency, driving safety and comfort of plug-in hybrid vehicles, and enhance the adaptability to different driving cycles.

[0006] In order to achieve the above object, the present invention adopts the following technical solution: The present invention provides a driving control method based on expert-priority multi-agent deep reinforcement learning, comprising the following steps: S1. Based on the ACC driving scenarios of the pilot vehicle and the main vehicle, a vehicle following model is constructed, and a vehicle power system model is constructed based on the structure of the main vehicle; S2. According to the vehicle following model and the vehicle power system model, the ACC agent and the EMS agent are constructed according to the MADDPG algorithm, and reward functions are set for the ACC agent and the EMS agent respectively; S3, establish an expert model based on MADRL according to the expert knowledge data set; perform offline training on the ACC agent and EMS intelligence in S2, introduce the expert model during the offline training process, and train the driving strategy after expert intervention; S4: Load the driving strategy into the vehicle controller of the vehicle, and the target domain vehicle executes the training to complete the driving strategy.

[0007] As a further improvement of the present invention, the specific process of S1 is: The pilot vehicle follows a predetermined speed curve. The host vehicle obtains information about the pilot vehicle based on the onboard sensors. The host vehicle navigates based on its own acceleration. A vehicle following model is established based on the acceleration data. The model definition is as follows:

[0008]

[0009]

[0010]

[0011] Where: v t For the main vehicle t Time speed; a t For the main vehicle t Time acceleration; L t For the main vehicle t Distance travelled at the time; v t+1 is the speed of the main vehicle at the next moment; L t+1 is the driving distance of the main vehicle at the next moment; Δ T is the sampling time interval; v h is the speed of the main vehicle; a h is the acceleration of the main vehicle; v max is the maximum speed of the main vehicle; a min is the minimum acceleration of the host vehicle; a max is the maximum acceleration of the main vehicle; D min is the minimum following distance between the pilot vehicle and the main vehicle; d 0 The static safe vehicle spacing is 2 to 3 meters. τ brk is the response time of the braking system, which is 0.5S; a brk is the braking deceleration, take 6~8 m / s 2 ; D max The maximum following distance between the pilot vehicle and the main vehicle; Jerk is the rate of change of acceleration; Taking hybrid electric vehicle as the main vehicle, a vehicle power system model is established according to energy consumption. The vehicle power system model includes a dual planetary gear system model, an engine consumption model, an electric motor consumption model, a generator consumption model and a battery model. The double planetary gear system model is defined as follows:

[0012]

[0013] Where: ω d The speed required by the main vehicle; T d The torque required by the main vehicle; k1 is the transmission ratio of the sun gear; k 2 is the transmission ratio of the planetary gear; i is the final transmission ratio of a single gear; Engine consumption model , Motor consumption model , Generator consumption model The definition is as follows:

[0014] Where: ω e The speed of the main vehicle engine; T e is the torque of the main vehicle engine; ω m The speed of the main vehicle motor; T m is the torque of the main vehicle motor; ω g The speed of the main vehicle generator; T g The torque of the main vehicle generator; f is the relationship function between efficiency and torque and speed respectively; The battery model is defined as follows:

[0015]

[0016]

[0017]

[0018] Where: SOC is the state of charge; U OCV is the open circuit voltage; U is the terminal voltage; I is the load current; R s is the ohmic internal resistance; C n is the nominal capacity of the battery; R D C D and R T C T Simulate the charge diffusion and transfer effects respectively, corresponding to the polarization voltage U D and U T .

[0019] As a further improvement of the present invention, the specific process of S2 is: Based on the vehicle following model and the vehicle power system model, the ACC agent corresponding to the vehicle following model and the EMS agent corresponding to the vehicle power system model are constructed through the MADDPG algorithm. Considering the speed, velocity, distance between the two vehicles and battery information of the ACC agent and the EMS agent, the state space expression is defined as follows:

[0020] Where: s is the state function; D h,L is the distance between the pilot vehicle and the main vehicle; v L is the speed of the pilot car; a L is the acceleration of the pilot car; P req The vehicle power requirement of the main vehicle; and are the speed and acceleration of the host vehicle respectively; I is the current, SoC is the state of charge of the battery; The action space expression is defined as follows:

[0021] Where: is the control action of the intelligent agent ACC, i.e., the acceleration of the main vehicle; is the control action of the EMS agent, i.e., the engine power.

[0022] As a further improvement of the present invention, the reward function expression of the ACC agent is as follows:

[0023]

[0024]

[0025] Where: r 1 , r 2 is the sub-reward function; α 1 , α 2 is the reward weight; d is the following distance between the host vehicle and the pilot vehicle; D safe is the safe following distance; The reward function of the EMS agent is defined as:

[0026] Where: β 1 , β 2 are the corresponding reward weights; m f is the fuel consumption; SOC tar is the SOC target value; It is the square of the difference between the current SOC and the SOC target value.

[0027] As a further improvement of the present invention, the offline training process of the ACC agent and the EMS agent in S3 is: Initialize the Actor network, Critic network and corresponding target network of the ACC agent and EMS agent, define and initialize a storage space as the experience replay pool. The expression of the experience replay pool is as follows:

[0028] Where: s t for t Observation values ​​of all agents in the environment at all times; a , a , …, a is a set of actions for all agents at the current moment; r , r , …, r Rewards for all agents after executing their current actions; s t+1 is the state of all agents at the next moment; During execution, actions can be obtained a :

[0029] Where: A policy network representing all agents; Critic Network The loss function :

[0030]

[0031] Where: represents the target network; represents the target value network; Actor Network θ i, The loss function :

[0032] Calculate the gradient of the updated Actor network:

[0033] Use the soft update method to update the target network parameters of the Actor and Critic networks:

[0034]

[0035] Where τ is the soft update factor; Training to get demonstration strategy ρπ E .

[0036] As a further improvement of the present invention, the process of establishing the expert strategy model based on MADRL according to the expert knowledge data set in S3 is: Expert strategy model established π H , the goal is to find an expert strategy , making it compatible with the presentation strategy ρπ E The difference is minimized, and its expression is as follows:

[0037] Obtain an expert dataset, which includes the intelligent driver model and the states and actions derived from dynamic programming, and input the expert dataset into the expert strategy model π H The trained expert model is obtained by training. The loss function of the expert model is expressed as follows:

[0038] The expert model copies the demonstration policy, so when the expert intervenes, the agent outputs the following expression for the action:

[0039] Where: a represents the output action of the expert intervention agent; represents the identity matrix associated with the action space dimension; Δ t is a Boolean variable, set to True when expert intervention occurs, otherwise set to False; The expert intervention step is added to the experience replay pool to form a new experience pool replay pool. The new experience pool replay pool expression is as follows:

[0040] During the learning process, the loss parameter of the Critic network is updated as follows:

[0041] Where: N 1 represents N transformation tuples from the demonstration strategy; N 2 represents N transition tuples from the expert strategy; γ is the decay rate; subscript Indicates that at time step t When the original reinforcement learning policy N 1 The first conversion l The variable corresponding to the experience; Indicates that at time step When from N 2 Expert Demonstration Conversion m sub-empirical variables; The loss parameter update of the Actor network is as follows:

[0042] Where: is with status s l,t and Actor Network Output The relevant Q value function; ω is the weight coefficient.

[0043] As a further improvement of the present invention, the difference between the Q value function of the expert strategy and the Q value function of the demonstration strategy is used as the priority in the priority experience playback, and the formula is expressed as follows:

[0044] Where: is the timing error; ε Is a positive constant.

[0045] A driving control system based on expert-first multi-agent deep reinforcement learning, including: The construction module is used to build a vehicle following model based on the ACC driving scenarios of the pilot vehicle and the host vehicle, and to build a vehicle power system model based on the structure of the host vehicle; A setting module is used to construct an ACC agent and an EMS agent according to the vehicle following model and the vehicle power system model according to the MADDPG algorithm, and set reward functions for the ACC agent and the EMS agent respectively; The strategy generation module is used to establish an expert strategy model based on MADRL according to the expert knowledge data set; the ACC agent and EMS intelligence in the setting module are trained offline, and the expert strategy model is introduced during the offline training process to train the driving strategy after expert intervention.

[0046] A driving control device based on expert-priority multi-agent deep reinforcement learning comprises a processor and a memory, wherein when the processor executes a computer program stored in the memory, the driving control method based on expert-priority multi-agent deep reinforcement learning as described above is implemented.

[0047] A computer-readable storage medium for storing a computer program, wherein when the computer program is executed by a processor, the driving control method based on expert-priority multi-agent deep reinforcement learning as described above is implemented.

[0048] Compared with the prior art, the present invention has achieved the following technical effects: The method of the present invention not only effectively solves the complexity problem of ecological driving control, but also significantly enhances the adaptability to various driving cycles; the present invention is a training method based on expert priority, which can integrate expert knowledge into the multi-agent deep reinforcement learning (MADRL) process, and by constructing a prioritized experience replay mechanism, ensures that key experiences containing expert wisdom can be used more frequently and more effectively; this mechanism not only accelerates the learning process, but also significantly enhances the influence of expert guidance in the training process, so that the intelligent agent can converge to a better strategy more quickly and achieve a significant improvement in driving performance.

[0049] The method of the present invention combines the expert priority guidance mechanism with the MPC framework, which not only brings out the strengths of MPC in dealing with multivariable constrained optimization problems, but also makes full use of the advantages of expert knowledge in guiding intelligent agent learning. The entire framework not only optimizes the training process, improves the learning efficiency and optimization performance of the intelligent agent, but also provides a new solution for the design and implementation of ecological driving strategies. BRIEF DESCRIPTION OF THE DRAWINGS Figure 1 It is a schematic diagram of the overall framework of the present invention; Figure 2 The ACC vehicle following model diagram of the present invention; Figure 3 A structural diagram of a plug-in hybrid electric vehicle of the present invention; Figure 4 is the distance trajectory curve of the main vehicle; Figure 5 is the acceleration change rate curve of the main vehicle; Figure 6 The SOC trajectory curve of the main vehicle; Figure 7 The distribution statistics of the engine operating points of the main vehicle; Figure 8 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION

[0050] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for ease of description, only parts related to the present invention, rather than all structures, are shown in the accompanying drawings.

[0051] like Figure 1 as well as Figure 8 As shown, the driving control method of the present invention based on expert-priority multi-agent deep reinforcement learning specifically includes the following steps: S1. Based on the ACC driving scenarios of the pilot vehicle and the main vehicle, a vehicle following model is constructed, and a vehicle power system model is constructed based on the structure of the main vehicle; S2. According to the vehicle following model and the vehicle power system model, the ACC agent and the EMS agent are constructed according to the MADDPG algorithm, and reward functions are set for the ACC agent and the EMS agent respectively; S3, establish an expert model based on MADRL according to the expert knowledge data set; perform offline training on the ACC agent and EMS intelligence in S2, introduce the expert model during the offline training process, and train the driving strategy after expert intervention; S4: Load the driving strategy into the vehicle controller of the vehicle, and the target domain vehicle executes the training to complete the driving strategy.

[0052] The present invention will be further explained below with reference to the accompanying drawings and specific embodiments: like Figure 2 As shown in the figure, step 1: construct a vehicle following model. t The state at the moment can be expressed as:

[0053]

[0054]

[0055] Where: v t For the main vehicle t Time speed; at For the main vehicle t Time acceleration; L t For the main vehicle t Distance travelled at the time; v t+1 is the speed of the main vehicle at the next moment; L t+1 is the driving distance of the main vehicle at the next moment; Δ T is the sampling time interval; v h is the speed of the main vehicle; a h is the acceleration of the main vehicle; v max is the maximum speed of the main vehicle; a min is the minimum acceleration of the host vehicle; a max The maximum acceleration of the main vehicle.

[0056] The pilot vehicle follows a predetermined speed profile, the host vehicle uses onboard sensors to obtain information about the pilot vehicle, and the host vehicle relies on its own acceleration signal as the only external input to navigate within the system. To ensure the reliability and effectiveness of the model, the maximum following distance varies according to the driver's behavior and is usually affected by the vehicle speed. Minimum following distance D min and the ideal maximum following distance D max The definition is as follows:

[0057]

[0058] Where: D min is the minimum following distance between the pilot vehicle and the main vehicle; d 0 The static safe vehicle spacing is 2 to 3 meters. τ brk is the response time of the braking system, which is 0.5S; a brk is the braking deceleration, take 6~8 m / s 2 ; D max It is the maximum following distance between the pilot vehicle and the main vehicle.

[0059] Jerk It represents the rate of change of acceleration and is a key indicator for evaluating driving comfort. The calculation formula is as follows:

[0060] Where: Jerkis the rate of change of acceleration.

[0061] Step 2: Figure 3 As shown, a hybrid vehicle is used as the main vehicle to construct a vehicle power system model, and a vehicle power system model is established, including a dual planetary gear system model, a diesel engine model, a drive motor model, a generator model and a battery model. Among them, the dual planetary gear system model is as follows:

[0062]

[0063] Where: ω d The speed required by the main vehicle; T d The torque required by the main vehicle; k 1 is the transmission ratio of the sun gear; k 2 is the transmission ratio of the planetary gear; i is the final transmission ratio of a single gear.

[0064] Engine consumption model , Motor consumption model , the generator consumption model is defined as follows:

[0065] Where: ω e The speed of the main vehicle engine; T e is the torque of the main vehicle engine; ω m The speed of the main vehicle motor; T m is the torque of the main vehicle motor; ω g The speed of the main vehicle generator; T g The torque of the main vehicle generator; f It is the relationship function between efficiency and torque and speed.

[0066] The battery model is as follows:

[0067]

[0068]

[0069]

[0070] Where: SOC is the state of charge; U OCV is the open circuit voltage; U is the terminal voltage; I is the load current; R s is the ohmic internal resistance; C n is the nominal capacity of the battery; R D C D and R T C T Simulate the charge diffusion and transfer effects respectively, corresponding to the polarization voltage U D and U T ; Step 2: Based on the MADDPG algorithm, construct the ACC agent corresponding to the vehicle following model and the EMS agent corresponding to the vehicle power system model.

[0071] For the ACC agent, information from the pilot vehicle and the host vehicle is used to evaluate vehicle dynamics, while the distance between vehicles is critical for evaluating following distance and safety. For the EMS agent, battery indicators such as current and SOC represent the health of the battery pack, while the vehicle power demand provides insight into the operating state. Therefore, the state space expression is defined as follows:

[0072] Where: s is the state function; D h,L is the distance between the pilot vehicle and the main vehicle; v L is the speed of the pilot car; a L is the acceleration of the pilot car; P req The vehicle power requirement of the main vehicle; and are the speed and acceleration of the host vehicle respectively; I is the current, and SoC is the state of charge of the battery.

[0073] Action Space: As mentioned above, the main control input of the ACC agent is the acceleration of the host vehicle. For the EMS agent, due to the planetary gear structure, any combination of speed and torque within the hardware constraints can be provided, thereby achieving the appropriate distribution of motor and generator speed and torque. The engine output power is then adjusted to a specific speed-torque pair based on the fuel consumption rate. Therefore, the engine power is specified as the action variable of the EMS agent, and the complete action space can be expressed as:

[0074] Where: is the control action of the intelligent agent ACC, i.e., the acceleration of the main vehicle; is the control action of the EMS agent, i.e., the engine power.

[0075] For the ACC agent, the policy goal is to keep the main vehicle at a safe distance and comfortable following behavior, so the reward function of the ACC agent is defined as:

[0076]

[0077]

[0078] Where: r 1 , r 2 is the sub-reward function; α 1 , α 2 is the reward weight. When the distance is outside this range, a cost term is applied, and if the distance exceeds the maximum allowed distance, a very high penalty is imposed; d is the following distance between the host vehicle and the pilot vehicle; D safe It is the safe following distance. To ensure driving safety, the driving distance should not be lower than the minimum safety limit.

[0079] The reward function of the EMS agent is defined as:

[0080] Where: β 1 , β 2 are the corresponding reward weights; m f is the fuel consumption; SOC tar is the SOC target value.

[0081] At present, the most representative algorithm in the field of MADRL is the MADDPG algorithm, which is an improvement of the DDPG algorithm for multi-agent scenarios. Each agent follows the structure of the DDPG algorithm, which solves the problem of environmental instability in multi-agent systems. The MADDPG algorithm uses offline mode for centralized training and online mode for distributed execution, allowing each critic network to access the state and action information of other agents, enabling it to evaluate the value of the current output action from the actor network, and each agent only relies on observed local information during task execution.

[0082] Therefore, a storage space is defined and initialized as the experience replay pool. The expression of the experience replay pool is as follows:

[0083] Where: s t for t Observation values ​​of all agents in the environment at all times; a , a , …, a is a set of actions for all agents at the current moment; r , r , …, r Rewards for all agents after executing their current actions; s t+1 is the state of all agents at the next moment.

[0084] Each agent maintains a network of actors μ θi and critic network θ i The basic structure of the target actor network μ and target critic network Q The critic network is trained in a centralized training approach:

[0085]

[0086] During distributed execution, a deterministic action can be directly obtained based on its local observation vector a :

[0087] Actor Network The loss function :

[0088] Compute the gradient of the updated actor network:

[0089] Where: ▽ represents the gradient function.

[0090] Use the soft update method to update the target network parameters of the Actor and Critic networks:

[0091]

[0092] Where τ is the soft update factor.

[0093] Repeat the training process until the end of the training, and output the demonstration strategy ρπ E .

[0094] Step 3: Adopt a novel expert-first-based training method to integrate expert knowledge into MADRL. The expert strategy model is expressed as π H , the goal is to find a strategy ρπ H , making it compatible with the presentation strategy ρπ E The difference d is minimized:

[0095] Obtain an expert dataset, which includes the intelligent driver model and the states and actions derived from dynamic programming, and input the expert dataset into the expert strategy model π H The trained expert model is obtained by training. The loss function of the expert model is expressed as follows:

[0096] The model will gradually develop the ability to accurately replicate the demonstrated policy, thus becoming an expert in supporting training. When the supervising expert decides to intervene, full control is granted, and then the actual action a output by the agent is completely determined by the expert's decision:

[0097] Where: a represents the output action of the expert intervention agent; represents the identity matrix associated with the action space dimension; Δ t is a Boolean variable that is set to True when expert intervention occurs and False otherwise.

[0098] The expert intervention step is added to the experience replay pool to form a new experience pool replay pool. The new experience pool replay pool expression is as follows:

[0099] During the learning phase, N transformation tuples are randomly drawn from the experience replay pool according to the batch size, where N 1 Part of it consists of the transformation of the MADDPG strategy from step 2, and the rest N 2 The conversion comes from the expert strategy. During the learning process of the evaluation network, the loss function of each agent is modified as follows:

[0100] Where: N 1 represents N transformation tuples from the demonstration strategy; N 2 represents N transition tuples from the expert strategy; γ is the decay rate; subscript Indicates that at time step t When the original reinforcement learning policy N 1 The first conversion l The variable corresponding to the experience; Indicates that at time step When from N 2 Expert Demonstration Conversion m sub-empirical variables.

[0101] The loss parameter update of the Actor network is as follows:

[0102] Where: is with status s l,t and Actor Network Output The relevant Q value function; ω is the weight coefficient.

[0103] Expert demonstrations are often more meaningful than most MADRL behavior policies due to the presence of prior knowledge and reasoning ability. Therefore, an effective way to prioritize expert demonstrations in the experience buffer is necessary. A new priority-based metric is proposed to replace the traditional TD error in the standard Prioritized Experience Replay (PER). Since the value function can evaluate the policy performance, the difference between the Q-value of the expert demonstration and the Q-value of the RL behavior can be calculated. For a given tuple of expert demonstration transitions, this difference acts as the advantage score of the PER, and the new priority is defined as:

[0104] Where: is the timing error; ε Is a positive constant.

[0105] In the offline training process of the embodiment, the agent randomly selects two driving cycles from the training dataset and seamlessly integrates them at the beginning and end. This curve is used as the training driving cycle curve for MADDPG to learn a near-optimal ecological driving strategy until convergence, and each cycle lasts 3600 seconds. The test cycle is used to evaluate the adaptability of the proposed strategy, spanning 3200s.

[0106] Keeping the training scenario parameters and algorithm hyperparameters unchanged, the PEiL-MADDPG (MADDPG strategy based on expert priority), EiL-MADDPG (MADDPG strategy based on expert guidance), MADDPG, DP and other algorithms of the present invention are used as comparison strategies, and the superiority of the proposed algorithm is verified by comparing the performance; the comfort index in the evaluation index is defined as the acceleration change rate of the vehicle, and the economy index is defined as the total driving cost. The simulation results are as follows: Figure 4 , Figure 5 , Figure 6 as well as Figure 7 shown.

[0107] The results show that if Figure 4 As shown in the figure, all strategies show acceptable tracking performance, effectively ensuring vehicle tracking and preventing collisions. In addition, for the variance of the absolute value of acceleration, the EiL-MADDPG strategy reduces it by 12.85% compared with MADDPG, while the PEiL-MADDPG strategy of the present invention further reduces the variance by 5.18% compared with the EiL-MADDPG strategy. Figure 5 The results of the jerk value are shown. For a given MADDPG-based algorithm, the |jerk| values ​​of all strategies are relatively low. Compared with the baseline MADDPG, PeiL-MADDPG shows better vehicle following performance and better driving comfort, highlighting the advantages of expert priority guidance and MPC framework in maintaining close following and providing comfortable driving.

[0108] Figure 6 It shows that all strategies keep the terminal SOC within the preferred operating margin of about 0.20 with a deviation of about 5%. Although the SOC trajectories generated by the MADRL method are different from those of DP, the MADRL-based trajectory shows greater variability, while the DP trajectory is closer to the target value, indicating that DP is able to find the global optimal solution based on complete future driving information. The SOC trajectories of DP and the PeiL-MADDPG of the present invention show smaller SOC changes and smoother fluctuations, which helps to extend battery life.

[0109] Figure 7The distribution of engine operating points is shown, showing that MADDPG exhibits a wider presence in the low- and medium-efficiency regions. However, under the guidance of experts, the engine operating points gradually shift to higher efficiency regions. Among these strategies, PEiL-MADDPG achieves the highest proportion of engine points in the high-efficiency range and the least engine points in the low BSFC region compared to MADDPG and EiL-MADDPG. These findings indicate that the present PEiL-MADDPG achieves excellent optimization in terms of engine operating efficiency, highlighting the advantages of prioritized expert guidance and the MPC framework. Based on the same inventive concept, an embodiment of the present invention also provides a driving control system based on expert-priority multi-agent deep reinforcement learning. Since the principle of solving the problem by the driving control system based on expert-priority multi-agent deep reinforcement learning is similar to the aforementioned driving control method based on expert-priority multi-agent deep reinforcement learning, the implementation of the driving control system based on expert-priority multi-agent deep reinforcement learning can refer to the implementation of the driving control method based on expert-priority multi-agent deep reinforcement learning, and the repeated parts will not be repeated.

[0110] In specific implementation, the driving control system based on expert-priority multi-agent deep reinforcement learning provided by the embodiment of the present invention specifically includes: The construction module is used to build a vehicle following model based on the ACC driving scenarios of the pilot vehicle and the host vehicle, and to build a vehicle power system model based on the structure of the host vehicle; A setting module is used to construct an ACC agent and an EMS agent according to the vehicle following model and the vehicle power system model according to the MADDPG algorithm, and set reward functions for the ACC agent and the EMS agent respectively; The strategy generation module is used to establish an expert strategy model based on MADRL according to the expert knowledge data set; the ACC agent and EMS intelligence in the setting module are trained offline, and the expert strategy model is introduced during the offline training process to train the driving strategy after expert intervention.

[0111] Correspondingly, an embodiment of the present invention also provides a driving control device based on expert-priority multi-agent deep reinforcement learning, comprising a processor and a memory, wherein when the processor executes the computer program stored in the memory, it implements the driving control method based on expert-priority multi-agent deep reinforcement learning as provided in the embodiment of the present invention.

[0112] For more specific processes of the above method, please refer to the corresponding contents disclosed in the aforementioned embodiments, which will not be repeated here.

[0113] Accordingly, an embodiment of the present invention also provides a computer-readable storage medium for storing a computer program, wherein, when the computer program is executed by a processor, the driving control method based on expert-priority multi-agent deep reinforcement learning as provided in the embodiment of the present invention is implemented.

[0114] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the systems, devices, and storage media disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method part.

[0115] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0116] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0117] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0118] The above is a detailed introduction to the driving control method, system, device and storage medium based on expert-priority multi-agent deep reinforcement learning provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for general technical personnel in this field, according to the idea of ​​the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. A driving control method based on expert-first multi-agent deep reinforcement learning, characterized in that: The following steps are involved: S1. Based on the ACC driving scenarios of the pilot vehicle and the main vehicle, a vehicle following model is constructed, and a vehicle power system model is constructed based on the structure of the main vehicle; S2. According to the vehicle following model and the vehicle power system model, the ACC agent and the EMS agent are constructed according to the MADDPG algorithm, and reward functions are set for the ACC agent and the EMS agent respectively; S3, establish an expert model based on MADRL according to the expert knowledge dataset; The ACC agent and EMS agent in S2 are trained offline. The expert model is introduced during the offline training process to train the driving strategy after expert intervention. S4: Load the driving strategy into the vehicle controller of the vehicle, and the target domain vehicle executes the training to complete the driving strategy.

2. The driving control method based on expert-priority multi-agent deep reinforcement learning according to claim 1, characterized in that: The specific process of S1 is: The pilot vehicle follows a predetermined speed curve. The host vehicle obtains information about the pilot vehicle based on the onboard sensors. The host vehicle navigates based on its own acceleration. A vehicle following model is established based on the acceleration data. The model definition is as follows: Where: v t For the main vehicle t Time speed; a t For the main vehicle t Time acceleration; L t For the main vehicle t Distance travelled at the time; v t+1 is the speed of the main vehicle at the next moment; L t+1 is the driving distance of the main vehicle at the next moment; Δ T is the sampling time interval; v h is the speed of the main vehicle; a h is the acceleration of the main vehicle; v max is the maximum speed of the main vehicle; a min is the minimum acceleration of the host vehicle; a max is the maximum acceleration of the main vehicle; D min is the minimum following distance between the pilot vehicle and the main vehicle; d 0 is the static safe vehicle spacing, which is 2 to 3 meters; τ brk is the response time of the braking system, which is 0.5S; a brk is the braking deceleration, take 6~8 m / s 2 ; D max The maximum following distance between the pilot vehicle and the main vehicle; Jerk is the rate of change of acceleration; Taking hybrid electric vehicle as the main vehicle, a vehicle power system model is established according to energy consumption. The vehicle power system model includes a dual planetary gear system model, an engine consumption model, an electric motor consumption model, a generator consumption model and a battery model. The double planetary gear system model is defined as follows: Where: ω d The speed required by the main vehicle; T d The torque required by the main vehicle; k 1 is the transmission ratio of the sun gear; k 2 is the transmission ratio of the planetary gear; i is the final transmission ratio of a single gear; Engine consumption model , Motor consumption model , Generator consumption model The definition is as follows: Where: ω e The speed of the main vehicle engine; T e is the torque of the main vehicle engine; ω m The speed of the main vehicle motor; T m is the torque of the main vehicle motor; ω g The speed of the main vehicle generator; T g The torque of the main vehicle generator; f is the relationship function between efficiency and torque and speed respectively; The battery model is defined as follows: Where: SOC is the state of charge; U OCV is the open circuit voltage; U is the terminal voltage; I is the load current; R s is the ohmic internal resistance; C n is the nominal capacity of the battery; R D C D and R T C T Simulate the charge diffusion and transfer effects respectively, corresponding to the polarization voltage U D and U T .

3. The driving control method based on expert-priority multi-agent deep reinforcement learning according to claim 1, characterized in that: The specific process of S2 is: Based on the vehicle following model and the vehicle power system model, the ACC agent corresponding to the vehicle following model and the EMS agent corresponding to the vehicle power system model are constructed through the MADDPG algorithm. Considering the speed, velocity, distance between the two vehicles and battery information of the ACC agent and the EMS agent, the state space expression is defined as follows: Where: s is the state function; D h,L is the distance between the pilot vehicle and the main vehicle; v L is the speed of the pilot car; a L is the acceleration of the pilot car; P req The vehicle power requirement of the main vehicle; and are the speed and acceleration of the vehicle, respectively; I is the current, and SoC is the state of charge of the battery; The action space expression is defined as follows: Where: is the control action of the intelligent agent ACC, i.e., the acceleration of the main vehicle; is the control action of the EMS agent, i.e., the engine power.

4. The driving control method based on expert-priority multi-agent deep reinforcement learning according to claim 3 is characterized in that: The reward function expression of the ACC agent is as follows: Where: r 1, r 2 is the sub-reward function; α 1, α 2 is the reward weight; d is the following distance between the host vehicle and the pilot vehicle; D safe is the safe following distance; The reward function of the EMS agent is defined as: Where: β 1, β 2 are the corresponding reward weights; m f is the fuel consumption; SOC tar is the SOC target value; It is the square of the difference between the current SOC and the SOC target value.

5. The driving control method based on expert-priority multi-agent deep reinforcement learning according to claim 1, characterized in that: The offline training process of the ACC agent and the EMS agent in S3 is: Initialize the Actor network, Critic network and corresponding target network of the ACC agent and EMS agent, define and initialize a storage space as the experience replay pool. The expression of the experience replay pool is as follows: Where: s t for t Observation values ​​of all agents in the environment at all times; a , a , …, a is a set of actions for all agents at the current moment; r , r , …, r Rewards for all agents after executing their current actions; s t+1 is the state of all agents at the next moment; During execution, actions can be obtained a : Where: A policy network representing all agents; Critic Network The loss function : Where: represents the target network; represents the target value network; Actor Network θ i, The loss function : Calculate the gradient of the updated Actor network: Use the soft update method to update the target network parameters of the Actor and Critic networks: Where τ is the soft update factor; Training Demonstration Strategy ρπ E .

6. The driving control method based on expert-priority multi-agent deep reinforcement learning according to claim 1, characterized in that: The process of S3 establishing an expert strategy model based on MADRL according to the expert knowledge data set is: Expert strategy model established π H , the goal is to find an expert strategy , making it compatible with the presentation strategy ρπ E The difference is minimized, and its expression is as follows: Obtain an expert dataset, which includes the intelligent driver model and the states and actions derived from dynamic programming, and input the expert dataset into the expert strategy model π H The trained expert model is obtained by training. The loss function of the expert model is expressed as follows: The expert model copies the demonstration policy, so when the expert intervenes, the agent outputs the following expression for the action: Where: a represents the output action of the expert intervention agent; represents the identity matrix associated with the action space dimension; Δ t is a Boolean variable, set to True when expert intervention occurs, otherwise set to False; The expert intervention step is added to the experience replay pool to form a new experience pool replay pool. The new experience pool replay pool expression is as follows: During the learning process, the loss parameter of the Critic network is updated as follows: Where: N 1 represents N transformation tuples from the demonstration strategy; N 2 represents N transition tuples from the expert strategy; γ is the decay rate; subscript Indicates that at time step t When the original reinforcement learning policy N 1st conversion l The variable corresponding to the experience; Indicates that at time step When from N 2nd expert demonstration conversion m sub-empirical variables; The loss parameter update of the Actor network is as follows: Where: is with status s l,t and Actor Network Output The relevant Q value function; ω is the weight coefficient.

7. A driving control method based on expert-priority multi-agent deep reinforcement learning according to claim 6, characterized in that: The difference between the Q-value function of the expert strategy and the Q-value function of the demonstration strategy is used as the priority in the priority experience replay. The formula is expressed as follows: Where: is the timing error; ε Is a positive constant.

8. A driving control system based on expert-first multi-agent deep reinforcement learning, characterized in that: include: The construction module is used to build a vehicle following model based on the ACC driving scenarios of the pilot vehicle and the host vehicle, and to build a vehicle power system model based on the structure of the host vehicle; A setting module is used to construct an ACC agent and an EMS agent according to the vehicle following model and the vehicle power system model according to the MADDPG algorithm, and set reward functions for the ACC agent and the EMS agent respectively; The strategy generation module is used to establish an expert strategy model based on MADRL according to the expert knowledge data set; the ACC agent and EMS intelligence in the setting module are trained offline, and the expert strategy model is introduced during the offline training process to train the driving strategy after expert intervention.

9. A driving control device based on expert-first multi-agent deep reinforcement learning, characterized in that: It includes a processor and a memory, wherein when the processor executes the computer program stored in the memory, it implements the driving control method based on expert-priority multi-agent deep reinforcement learning as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: Used to store a computer program, wherein when the computer program is executed by a processor, it implements the driving control method based on expert-priority multi-agent deep reinforcement learning as described in any one of claims 1 to 7.