An Optimization Method for Automotive Adaptive Cruise Control Based on SAC

By adopting SAC deep reinforcement learning method in adaptive cruise control, combining dynamic reward function and multi-task experience classification module, the problem of difficulty in dealing with complex urban traffic scenarios in the existing technology is solved, and more efficient and safer adaptive cruise control is achieved.

CN119329519BActive Publication Date: 2025-05-27CHANGCHUN UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411768854.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2025-05-27
Estimated Expiration
2044-12-04

AI Technical Summary

Technical Problem

The existing rules-based adaptive cruise control method is difficult to cope with complex and changeable urban traffic scenarios, and the fixed proportion of reward functions and a single experience pool lead to low learning efficiency, difficulty in convergence and insufficient generalization ability.

Method used

The deep reinforcement learning method based on SAC is adopted, and the state information processing module is integrated with high and low dimensional information, and a dynamic reward function module, an experience classification module and an experience sampling module based on different driving environments are designed to improve the pertinence of learning and the convergence speed, robustness and generalization ability of the algorithm.

Benefits of technology

Adaptive cruise control under urban operating conditions is realized, the safety, driving efficiency and comfort of the car are improved, and the convergence speed and generalization ability of the algorithm are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119329519B_ABST
    Figure CN119329519B_ABST
Patent Text Reader

Abstract

A SAC-based automobile adaptive cruise control optimization method belongs to the field of automatic driving, and is characterized in that the method includes the following modules: driving environment, state information processing module, SAC reinforcement learning module, dynamic reward function module, experience classification module and experience sampling module. First, two-dimensional fusion information is obtained from the driving environment to obtain the current state. Then, the SAC reinforcement learning module decides the control action based on the current state and applies it to the driving environment, updates the environment and obtains the state at the next moment. Among them, the dynamic reward function module calculates the reward value according to the action effect and importance difference; the experience classification module stores the experience samples in different regions according to the driving environment; the experience sampling module uses fixed experience sampling and local priority experience playback methods to sample samples for training the SAC reinforcement learning module, and decides the optimal control action to achieve adaptive cruise control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical field:

[0001] The present invention belongs to the field of autonomous driving, and specifically provides a control optimization method for adaptive cruising of a smart car based on SAC deep reinforcement learning under urban conditions. Background technology:

[0002] With the rapid development of intelligent network technology, the global automotive intelligence process continues to accelerate. Intelligent network vehicles have become one of the important directions for the development of the domestic and foreign automotive industry with their potential capabilities, such as reducing safety accidents and environmental pollution, alleviating traffic congestion, and reducing energy consumption. Cruise control is an important part of the decision-making control system of autonomous vehicles. Frequent acceleration, deceleration, emergency braking and other control behaviors not only affect traffic efficiency and stability, but also affect the safety and comfort of the vehicle during driving, and even cause serious collision accidents. Therefore, it is of great significance to adopt reasonable and effective cruise control methods.

[0003] Existing rule-based adaptive cruise control methods are limited by the designer's prior knowledge and are difficult to cope with complex and changing urban traffic scenarios. Reinforcement learning algorithms can autonomously learn optimal strategies through continuous interaction with the environment, thereby solving complex decision-making and control problems. Among them, the Soft Actor-Critic algorithm provides more flexible and refined control through continuous actions, which helps to achieve smooth adjustment of vehicle speed, thereby optimizing the performance of adaptive cruise control.

[0004] However, due to the dynamic changes in the urban traffic environment and the differences in the importance of performance requirements at different time points, the use of a fixed-weight reward function will reduce learning efficiency, resulting in a long algorithm training cycle and difficulty in convergence. In addition, if the reinforcement learning network is trained based on a single experience pool, the targeted nature of the learning will be reduced, resulting in poor network adaptability and slow convergence. In terms of experience pool sampling, if the agent uses the traditional uniform random sampling method, it will lead to low experience utilization, slow algorithm convergence, and weaken the network's generalization ability.

[0005] In the field of autonomous driving decision-making based on deep reinforcement learning, some scholars have studied the use of semantic segmentation maps and low-dimensional driving features as input, the acquisition of vehicle actions as network output through generator learning, and the use of the SAC algorithm to achieve safe driving within a specified time. However, the system's cruising scene is relatively simple, the reward function cannot change dynamically with the driving scene, and the traditional random sampling method is used. The network converges slowly and has low robustness. The SAC network finally trained also has certain limitations in generalization ability. Summary of the invention:

[0006] In order to solve the above technical problems, the present invention comprehensively considers safety, traffic efficiency and comfort, and proposes a vehicle adaptive cruise control optimization method based on SAC.

[0007] The technical solution adopted by the present invention to solve the technical problem is as follows:

[0008] First, environmental information is obtained from the driving environment, and information fusion of two dimensions is performed through the state information processing module to obtain the current state; secondly, the SAC reinforcement learning module decides the control action based on the current state and applies it to the driving environment, updates the driving environment and obtains the state at the next moment, and the dynamic reward function module calculates the reward value according to the action effect; finally, the experience classification module stores the experience samples in different regions according to the driving environment, and the experience sampling module samples the experience samples for training the SAC reinforcement learning module. The trained SAC reinforcement learning module decides the optimal control action to realize the adaptive cruise control of the car. The method includes the following steps:

[0009] Step 1: Build the driving environment:

[0010] The driving environment of the present invention includes stable following, front vehicle cut-in, front vehicle cut-out and constant speed cruising. The overall training process is as follows:

[0011] The first scene in the training scenario is following the front vehicle, with an initial speed of 40km / h; then it enters the acceleration following scene, where the front vehicle quickly accelerates to 50km / h; then it enters the deceleration following scene, where the front vehicle decelerates from 50km / h to 20km / h; then it enters the front vehicle cutting out scene, where the target cruising speed of the main vehicle is 50km / h; then it enters the front vehicle cutting in scene, where the front vehicle cuts into the lane where the main vehicle is located at a speed of 30km / h; finally it enters the cruising scene, where the main vehicle should cruise at a speed of 50km / h.

[0012] Step 2: Design the status information processing module:

[0013] The present invention decomposes the autonomous driving goals in the simulation scenario into three parts: safety, traffic efficiency and comfort. The state information includes the position information of the autonomous driving vehicle itself and the surrounding environment information (such as vehicles and lane lines). Therefore, the state information data includes two items: high-dimensional image data and low-dimensional motion measurement data. Among them, the low-dimensional motion measurement data includes the difference Δs between the actual distance between the main vehicle and the front vehicle and the ideal following distance and the speed difference Δv between the main vehicle and the front vehicle. The image data comes from the semantic segmentation top view, and the image processing mainly includes two steps: cropping the image and color simplification.

[0014] Then the motion measurement data and the image feature information processed by the convolutional neural network are spliced ​​as the state input of the intelligent agent, which helps to improve the intelligent agent's ability to understand the scene.

[0015] Step 3: Design SAC reinforcement learning module:

[0016] At the current moment, the processed state is input into the Actor neural network to output the action, and then stored in the experience classification module. The experience sampling module extracts the experience, and then inputs it into the Critic network to output the evaluation value of the action taken under the state. Then, the output action is used for the vehicle to obtain the state at the next moment, and the cycle is repeated. The network parameter optimization process is as follows: the current state, action, reward and the next state will be saved in the experience pool in the form of a tuple. A portion of the experience in the experience pool will be sampled for neural network gradient descent, and the parameters of the Actor network (action strategy network) and the Critic network (action value estimation network) will be updated respectively. The process includes the following sub-steps:

[0017] Step 3.1, action space design:

[0018] Considering the requirements of the continuous action space of the adaptive cruise control task, the longitudinal continuous action space [-1, 1] is adopted, where [-1, 0] is the braking deceleration action, and [0, 1] is the acceleration action. Because the present invention studies the adaptive cruise control task under urban conditions, the maximum speed of the main vehicle is set to no more than 60km / h.

[0019] Step 3.2, neural network design:

[0020] Step 3.2.1, Actor network design:

[0021] The Actor network is an action strategy network. It needs to use the state information processing module to process the original image information, and then further perform nonlinear mapping on the input information through the fully connected network. The output layer outputs the Gaussian distribution of the mean and variance in this state, and projects the action to [-1,1] to obtain the real throttle or brake action value.

[0022] Step 3.2.2, Critic network design:

[0023] The Critic network is an action value estimation network. It uses the state information processing module to process the original image information, flattens the feature information and concatenates it with the motion measurement information to obtain a one-dimensional vector with a length of 259. Then, the fully connected layer performs nonlinear mapping on the combination of state information and action, and the output layer outputs the state action value Q(s t ,a t ).

[0024] Step 4: Reward function design:

[0025] The reward function design of the present invention consists of safety reward, driving efficiency reward and comfort reward, which are as follows:

[0026] Step 4.1, Security Reward:

[0027] When following a vehicle, the vehicle needs to maintain an ideal following distance from the vehicle in front. The specific calculation formula for the safety reward is as follows:

[0028]

[0029] Among them, d actual is the actual distance between the main vehicle and the front vehicle, d ideal is the ideal following distance, η, ω is a hyperparameter, which is set to: η = 1, which is the maximum reward value in the ideal following interval, is the reward-penalty ratio coefficient when the actual following distance is greater than the ideal following distance, and ω=0.1 is the reward-penalty ratio coefficient when the actual following distance is less than the ideal following distance.

[0030] Step 4.2, driving efficiency reward:

[0031] Driving efficiency is reflected in the rapid adjustment of the vehicle speed to the speed of the vehicle in front or the set cruising speed. The specific calculation formula is as follows:

[0032]

[0033] Among them, v actual is the current actual speed of the main vehicle, v target is the target driving speed. Since the present invention studies adaptive cruise driving under urban conditions, the vehicle speed range studied is 0-60km / h. As with the safety reward, since the sensor measurement data may have errors, a 1-meter buffer zone is set for the ideal speed area to improve the fault tolerance of the intelligent agent to learn the optimal strategy.

[0034] Step 4.3, comfort reward:

[0035] On the basis of ensuring safety, the specific calculation formula of the reward function for controlling the smooth changes of the car's movements when accelerating or decelerating and the comfort is as follows:

[0036] reward3=-ω·|a actual |+ξ

[0037] Among them, a actual represents the actual acceleration value of the main vehicle, ξ is a hyperparameter, representing the maximum value of the comfort reward, set to 0.2, and the goal is to control the acceleration change of the car to 2m / s 2 Within this range, the passengers have a better comfort experience.

[0038] Step 4.4, design dynamic reward function module:

[0039] In the adaptive cruise control task in the urban traffic environment, due to the dynamic nature of the environment, the size of each part of the reward function is dynamically adjusted according to the current state to adapt to the changing traffic environment, as follows:

[0040] 1) When the main vehicle takes an action, if the relative distance is greater than the ideal following distance, increase the weight of the driving efficiency reward.

[0041] 2) When the main vehicle takes an action, if the relative distance is less than the ideal following distance, the weight of the safety reward is increased.

[0042] 3) When the main vehicle takes an action, if the relative distance is equal to the ideal following distance, the action taken by the guiding car at this time improves the comfort of following the car, so the weight of the comfort reward is increased.

[0043] 4) When the main vehicle takes an action, if the relative distance reaches the maximum following distance, the driving efficiency reward weight is increased.

[0044] 5) When the host vehicle takes an action, if the relative distance is equal to the minimum safety distance, the weight of the safety reward is increased.

[0045] 6) When the main vehicle is in cruising state, the main task of the main vehicle is to reach the set cruising speed, and there is no following distance requirement. At this time, the weights of speed reward and comfort reward are increased.

[0046] According to the above rules, the present invention proposes a reward function that changes dynamically during training. The following are the adjustment factors in different situations:

[0047]

[0048] Among them, d max is the maximum following distance, d min is the minimum safety distance, e, f, and g are the weights of safety reward, driving efficiency reward, and comfort reward in the total reward respectively.

[0049] The sum of the three rewards multiplied by the dynamic weights gives the total reward function, as shown in the following formula:

[0050] Reward=e·reward1+f·reward2+g·reward3

[0051] Step 5: Design a multi-task based experience classification module:

[0052] The present invention has designed four driving scenarios. If all scenarios share one experience pool, then when entering a specific scenario, the extracted experience will include the experience of other scenarios, resulting in low pertinence of learning. Therefore, the present invention designs four different experience playback pools to store samples of different driving environments, namely, the stable following experience pool, the front vehicle cut-in experience pool, the front vehicle cut-out experience pool, and the cruise control experience pool. During the training process, when the vehicle enters different driving environments, experience is extracted from the corresponding experience pool accordingly to conduct targeted training and improve its performance and generalization ability.

[0053] Step 6: Design the experience sampling module:

[0054] The traditional SAC algorithm extracts experience from the experience pool uniformly and randomly. After each action performed by the agent to obtain reward feedback from the environment, the experience is stored in the experience pool in the form of a tuple (s, a, r, s'). The experience sampling module includes two methods: fixed experience sampling and local priority experience playback.

[0055] Step 6.1, sampling fixed experience:

[0056] In the traditional uniform sampling of the SAC algorithm, some experiences in the experience pool may not be used at all, resulting in low utilization of the experience. The present invention makes an improvement when extracting experience, and the most recently generated experience needs to be fixedly extracted from the experience pool each time.

[0057] Step 6.2: Local priority experience playback:

[0058] In order to solve the problem of insufficient use of the importance of experience in the random uniform sampling method and take into account a certain efficiency, the present invention proposes a local priority experience playback method based on uniform sampling. First, 2M experiences are extracted from the experience pool by uniform random extraction, M represents the number of random experiences required for each gradient descent calculation, and these experiences are temporarily stored in the temporary experience pool, and then these experiences are sorted according to the total reward value of each experience, and the sorted experiences are divided into three parts in order, the first to M / 2 parts are experiences with high learning effects, the M / 2 to M parts are experiences with medium learning effects, and the M to 2M parts are experiences with low learning effects. Finally, experiences are extracted from these three parts in proportion, the first M / 4 experiences are extracted from the high learning effect part, the first M / 4 experiences are extracted from the medium learning effect part, and the last M / 2 experiences are extracted from the low learning effect part. In this way, the total number of experiences required for each gradient descent calculation is M+1.

[0059] The beneficial effects of the present invention are:

[0060] The present invention proposes a SAC-based automobile adaptive cruise control optimization method, which is characterized in that, firstly, in combination with the control task requirements, a state information processing module that integrates high and low dimensional information is constructed to improve the intelligent agent's ability to understand the scene; secondly, a dynamic reward function module is constructed to improve the safety, driving efficiency and comfort of the automobile during driving; finally, in combination with the specific control task, an experience classification module and an experience sampling module based on different driving environments are proposed, which increase the pertinence of learning, improve the convergence speed, robustness and generalization ability of the algorithm, and the trained SAC reinforcement learning module decides the optimal control action to realize the adaptive cruise control of the automobile. Description of the drawings:

[0061] Figure 1 It is the overall framework diagram.

[0062] Figure 2 It is a network structure diagram.

[0063] Figure 3 This is the flow chart of the multi-task based experience classification module.

[0064] Figure 4 It is the flow chart of the experience sampling module. Specific implementation method:

[0065] In step 1, build the driving environment:

[0066] The driving environment of the present invention includes stable following, front vehicle cut-in, front vehicle cut-out and constant speed cruising. The overall training process is as follows:

[0067] The first training scenario is following the front car, with an initial speed of 40km / h. Then the acceleration following scenario is entered, with the front car accelerating rapidly to 50km / h. Then the deceleration following scenario is entered, with the front car decelerating from 50km / h to 20km / h. Then the front car cuts out, and the target cruising speed of the main car should be 50km / h. Then the front car cuts in, with the front car cutting into the lane where the main car is at a speed of 30km / h. Finally, the cruising scenario is entered, and the main car should cruise at a speed of 50km / h. The entire experimental training scenario includes four main adaptive cruise driving environments: following, cutting in, cutting out, and cruising, which can meet the adaptive cruise driving scenarios under general urban conditions.

[0068] In step 2, design the status information processing module:

[0069] The present invention decomposes the autonomous driving goals in the simulation scenario into three parts: safety, traffic efficiency and comfort. During the adaptive cruising process, the car will be affected by the relative distance and relative speed from the vehicle in front. The addition of spatial feature information helps the car intelligent body to better understand the scene. The state information includes the position information of the autonomous driving vehicle itself and the surrounding environment information (such as vehicles and lane lines). Therefore, the state information data includes two items: high-dimensional image data and low-dimensional motion measurement data. Among them, the low-dimensional motion measurement data includes the difference Δs between the actual distance between the main vehicle and the front vehicle and the ideal following distance and the speed difference Δv between the main vehicle and the front vehicle. The image data comes from the semantic segmentation top view, and the processing of the image mainly includes two steps: cropping the image and color simplification.

[0070] The default image size of Carla semantic segmentation camera is 487×487, and after image cropping, the image size becomes 120×120. The specific process of color processing is as follows: first, the perspective of the semantic segmentation camera is adjusted to a top-down perspective, and the horizontal position of the camera is adjusted to obtain a semantic segmentation bird's-eye view with the main vehicle position as the center of the image. After color simplification, only four different colors remain in the image, namely, lane lines, roads, cars, and other objects.

[0071] Then the motion measurement data and the image feature information processed by the convolutional neural network are spliced ​​as the state input of the intelligent agent, which helps to improve the intelligent agent's ability to understand the scene.

[0072] In step 3, design the SAC reinforcement learning module:

[0073] Figure 1 This is the overall framework diagram of the present invention. After the state information processing module, the high- and low-dimensional information obtained from the environment is spliced ​​as the state input. At the same time, the current state is input into the Actor neural network output action and stored in the experience classification module. After the experience sampling module extracts experience, it is input into the Critic network to output the evaluation value of the action taken under the state; then, the output action is used for the vehicle to obtain the state at the next moment, and the cycle is iterated. The current state, action, reward and next moment state are saved in the form of tuples in the experience pool. The loss function is calculated by extracting the experience in the experience pool, and the parameters of the Actor network (action strategy network) and the Critic network (action value estimation network) are updated respectively. The process includes the following sub-steps:

[0074] In step 3.1, action space design:

[0075] Considering the requirements of the continuous action space of the adaptive cruise control task, the longitudinal continuous action space [-1, 1] is adopted, where [-1, 0] is the braking deceleration action, and [0, 1] is the acceleration action. Because the present invention studies the adaptive cruise control under urban conditions, the maximum speed of the main vehicle is set to 0-60km / h.

[0076] In step 3.2, neural network design:

[0077] In step 3.2.1, Actor network design:

[0078] Figure 2 It is the network structure diagram of the present invention, and the Actor action strategy network is designed. The action network consists of three layers of fully connected neural networks, and the ReLU function is selected as the activation function; the third layer of the network structure is two [128, 1] fully connected layer networks, one for outputting the mean μ of the Gaussian distribution, and the other for outputting the standard deviation std of the Gaussian distribution. When obtaining the action, the Gaussian distribution with a mean of μ and a standard deviation of std is first sampled, and then the sampling result is activated using a hyperbolic tangent (tanh) function, and the action is projected to [-1, +1], thereby obtaining the real throttle or brake action value.

[0079] In step 3.2.2, Critic network design:

[0080] The Critic network is an action evaluation network that can output the evaluation of whether the action taken by the agent in the current state is good or bad. The input of the action evaluation network is high-dimensional image information, low-dimensional motion information and the current action a t Combined input. The distance difference Δs and the speed difference Δv in the state space are both single elements. After the image data is processed by three layers of convolutional layers, the extracted feature information is flattened into a one-dimensional vector. Then, the splicing function Cat in the neural network library is used to splice the two types of information (image information after feature extraction and motion measurement information) to obtain a one-dimensional vector of length 259. Then, the fully connected layer performs nonlinear mapping on the combination of state information and action, and the output layer outputs the state action value Q (s t ,a t ).

[0081] In step 4, design the reward function:

[0082] The reward function design of the present invention consists of safety reward, driving efficiency reward and comfort reward, which are as follows:

[0083] In step 4.1, design security rewards:

[0084] When following a vehicle, the vehicle needs to maintain an ideal following distance from the vehicle in front. The specific calculation formula for the safety reward is as follows:

[0085]

[0086] Among them, d actual is the actual distance between the main vehicle and the front vehicle, d ideal is the ideal following distance, η, ω is a hyperparameter, which is set to: η = 1, which is the maximum reward value in the ideal following interval, is the reward-penalty ratio coefficient when the actual following distance is greater than the ideal following distance, and ω=0.1 is the reward-penalty ratio coefficient when the actual following distance is less than the ideal following distance.

[0087] The adaptive cruise control system has two types of following distances: fixed headway and variable headway. Since the present invention studies the adaptive cruise control under urban conditions, it does not involve high-speed following scenarios, and the vehicle speed range is 0-60km / h. Considering the complexity of urban conditions and the requirements for real-time performance, the fixed headway of the variable safety strategy is selected as the calculation formula for the ideal following distance of the car in the scenario of stable following the front vehicle, which is defined as follows:

[0088]

[0089] Among them, τ h is the headway, D is the relative longitudinal distance between the controlled vehicle and the following vehicle, and v is the speed of the controlled vehicle.

[0090] The ideal following distance calculation formula of the vehicle of the present invention in the scenario of stably following the vehicle ahead is as follows:

[0091] d ideal =τ h v+d 0

[0092] Among them, d ideal is the ideal following distance, τ h is the set headway, d 0 is the minimum safe distance between the main vehicle and the front vehicle when following the front vehicle. In the stable following vehicle scenario, τ h The headway time is set to 1.5 seconds, and the minimum safe distance d 0 Set to 10 meters.

[0093] In step 4.2, design the driving efficiency reward:

[0094] Under the premise of ensuring the safety of car driving, improving driving efficiency is directly reflected in quickly adjusting the speed to the speed of the vehicle in front or the set cruising speed, otherwise there will be a penalty. The driving efficiency calculation formula is as follows:

[0095]

[0096] Among them, v actual is the current actual speed of the main vehicle, v target is the target driving speed. Since the present invention studies adaptive cruise driving under urban conditions, the vehicle speed range studied is 0-60km / h. As with the safety reward, due to the measurement data error of the sensor, a 1-meter buffer zone is set for the ideal speed area to improve the fault tolerance of the intelligent body to learn the optimal strategy.

[0097] In step 4.3, design the comfort bonus:

[0098] On the basis of ensuring safety, the car's movements must be controlled to be smooth when accelerating or decelerating. The specific calculation formula of the comfort reward function is as follows:

[0099] reward3=-ω·|a actual |+ξ

[0100] Among them, a actual represents the actual acceleration value of the main vehicle, ξ is a hyperparameter, representing the maximum value of the comfort reward, set to 0.2, and the goal is to control the acceleration change of the car to 2m / s 2 Within this range, the passengers have a better comfort experience.

[0101] In step 4.4, design the dynamic reward function module:

[0102] In the adaptive cruise control task in the urban traffic environment, due to the dynamic change of the environment, compared with the traditional reward function method with fixed weights, this paper proposes a dynamically changing reward function method. In this method, the size of each part of the reward function is dynamically adjusted according to the current state to adapt to the changing traffic environment, as follows:

[0103] 1) When the main vehicle takes a maneuver, if the relative distance between the two vehicles is greater than the ideal following distance, the following distance is too large, which will affect the traffic efficiency of urban roads, so the weight of the driving efficiency reward is increased.

[0104] 2) When the main vehicle takes a maneuver, if the relative distance between the two vehicles is less than the ideal following distance, further reduction of the following distance will cause safety hazards, so the weight of the safety reward is increased.

[0105] 3) When the main vehicle takes a maneuver, if the relative distance between the two vehicles is equal to the ideal following distance, the action taken by the guiding car at this time is to maintain smooth driving as much as possible to improve the comfort of following the vehicle, so the weight of the comfort reward is increased.

[0106] 4) When the main vehicle takes a maneuver, if the relative distance between the two vehicles reaches the maximum critical following distance, the main vehicle is about to be left behind, resulting in the end of the round. At this time, the weight of the driving efficiency reward is increased.

[0107] 5) When the main vehicle takes a maneuver, if the relative distance between the two vehicles is equal to the minimum safe distance, a collision will occur if the following vehicle distance is further reduced, resulting in the end of the round. At this time, the weight of the safety reward is increased.

[0108] 6) When the main vehicle is in cruising state, the main task of the main vehicle is to follow the speed and reach the set cruising speed. There is no following distance requirement. At this time, the weights of speed reward and comfort reward are increased.

[0109] According to the above rules, a reward function that changes dynamically during training is proposed to adapt to the ever-changing traffic environment. The following are the adjustment factors of the reward function:

[0110]

[0111] Among them, d max is the maximum following distance, d min is the minimum safety distance, e, f, and g are the weights of safety reward, driving efficiency reward, and comfort reward in the total reward respectively.

[0112] The sum of the three rewards multiplied by the dynamic weights gives the total reward function, as shown in the following formula:

[0113] Reward=e·reward1+f·reward2+g·reward3

[0114] In step 5, design a multi-task based experience classification module:

[0115] Figure 3 This is a flow chart of the experience classification module based on multi-tasks. If all scenes share one experience pool, then when entering a specific scene, the extracted experience will include the experience of other scenes, resulting in low pertinence of learning. Therefore, the present invention designs four different experience playback pools to store samples of different driving environments, namely, the stable following experience pool, the front car cut-in experience pool, the front car cut-out experience pool, and the cruise control experience pool. During the training process, when the vehicle enters different driving environments, experience is extracted from the corresponding experience pool to conduct targeted training and improve its performance and generalization ability.

[0116] In step 6, design the experience sampling module:

[0117] The traditional SAC algorithm extracts experience uniformly and randomly from the experience pool. The experience sampling module design includes sampling fixed experience and local priority experience playback method, which are as follows:

[0118] In step 6.1, sample fixed experience:

[0119] In the traditional uniform sampling of the SAC algorithm, some experiences in the experience pool may not be used at all, resulting in low utilization of the experience. The present invention makes an improvement when extracting experience, and the most recently generated experience needs to be fixedly extracted from the experience pool each time.

[0120] In step 6.2, local priority experience playback:

[0121] In order to solve the problem of insufficient use of the importance of experience in random uniform sampling and to take into account a certain degree of efficiency, the present invention proposes a local priority experience playback method based on uniform sampling. Figure 4 This is the flow chart of the experience sampling module. First, 2M experiences are extracted from the experience pool by uniform random sampling. M represents the number of random experiences required for each gradient descent calculation. These experiences are temporarily stored in the temporary experience pool. Then, these experiences are sorted according to the total reward value of each experience. The sorted experiences are divided into three parts in order. The 1st to M / 2 parts are experiences with high learning effects, the M / 2 to M parts are experiences with medium learning effects, and the M to 2M parts are experiences with low learning effects. Finally, experiences are extracted from these three parts in proportion. The first M / 4 experiences are extracted from the high learning effect part, the first M / 4 experiences are extracted from the medium learning effect part, and the last M / 2 experiences are extracted from the low learning effect part. In this way, the total number of experiences required for each gradient descent calculation is M+1.

[0122] Step 7: Parameter setting and iterative optimization training:

[0123] Step 7.1, SAC parameter setting:

[0124] The Actor network learning rate is set to 0.0001, the Critic network learning rate is set to 0.00001, the discount factor γ is 0.95, the soft update coefficient τ is 0.001, the experience pool size is set to 25000, the batch sampling size is 256, and the simulation time step (in seconds) is set to 0.1.

[0125] Step 7.2, iterative optimization training:

[0126] At the current moment, the state information is first obtained from the driving environment, and the low-dimensional motion information and high-dimensional image information are spliced ​​through the state information processing module to output the state s t, stored in the experience classification module, and then the SAC reinforcement learning module decides the control action a according to the current state. t And execute it into the environment, the driving environment is updated accordingly, and the state s at the next moment t+1 is obtained, and the dynamic reward function module calculates the reward value r according to the effect of the action execution; depending on the task type, the above state s t 、Action a t , reward r and next moment state s t+1 As experience samples, they are saved in different regions through the experience classification module; finally, the experience sampling module performs experience sampling according to different task types, and the obtained experience is used to train the SAC reinforcement learning module. In the process of interaction between the agent and the environment, the smaller one of the target Q values ​​output by the Critic network participates in the gradient update of the Critic network, and the smaller one of the actual Q values ​​output by the Actor network participates in the gradient update of the Actor network. The Actor network parameter update is obtained by KL divergence and is related to Q(s t ,a t ) is related to entropy. The specific iterative optimization training steps are as follows:

[0127] Step 7.2.1, Initialization. First, initialize the experience replay pools D1, D2, D3, and D4, which correspond to the stable following experience pool, the front car cut-in experience pool, the front car cut-out experience pool, and the cruise control experience pool, with a capacity of N. Initialize the reality Q 1 , Q 2 Network, randomly generate weights θ 1 ,θ 2 ; Initialize target Q - 1 , Q - 2 network, copy the parameters of the real network to the target network, and the weights are θ - 1 ,θ - 2 ; Initialize the Actor network parameters, the weight is

[0128] Step 7.2.2, loop through each round from 1 to M+1, where M+1 is the total number of experiences required for each gradient descent calculation;

[0129] Step 7.2.3: Initialize state s in each round 1 ;

[0130] Step 7.2.4: Get the current state s from the environment t , input to the strategy function of the Actor network and select s t Action in state a tIn the Actor network, the regularization coefficient α can maximize the expected reward and entropy of each strategy output, making the strategy selection more random;

[0131] Step 7.2.5: The main vehicle performs action a t , interact with the Carla environment and receive an immediate reward r t and the new state s t+1 ;

[0132] Step 7.2.6, the sample (s t ,a t ,r t ,s t+1 ) According to different driving scenarios, they are stored in the corresponding experience replay pool as a data set for training the neural network;

[0133] Step 7.2.7, when entering different task scenarios, extract the experience of the corresponding scenarios through the experience classification module and the experience sampling module for targeted learning and training;

[0134] Step 7.2.8: Optimize and update the parameters of the two real Q networks using the neural network stochastic gradient descent method. The update formula is as follows:

[0135]

[0136]

[0137]

[0138] Among them, λ Q is the learning rate, θ is the parameter of the actual Q network, TD_error is the temporal difference error, The Q network performs stochastic gradient descent on θ, L Q (θ) is the loss function of the Q network, γ is the discount factor, logπ(a t+1 |s t+1 ) is the log probability of the strategy, which measures the certainty of the strategy, is the timing error randomly sampled from the experience replay buffer, TD_error new is the timing error corresponding to the latest sampling data;

[0139] Step 7.2.9, update the Actor network parameters. The loss function of strategy π is obtained by KL divergence. The loss function simplified by reparameterization technique is:

[0140]

[0141] in, is the parameter of the Actor strategy network, α is the regularization coefficient, which is used to control the importance of entropy, ∈ t is a noise random variable, The reparameterization technique is used to make the policy function differentiable after sampling from the policy Gaussian distribution;

[0142] Step 7.2.10: Automatically adjust the entropy regularization term, that is, maximize the expected return while constraining the mean entropy to be greater than After simplifying through mathematical techniques, the loss function of α is obtained as follows:

[0143]

[0144] in, It is a man-made target entropy value used to constrain the mean value of entropy. The value is set to -1 in the present invention.

[0145] Step 7.2.11, update the target Q network parameters soft update, that is, copy the parameters of the actual Q network to the target network, the update formula is as follows,

[0146] θ - 1 ←τθ 1 +(1-τ)θ - 1

[0147] θ - 2 ←τθ 2 +(1-τ)θ - 2

[0148] Among them, θ and θ - are the actual and target network parameters respectively, τ is the soft update coefficient;

[0149] Step 7.2.12, repeat the training until the algorithm converges;

[0150] Step 7.3: After the algorithm converges through iterative training, the parameters in the network are saved, and the optimized parameters are loaded into the neural network controller. By collecting the environmental status under the current driving environment, the neural network controller is used to output control actions in real time to achieve adaptive cruise control of the car.

[0151] In summary: the present invention proposes a SAC-based automobile adaptive cruise control optimization method. Combined with the control task requirements, a state information processing module that integrates high- and low-dimensional information is constructed to improve the intelligent agent's ability to understand the scene; secondly, a dynamic reward function module is constructed to improve the safety, driving efficiency and comfort of the car during driving; finally, combined with specific control tasks, an experience classification module and an experience sampling module based on different driving environments are proposed, which increases the pertinence of learning, improves the convergence speed, robustness and generalization ability of the algorithm, and the trained SAC reinforcement learning module decides the optimal control action, produces positive effects, and has good application prospects in the field of intelligent vehicle adaptive cruise control. The experimental results show that in scenes such as cruise control, stable following, and front vehicle cut-in and cut-out, the intelligent car can achieve fast speed tracking while ensuring a safe following distance, and keep the acceleration change within the ideal range. The following distance and speed errors of the car are within 1m and 1m / s respectively, and the acceleration peak does not exceed 3m / s 2 .

Claims

1. A SAC-based automobile adaptive cruise control optimization method, belonging to the field of automatic driving, characterized in that: The method includes the following modules: driving environment, state information processing module, SAC reinforcement learning module, dynamic reward function module, experience classification module and experience sampling module; First, the fusion information of two dimensions is obtained from the driving environment to obtain the current state. Then, the SAC reinforcement learning module decides the control action based on the current state and applies it to the driving environment, updates the environment and obtains the state at the next moment. The dynamic reward function module includes safety reward, driving efficiency reward and comfort reward; among which, the safety reward means that the vehicle needs to maintain an ideal following distance with the vehicle in front. The specific calculation formula is as follows: Among them, d actual is the actual distance between the main vehicle and the front vehicle, d ideal is the ideal following distance, η, ω is a hyperparameter, η is the maximum reward value in the ideal following interval, is the reward-penalty ratio coefficient when the actual following distance is greater than the ideal following distance, and ω is the reward-penalty ratio coefficient when the actual following distance is less than the ideal following distance; The driving efficiency reward is used to quickly adjust the vehicle speed to the speed of the vehicle in front or the set cruising speed. The specific calculation formula is as follows: Among them, v actual is the current actual speed of the main vehicle, v target is the target driving speed; The comfort reward refers to the smoothness of the action change when controlling the car to accelerate or decelerate. The specific calculation formula is as follows: reward3 = -ω·|a actual |+ξ Among them, a actual represents the actual acceleration value of the main vehicle, ξ is a hyperparameter, which is the maximum value of the comfort reward; The size of each part of the reward function is dynamically adjusted according to the current state to adapt to the changing traffic environment. The following are the adjustment factors of the reward function: The sum of the three rewards multiplied by the dynamic weights gives the total reward function, where d max is the maximum following distance, d min is the minimum safety distance, e, f, and g are the weights of safety reward, driving efficiency reward, and comfort reward in the total reward respectively; The experience classification module stores the experience samples in different regions according to the driving environment; the experience sampling module samples the samples using fixed experience sampling and local priority experience playback methods for training the SAC reinforcement learning module and deciding the optimal control action to achieve adaptive cruise control.

2. The SAC-based vehicle adaptive cruise control optimization method according to claim 1, characterized in that: The multi-task-based experience classification module is designed with four different experience replay pools, which are used to store samples in different driving environments. The driving environments include stable following, front vehicle cutting in, front vehicle cutting out, and cruise control. When the vehicle enters different driving environments, experience is extracted from the corresponding experience pool for targeted training.

3. The SAC-based vehicle adaptive cruise control optimization method according to claim 1, characterized in that: The experience sampling module includes two methods: fixed experience sampling and local priority experience replay; wherein, fixed experience sampling means that the latest generated experience is fixedly extracted from the experience pool each time; the local priority experience replay method means first extracting 2M experiences from the experience pool by uniform random sampling, M represents the number of random experiences required for each gradient descent calculation, temporarily storing these experiences in a temporary experience pool, and then sorting these experiences according to the total reward value of each experience, and dividing the sorted experiences into three parts in order, the 1st to M / 2 parts are experiences with high learning effects, the M / 2 to M parts are experiences with medium learning effects, and the M to 2M parts are experiences with low learning effects, and finally extracting experiences from these three parts in proportion, extracting the first M / 4 experiences from the high learning effect part, the first M / 4 experiences from the medium learning effect part, and the last M / 2 experiences from the low learning effect part, so that the total number of experiences required for each gradient descent calculation is M+1.

Citation Information

Patent Citations

  • Automatic driving decision-making system based on deep reinforcement learning

    CN117962926A