Learning type emergency collision avoidance control method and system for automatic driving commercial vehicle and storage medium
By adopting Bayesian strategy network and attenuation expert experience technology in autonomous driving commercial vehicles, the problem of difficult learning of the logic relationship of control action output under emergency collision avoidance conditions is solved, the learning efficiency and strategy quality are improved, and efficient emergency collision avoidance control is achieved.
Patent Information
- Application Number
- CN202510222937.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-02-27
AI Technical Summary
Autonomous driving commercial vehicles face complex needs for scene perception and motion control under emergency collision avoidance conditions. In the prior art, the logical relationship of control action output is difficult to learn and the rewards are sparse, resulting in low training efficiency and poor strategic performance.
Using Bayesian policy network and attenuation expert experience guidance technology, we use reinforcement learning network models, including policy network, value network and Q-value network, and trained in combination with data from the replay buffer to generate joint control strategies that adapt to emergency collision avoidance conditions.
It improves the learning efficiency of emergency collision avoidance control for commercial vehicles for autonomous driving, ensures the optimality and robustness of the final driving strategy, solves the problem of sparse rewards, does not require complex system modeling, and the algorithm architecture is streamlined.
Smart Images

Figure CN119953360A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent transportation technology, and in particular to a learning-based emergency collision avoidance control method, system and storage medium for an autonomous commercial vehicle. Background Art
[0002] With the rapid development of artificial intelligence and automotive technology, autonomous vehicles (smart vehicles) are changing the control methods of traditional vehicles, freeing drivers from the "vehicle-road-person" closed-loop control mode, thereby improving the efficiency and safety of the transportation system.
[0003] At present, commercial vehicle autonomous driving technology is the most feasible for deployment due to its economic cost and the particularity of application scenarios. However, due to the large size, large mass, and large blind spots of driving vision of commercial vehicles, the demand for emergency collision avoidance conditions is much higher than that of other vehicles. In addition, since commercial vehicles are more difficult to control than ordinary vehicles, there are higher requirements for vehicle scene perception and motion control in emergency collision avoidance conditions.
[0004] Among the existing technologies, the methods for solving the upper-level action decision-making problem of autonomous commercial vehicles are mainly divided into modular autonomous driving technology and learning autonomous driving technology. Modular autonomous driving technology divides autonomous driving problems into modules such as perception, decision-making, planning, and control. Commonly used planning and control algorithms such as dynamic A* and MPC have a large amount of calculation, resulting in a long decision-making control time and occupying a large computing power burden of the on-board domain controller. They also have high requirements for vehicle modeling, and the algorithm is not very applicable. Learning autonomous driving technology directly maps the input information of on-board sensors to the output of vehicle control actions through a neural network model, such as imitation learning, deep reinforcement learning and other algorithms. This type of method completes self-learning based on data or exploration, without the need for manual rule-making.
[0005] The invention with publication number CN119150652A discloses a safe driving behavior planning method inspired by defensive driving experience, including S1: initializing the traffic environment and setting the vehicle state based on the Gym library; S2: designing the state space; S3: designing the action space; S4: designing the reward function; S5: training the algorithm until convergence. The invention improves the autonomous driving function by applying soft safety rules such as "defensive driving" and the "road sense" and "driving intuition" of human drivers to the current autonomous driving, increases the ability to plan the motion of intelligent vehicles, and enables the autonomous driving system to respond to different emergency road conditions more intelligently. The invention with publication number CN116176572A discloses a vehicle emergency collision avoidance control method based on DQN deep reinforcement learning, which can solve the problem that the current AEB system cannot avoid collision in an emergency situation where the distance to the obstacle is less than the total longitudinal braking distance, and can only make the same braking action when facing different obstacles only by relying on sensor information. The T-type emergency collision avoidance control method, system, medium and equipment for autonomous driving vehicles with publication number CN113460090B can maximize the collision avoidance potential of autonomous driving vehicles and improve the performance of autonomous driving vehicles in high-speed emergency avoidance and extreme driving conditions. However, in the aforementioned deep reinforcement learning-based methods, when faced with complex control tasks, the logical relationship of the control action output is rarely considered, resulting in problems such as slow training efficiency and poor strategy performance. Summary of the invention
[0006] The purpose of the present invention is to provide a learning-based emergency collision avoidance control method, system and storage medium for an autonomous commercial vehicle, so that the autonomous commercial vehicle can learn the emergency collision avoidance control strategy through exploration in training and efficiently complete the emergency collision avoidance task in the application scenario.
[0007] In order to achieve the above technical objectives, the technical solution adopted by the present invention is:
[0008] In a first aspect, the present invention discloses a learning-based emergency collision avoidance control method for an autonomous commercial vehicle, the method comprising the following steps:
[0009] Step 1: Collect driving data of human drivers, build an expert experience dataset, and add the expert experience dataset to the replay buffer in the form of a decay ratio;
[0010] Step 2: construct a reinforcement learning network model based on the soft actor-critic architecture, wherein the reinforcement learning network model includes a policy network, a value network, and a Q-value network. The policy network generates continuous driving actions based on the observation data of the sensor, and the value network and the Q-value network evaluate the driving actions generated by the policy network. The policy network, the value network, and the Q-value network are trained by using the observation-action pairs in the replay buffer to update the corresponding network parameters.
[0011] Step 3: Based on the reinforcement learning network model built in step 2, a corresponding Bayesian strategy network is constructed for commercial vehicle driving tasks and control actions. The strategy network structure and network update process in step 2 are expanded to generate a joint control strategy that adapts to emergency collision avoidance conditions. The node sets of the Bayesian strategy network are steering δ k 、Throttle a k and brake k A three-action strategy;
[0012] Step 4: Apply the driving actions generated by the Bayesian policy network to the vehicle's driving task, and through interaction with the driving environment, with the goal of collision avoidance without collision at high speeds, perform model training under the guidance of the reward function to generate the optimal strategy.
[0013] Furthermore, in step 1, the collected expert experience data set is arranged into a replay buffer with an ε-ratio priority, so as to help the autonomous commercial vehicle to map the state to the action for initial exploration based on the demonstration provided by the expert; wherein ε is set as a gradually decaying proportional coefficient, which gradually decays to zero as the training progresses.
[0014] Furthermore, in step 2, the reinforcement learning network model includes a policy network, a value network and two Q-value networks. The policy network designed according to the dimension of state information is used to generate continuous driving actions through the observation values of the current sensor input; the state space of the reinforcement learning network model is s=[θ k ,d k ,d axis ,v k ,w k ,n k ,f c ,op k ], where θ k is the angle between the vehicle direction and the center of the road, d k is the distance between the vehicle and the edge of the road within 200 meters returned by the sensor, d axis is the normalized distance between the vehicle and the centerline of the road; v k It contains three speed scalars, namely the vehicle speed along the travel direction, the speed transverse to the travel direction axis, and the speed along the Z axis of the vehicle; w k Contains n scalars, which are the rotation speeds of the n wheels of the vehicle, n k is the engine speed per minute, f c is a binary flag indicating that the vehicle has collided with other traffic participants, op k Store the information of the vehicle closest to the vehicle, including the relative distance and angle; the action space is a = [δ k ,ak ,br k ], where δ k is the turning action, the range is [-1, +1], a k For acceleration action, the range is [0,1], br k It is the braking action, and the range is [0,1].
[0015] Furthermore, in step 2, the strategy network, value network and Q-value network are all constructed by a 5-layer neural network, which are a state input layer, two hidden layers containing 256 neurons, a normalization layer and a driving action output layer;
[0016] The observation information obtained by the sensor is normalized and input into the state input layer in sequence. Feature extraction and processing are performed through two hidden layers and activated through the ReLu function to obtain the tensor corresponding to the observation information. In the policy network, the tensor is activated through the linear function of the driving action output layer to obtain the final action space. In the value network and the Q-value network, the tensor is activated through the linear function of the driving action output layer to obtain the final action evaluation.
[0017] Furthermore, in step 2, the learning rate of the policy network is 0.0003, the learning rate of the Q value network is 0.003, the influence ratio of future action rewards is 0.96, the soft update parameter of the target network is 0.005, the playback buffer size is 1000000, and the sampling batch is 256.
[0018] Furthermore, in step 3, the Bayesian strategy network outputs the driving action strategy T = (t i ) i∈I It can be described as the acceleration control strategy t a , Steering control strategy δ And the braking control strategy br Three sub-strategy action probabilities π T The joint probability function is composed of the following relations:
[0019] π T (t δ ,t a ,t br )=π 1 (t a )π 2 (t δ |t a )π 3 (t br |t a ).
[0020] Furthermore, in step 2, the joint control strategy adopted by the Bayesian strategy network is:
[0021]
[0022] Where A is the joint action distribution of strategy π, is each sub-strategy π i The action space, θ i are the corresponding parameters;
[0023] The value network uses the value function V(A t ,s) Evaluate the performance of the joint control strategy:
[0024]
[0025] Among them, π(A t |s;θ) is the action A taken in state s t The probability of q(s,A t ; ω) is the action A taken in state s t The expected return of θ and ω are the corresponding network parameters.
[0026] The gradient descent of the corresponding parameters is calculated by the following formula:
[0027]
[0028] The joint control strategy uses the objective function J v (π) Calculate the value to maximize the sub-strategy π i The reward r(s t ,A t ) and entropy sum:
[0029]
[0030] Here, α is a hyperparameter used to control the weight of the entropy regularization term.
[0031] In the evaluation process of soft policy iteration, the soft Q value is iteratively calculated starting from an arbitrary function Q. and repeatedly apply the modified Bellman backup operator The soft value function of the updated sub-strategy, the Bellman backup operator is given by:
[0032]
[0033] In the improvement step of each sub-strategy, the corresponding strategy is updated according to the following formula for the state at each moment:
[0034]
[0035] in, is the old policy action value function, is the normalization factor.
[0036] In each sub-strategy iteration, soft policy evaluation and soft policy improvement are performed alternately to ensure the optimal maximum entropy convergence between sub-strategy combinations.
[0037] Furthermore, in step 4, the reward function is:
[0038]
[0039] in, is the driving efficiency index, It is an indicator of driving safety. is the driving comfort index, It is the completion index of the emergency collision avoidance task;
[0040]
[0041]
[0042] In the formula, v k is the vehicle speed, v limit With v max are the maximum speed limit of the road and the maximum speed that the vehicle can reach, respectively; ttc is the expected time of collision at the current speed, d lat is the lateral distance between vehicles, f c is a binary indicator of whether the vehicle has collided, v x is the vehicle's current speed along the axis of travel. ttc , lat , k c are all hyperparameters; thr and ρ ste are the hyperparameters of the longitudinal acceleration term and the lateral acceleration term respectively; is the Euclidean distance between the autonomous commercial vehicle and the background vehicle, c r is a hyperparameter.
[0043] The second invention, the present invention discloses a learning-based emergency collision avoidance control system for an autonomous commercial vehicle according to the method described above, characterized in that the system comprises:
[0044] A data set construction module is used to collect driving data of human drivers, construct an expert experience data set, and add the expert experience data set to the replay buffer in the form of a decay ratio;
[0045] A reinforcement learning network model construction module is used to construct a reinforcement learning network model based on a soft actor-critic architecture, wherein the reinforcement learning network model includes a policy network, a value network, and a Q-value network, wherein the policy network generates continuous driving actions based on sensor observation data, and the value network and the Q-value network evaluate the driving actions generated by the policy network; the policy network, the value network, and the Q-value network are trained by using observation-action pairs in a replay buffer to update corresponding network parameters;
[0046] The joint control strategy module is used to build a corresponding Bayesian strategy network for commercial vehicle driving tasks and control actions based on the built reinforcement learning network model, expand the strategy network structure and network update process, and generate a joint control strategy that adapts to emergency collision avoidance conditions; the node sets of the Bayesian strategy network are steering δ k 、Throttle a k and brake k A three-action strategy;
[0047] The optimal strategy output module is used to apply the driving actions generated by the Bayesian strategy network to the vehicle's driving tasks, and through interaction with the driving environment, with the goal of collision avoidance without collision at high speeds, conduct model training under the guidance of the reward function to generate the optimal strategy.
[0048] In a third aspect, the present invention discloses a storage medium, wherein a computer program is stored in the storage medium. When the computer program is executed in a computer, the computer is caused to execute the emergency collision avoidance control method as described above.
[0049] Compared with the prior art, the present invention has the following beneficial effects:
[0050] First, the learning emergency collision avoidance control method, system and storage medium for autonomous driving commercial vehicles of the present invention, by constructing a Bayesian strategy network to model the interactive relationship of vehicle control instructions, and designing a decaying expert experience re-entry buffer to guide the learning of the intelligent agent (autonomous driving commercial vehicle), solves the problem that the logical relationship of the control action output of the existing learning autonomous driving method is difficult to learn and the rewards are sparse in complex scenarios, improves the learning efficiency of reinforcement learning and ensures the optimality and robustness of the final driving strategy.
[0051] Second, the learning emergency collision avoidance control method, system and storage medium for autonomous commercial vehicles of the present invention utilize Bayesian strategy networks and attenuated expert experience guidance technology to achieve learning efficiency and driving strategy improvements for emergency collision avoidance control of autonomous commercial vehicles, which is beneficial for solving long-tail safety issues of autonomous commercial vehicles and promoting the application of autonomous commercial vehicles. It does not require complex system modeling and has a streamlined algorithm architecture. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 is a schematic diagram of the deep reinforcement learning network architecture of the present invention;
[0053] Figure 2 It is a schematic diagram of the driving strategy decomposition method and the probability distribution of the joint strategy of the present invention;
[0054] Figure 3 It is a Bayesian strategy network structure diagram of the soft actor critic algorithm of the present invention. DETAILED DESCRIPTION
[0055] The embodiments of the present invention are further described in detail below with reference to the accompanying drawings.
[0056] The present invention discloses a learning-based emergency collision avoidance control method for an autonomous commercial vehicle, the method comprising the following steps:
[0057] Step 1: Collect driving data of human drivers, build an expert experience dataset, and add the expert experience dataset to the replay buffer in the form of a decay ratio;
[0058] Step 2: construct a reinforcement learning network model based on the soft actor-critic architecture, wherein the reinforcement learning network model includes a policy network, a value network, and a Q-value network. The policy network generates continuous driving actions based on the observation data of the sensor, and the value network and the Q-value network evaluate the driving actions generated by the policy network. The policy network, the value network, and the Q-value network are trained by using the observation-action pairs in the replay buffer to update the corresponding network parameters.
[0059] Step 3: Based on the reinforcement learning network model built in step 2, a corresponding Bayesian strategy network is constructed for commercial vehicle driving tasks and control actions. The strategy network structure and network update process in step 2 are expanded to generate a joint control strategy that adapts to emergency collision avoidance conditions. The node sets of the Bayesian strategy network are steering δ k 、Throttle a k and brake k A three-action strategy;
[0060] Step 4: Apply the driving actions generated by the Bayesian policy network to the vehicle's driving task, and through interaction with the driving environment, with the goal of collision avoidance without collision at high speeds, perform model training under the guidance of the reward function to generate the optimal strategy.
[0061] The emergency collision avoidance control method of the present invention specifically comprises the following steps:
[0062] Step 1: Collect driving data from human drivers and build an expert experience dataset.
[0063] Specifically, a driving simulator and simulation platform are used to collect the sensor perception status information of the autonomous driving vehicle at each moment in high-speed emergency collision avoidance situations and the corresponding human driver control action information at that moment (including steering wheel, throttle opening, and brake pedal opening) to form a human expert demonstration data set, which is stored in the form of a matrix to facilitate the use of human expert driving experience in subsequent steps.
[0064] Step 2: Based on the soft actor-critic architecture, a reinforcement learning network model is constructed. The reinforcement learning network model specifically includes a policy network, a value network, and a Q-value network. The policy network generates continuous driving actions through sensor observation data, and the value network and the Q-value network evaluate the actions. The policy network, value network, and Q-value network are trained by using observation-action pairs in the replay buffer to update the corresponding network parameters.
[0065] Specifically, the state space is s = [θ k ,d k ,d axis ,v k ,w k ,n k ,f c ,op k ], where θ k is the angle between the vehicle direction and the center of the road, d k is the distance between the vehicle and the edge of the road within 200 meters returned by the sensor, d axis is the normalized distance between the vehicle and the road centerline, v k It contains three speed scalars, namely the vehicle speed along the driving direction, the speed transverse to the driving direction axis, and the speed along the Z axis of the vehicle, w k Contains n scalars, which are the rotation speeds of the n wheels of the vehicle. The specific value is determined by the number of wheels actually used in the commercial vehicle. k is the engine's speed per minute; the action space is a=[δ k ,a k ,br k ], δ k is the turning action, the range is [-1, +1], a k For acceleration action, the range is [0,1], br k is the braking action, the range is [0,1]. c It is a binary flag indicating that the vehicle has collided with other traffic participants. op k Stores the information of the vehicle currently closest to the vehicle, including the relative distance and angle.
[0066] The reinforcement learning network model of the present invention includes a policy network, a value network and two Q-value networks. The policy network designed according to the dimension of state information is used to generate continuous driving actions through the observation values of the current sensor input. The value function is used to estimate the value function of the current driving strategy. The present invention includes two estimates, namely the value network and the Q-value network. The policy network updates the network parameters according to the value function. The unmanned commercial vehicle (i.e., the intelligent body) applies continuous driving actions to enter the next state, obtains reward feedback, and stores the state-action pair between the obtained observation value and the current action in the replay buffer. The replay buffer is sampled, and the sampling information is used to update the policy network, the value network and the Q-value network.
[0067] In this embodiment, each neural network includes 5 layers, namely, a state input layer, two hidden layers containing 256 neurons, a normalization layer, and a driving action output layer; first, the observation information obtained by the sensor is normalized and input into the state input layer in turn, and features are extracted and processed through two hidden layers, and activated by the ReLu function. In the policy network, the above tensor is activated by the linear function of the output layer to obtain the final action space. In the value network, the above tensor is activated by the linear function of the output layer to obtain the final action evaluation.
[0068] The hyperparameters of the reinforcement learning network model are set as those of the soft actor-critic algorithm, where the policy network learning rate is 0.0003, the Q network learning rate is 0.003, the future action reward influence ratio (discount factor) is 0.96, the target network soft update parameter is 0.005, the replay buffer size is 1000000, and the sampling batch is 256.
[0069] Step 3: Based on the soft actor-critic reinforcement learning architecture built in step 2, a corresponding Bayesian policy network is constructed for commercial vehicle driving tasks and control actions, and the policy network structure and network update process of the soft actor-critic reinforcement learning architecture are expanded to generate a joint control strategy that adapts to emergency collision avoidance conditions.
[0070] Specifically, the three actions that a driver can take when driving a vehicle, namely steering, accelerator and brake, are independent random variables. In order to model the process of skilled drivers selecting joint driving actions during emergency collision avoidance and improve the efficiency of agent learning, the present invention decomposes the complex strategy combination into three simple sub-strategies according to driving actions, and the three strategy networks learn their respective action outputs, and construct the three sub-strategies into a joint strategy form with a Bayesian network structure.
[0071] In the present invention, the node sets of the Bayesian network are respectively k 、Throttle a k and brakek The strategy of three actions, the specific relationship between them is as follows Figure 2 As shown, the final output driving action strategy T = (t i ) i∈I It can be described as the acceleration control strategy t a , Steering control strategy δ And the braking control strategy br Three sub-strategy action probabilities π T The joint probability function composed by this relationship is:
[0072] π T (t δ ,t a ,t br )=π 1 (t a )π 2 (t δ |t a )π 3 (t br |t a ).
[0073] The present invention uses the corresponding deep convolutional neural network to approximate the joint strategy in order to learn the optimal strategy in the DRL algorithm:
[0074]
[0075] Where A is the joint action distribution of strategy π, is each sub-strategy π i The action space, θ i are the corresponding parameters.
[0076] The value network uses the value function V(A t ,s) Evaluate the performance of a specific joint strategy A.
[0077]
[0078] Among them, π(A t |s;θ) is the action A taken in state s t The probability of q(s,A t ; ω) is the action A taken in state s t The expected return of θ and ω are the corresponding network parameters.
[0079] The gradient descent of the corresponding parameters can be calculated as follows.
[0080]
[0081] Aiming at the maximum entropy concept of soft actor critic, due to the additivity of entropy, the present invention expresses the entropy of the joint strategy system as the sum of the entropies of independent sub-strategies. V (π) will use the objective function of the following formula to calculate the value to maximize the sub-strategy π i The reward r(s t ,A t ) and entropy sum.
[0082]
[0083] Here, α is a hyperparameter used to control the weight of the entropy regularization term.
[0084] In the evaluation process of soft policy iteration, the soft Q value can be iteratively calculated starting from any function Q, that is, and repeatedly apply the modified Bellman backup operator The soft value function of the updated sub-strategy, the Bellman backup operator is given by:
[0085]
[0086] In the improvement step of each sub-strategy, the corresponding strategy is updated according to the following formula for the state at each moment:
[0087]
[0088] in, is the old policy action value function, is the normalization factor.
[0089] In each sub-strategy iteration, soft policy evaluation and soft policy improvement are performed alternately to ensure the optimal maximum entropy convergence between sub-strategy combinations. The present invention utilizes two Q networks to mitigate the positive deviation in the policy improvement step, and alternately uses the experience collected from the environment by the current policy and the stochastic gradient of the mini-batch sampled from the replay buffer to update the function approximator.
[0090] Step 4: The vehicle applies the driving action generated by the Bayesian policy network to the driving task, and interacts with the driving environment to achieve emergency collision avoidance in dangerous high-speed driving, and performs model training under the guidance of the reward function. The expert data set in step 1 is added to the replay buffer in the form of a decay ratio to optimize the model training effect.
[0091] Specifically, in using reinforcement learning algorithms to guide autonomous commercial vehicles (i.e., intelligent agents) to learn high-speed driving emergency collision avoidance tasks, the design of the reward function is crucial. In the present invention, a total reward function is set for multiple indicators such as driving efficiency, driving safety, driving comfort, and emergency collision avoidance task completion to guide the training and learning of autonomous commercial vehicles (i.e., intelligent agents).
[0092] The driving efficiency as the basic structure of the reward function is set as:
[0093]
[0094] Among them, v k is the vehicle speed, v limit With v max They are the maximum speed limit of the road and the maximum speed that the vehicle can reach.
[0095] Driving safety as a constraint reward function is set as:
[0096]
[0097] Among them, ttc is the expected collision time at the current vehicle speed, d lat is the lateral distance between vehicles, f c is a binary indicator of whether the vehicle has collided, v x is the vehicle's current speed along the axis of travel. ttc , lat , k c are all hyperparameters.
[0098] Driving comfort as a constraint reward function is set as:
[0099]
[0100] Among them, ρ thr and ρ ste are the hyperparameters of the longitudinal acceleration term and the lateral acceleration term respectively.
[0101] The incentive reward function for the completion of the emergency collision avoidance task is set as:
[0102]
[0103] in, is the Euclidean distance between the autonomous commercial vehicle and the background vehicle, c r is a hyperparameter.
[0104] On this basis, the total reward function of the present invention is set as:
[0105]
[0106] In the process of model training under the guidance of the reward function established above, in view of the sparse reward problem existing in the application scenario of the present invention, the expert experience data collected in step 1 is arranged in the replay buffer with ε ratio priority, so as to help the self-driving commercial vehicle (i.e., the intelligent agent) to map the state to the action for initial exploration based on the demonstration provided by the expert. ε is set as a gradually decaying proportional coefficient, which gradually decays to zero as the training progresses.
[0107] By adopting the reinforcement learning algorithm architecture provided in the above steps, a driving strategy for an autonomous commercial vehicle that can achieve high-speed emergency collision avoidance tasks can be obtained by conducting model training on a computer and screening the optimal strategy.
[0108] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of complete hardware embodiments, complete software embodiments, or embodiments in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The scheme in the embodiments of the present application can be implemented in various computer languages, for example, object-oriented programming language Java and literal scripting language JavaScript, etc.
[0109] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0110] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0111] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0112] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0113] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A learning-based emergency collision avoidance control method for an autonomous commercial vehicle, characterized in that: The method comprises the following steps: Step 1: Collect driving data of human drivers, build an expert experience dataset, and add the expert experience dataset to the replay buffer in the form of a decay ratio; Step 2: construct a reinforcement learning network model based on the soft actor-critic architecture, wherein the reinforcement learning network model includes a policy network, a value network, and a Q-value network. The policy network generates continuous driving actions based on the observation data of the sensor, and the value network and the Q-value network evaluate the driving actions generated by the policy network. The policy network, the value network, and the Q-value network are trained by using the observation-action pairs in the replay buffer to update the corresponding network parameters. Step 3: Based on the reinforcement learning network model built in step 2, a corresponding Bayesian strategy network is constructed for commercial vehicle driving tasks and control actions. The strategy network structure and network update process in step 2 are expanded to generate a joint control strategy that adapts to emergency collision avoidance conditions. The node sets of the Bayesian strategy network are steering δ k 、Throttle a k and brake k A three-action strategy; Step 4: Apply the driving actions generated by the Bayesian policy network to the vehicle's driving task, and through interaction with the driving environment, with the goal of collision avoidance without collision at high speeds, perform model training under the guidance of the reward function to generate the optimal strategy.
2. The learning-based emergency collision avoidance control method for an autonomous commercial vehicle according to claim 1, characterized in that: In step 1, the collected expert experience dataset is arranged into the replay buffer with ε-proportional priority to help the autonomous commercial vehicle to conduct initial exploration by mapping states to actions based on the demonstrations provided by the experts; wherein ε is set as a gradually decaying proportional coefficient, which gradually decays to zero as the training progresses.
3. The learning-based emergency collision avoidance control method for an autonomous commercial vehicle according to claim 1, characterized in that: In step 2, the reinforcement learning network model includes a policy network, a value network and two Q value networks. The policy network designed according to the dimension of state information is used to generate continuous driving actions through the observation values of the current sensor input; the state space of the reinforcement learning network model is s=[θ k ,d k ,d axis ,v k ,w k ,n k ,f c ,op k ], where θ k is the angle between the vehicle direction and the center of the road, d k is the distance between the vehicle and the edge of the road within 200 meters returned by the sensor, d axis is the normalized distance between the vehicle and the centerline of the road; v k It contains three speed scalars, namely the vehicle speed along the travel direction, the speed transverse to the travel direction axis, and the speed along the Z axis of the vehicle; w k Contains n scalars, which are the rotation speeds of the n wheels of the vehicle, n k is the engine speed per minute, f c is a binary flag indicating that the vehicle has collided with other traffic participants, op k Store the information of the vehicle closest to the vehicle, including the relative distance and angle; the action space is a = [δ k ,a k ,br k ], where δ k is the turning action, the range is [-1, +1], a k For acceleration action, the range is [0,1], br k It is the braking action, and the range is [0,1].
4. The learning-based emergency collision avoidance control method for an autonomous commercial vehicle according to claim 3, characterized in that: In step 2, the strategy network, value network and Q-value network are all constructed by a 5-layer neural network, which are a state input layer, two hidden layers containing 256 neurons, a normalization layer and a driving action output layer; The observation information obtained by the sensor is normalized and input into the state input layer in sequence. Feature extraction and processing are performed through two hidden layers and activated through the ReLu function to obtain the tensor corresponding to the observation information. In the policy network, the tensor is activated through the linear function of the driving action output layer to obtain the final action space. In the value network and the Q-value network, the tensor is activated through the linear function of the driving action output layer to obtain the final action evaluation.
5. The learning-based emergency collision avoidance control method for an autonomous commercial vehicle according to claim 1, characterized in that: In step 2, the learning rate of the policy network is 0.0003, the learning rate of the Q value network is 0.003, the influence ratio of future action rewards is 0.96, the soft update parameter of the target network is 0.005, the playback buffer size is 1000000, and the sampling batch is 256.
6. The learning-based emergency collision avoidance control method for an autonomous commercial vehicle according to claim 1, characterized in that: In step 3, the Bayesian policy network outputs the driving action strategy T = (t i ) i∈I Described as the acceleration control strategy t a , Steering control strategy δ And the braking control strategy br Three sub-strategy action probabilities π T The joint probability function is composed of the following relations: π T (t δ ,t a ,t br )=π1(t a )π2(t δ |t a )π3(t br |t a )。 7. The learning-based emergency collision avoidance control method for an autonomous commercial vehicle according to claim 6, characterized in that: In step 2, the joint control strategy adopted by the Bayesian strategy network is: Where A is the joint action distribution of strategy π, is each sub-strategy π i The action space, θ i are the corresponding parameters; The value network uses the value function V(A t ,s) Evaluate the performance of the joint control strategy: Among them, π(A t |s;θ) is the action A taken in state s t The probability of q(s,A t ; ω) is the action A taken in state s t The expected return of , θ, ω are the corresponding network parameters; The gradient descent of the corresponding parameters is calculated by the following formula: The joint control strategy uses the objective function J V (π) Calculate the value to maximize the sub-strategy π i The reward r(s t ,A t ) and entropy sum: Among them, α is a hyperparameter used to control the weight of the entropy regularization term; In the evaluation process of soft policy iteration, the soft Q value is iteratively calculated starting from an arbitrary function Q. and repeatedly apply the modified Bellman backup operator The soft value function of the updated sub-strategy, the Bellman backup operator is given by: In the improvement step of each sub-strategy, the corresponding strategy is updated according to the following formula for the state at each moment: in, is the old policy action value function, is the normalization factor. In each sub-strategy iteration, soft policy evaluation and soft policy improvement are performed alternately to ensure the optimal maximum entropy convergence between sub-strategy combinations.
8. The learning-based emergency collision avoidance control method for an autonomous commercial vehicle according to claim 1, characterized in that: In step 4, the reward function is: in, is the driving efficiency index, It is an indicator of driving safety. is the driving comfort index, It is the completion index of the emergency collision avoidance task; In the formula, v k is the vehicle speed, v limit With v max are the maximum speed limit of the road and the maximum speed that the vehicle can reach, respectively; ttc is the expected time of collision at the current speed, d lat is the lateral distance between vehicles, f c is a binary indicator of whether the vehicle has collided, v x is the vehicle's current speed along the axis of travel. ttc , lat , k c are all hyperparameters; thr and ρ ste are the hyperparameters of the longitudinal acceleration term and the lateral acceleration term respectively; is the Euclidean distance between the autonomous commercial vehicle and the background vehicle, c r is a hyperparameter.
9. A learning-based emergency collision avoidance control system for an autonomous commercial vehicle based on the method described in any one of claims 1 to 8, characterized in that: The system comprises: A data set construction module is used to collect driving data of human drivers, construct an expert experience data set, and add the expert experience data set to the replay buffer in the form of a decay ratio; A reinforcement learning network model construction module is used to construct a reinforcement learning network model based on a soft actor-critic architecture, wherein the reinforcement learning network model includes a policy network, a value network, and a Q-value network, wherein the policy network generates continuous driving actions based on sensor observation data, and the value network and the Q-value network evaluate the driving actions generated by the policy network; the policy network, the value network, and the Q-value network are trained by using observation-action pairs in a replay buffer to update corresponding network parameters; The joint control strategy module is used to build a corresponding Bayesian strategy network for commercial vehicle driving tasks and control actions based on the built reinforcement learning network model, expand the strategy network structure and network update process, and generate a joint control strategy that adapts to emergency collision avoidance conditions; the node sets of the Bayesian strategy network are steering δ k 、Throttle a k and brake k A three-action strategy; The optimal strategy output module is used to apply the driving actions generated by the Bayesian strategy network to the vehicle's driving tasks, and through interaction with the driving environment, with the goal of collision avoidance without collision at high speeds, conduct model training under the guidance of the reward function to generate the optimal strategy.
10. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed in a computer, the computer is caused to execute the emergency collision avoidance control method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Methods, systems, media and equipment for T-shaped emergency collision avoidance control of autonomous vehicles
CN113460090B
Automobile emergency collision avoidance control method based on DQN deep reinforcement learning
CN116176572A
Driverless car behavior decision making method based on type-2 fuzzy comprehensive evaluation
CN110308718A
Safe driving decision-making method for highly human-like automatic driving commercial vehicle
CN114407931A
Urban scene end-to-end automatic driving control method and device based on attention mechanism and graph model reinforcement learning
CN115303297A
Cited By
Reinforcement learning semi-active suspension intelligent control method based on Bayesian optimization
CN120792407A