An end-to-end autonomous driving decision-making method integrating prior knowledge

By constructing a control guidance model based on prior knowledge and P3O reinforcement learning algorithm, combined with the VC-BEVFormer network, the problem of slow learning speed and low anthropomorphism of autonomous driving under urban operating conditions is solved, and the end-to-end autonomous driving decisions with rapid anthropomorphism are realized.

CN119975417BActive Publication Date: 2025-07-18CHANGSHA AUTOMOBILE INNOVATION RES INST
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510483417.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-18
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

It is difficult to achieve natural integration into the human traffic environment under urban operating conditions. The traditional method has a slow learning speed and a low degree of anthropomorphism. The reinforcement learning method has problems such as low sample efficiency and slow strategy convergence.

Method used

By constructing a control-guided model based on prior knowledge, combining the VC-BEVFormer network and P3O reinforcement learning algorithm, the experience of manual driving mode is used for training, reducing the exploration space and improving the degree of anthropomorphism.

Benefits of technology

The learning speed of end-to-end autonomous driving decision model has been accelerated, the degree of anthropomorphism and training speed have been improved, and the performance of autonomous driving under urban operating conditions has been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119975417B_ABST
    Figure CN119975417B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of autonomous driving technology, and discloses an end-to-end autonomous driving decision-making method integrating prior knowledge, including: establishing a vehicle control and guidance model; in a simulation environment, using the vehicle control and guidance model to control the vehicle, and taking the control variables output by the vehicle control and guidance model as the first action variables; and taking the state variables and the corresponding first action variable groups as prior experiences; using the prior experiences to train a reinforcement learning network model to obtain a preliminarily trained reinforcement learning network model; using the preliminarily trained reinforcement learning network model to control the vehicle to drive, and switching to the manual driving mode when it is determined that there is a collision risk; storing the output control variables of the preliminarily trained reinforcement learning network model and the control variables output by the manual driving mode as the second action variables; using the state variables and their corresponding second action variables to iteratively optimize the preliminarily trained reinforcement learning network model to obtain an autonomous driving decision-making model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of autonomous driving, and particularly relates to an end-to-end autonomous driving decision-making method integrating prior knowledge. Background Art

[0002] At present, the technology in the field of autonomous driving is booming, but the algorithm implementation of autonomous driving vehicles under urban conditions still faces many problems. Among them, the intelligent decision-making technology problem is the focus of many problems and has received extensive attention in the academic community. Traditional decision-making and control methods rely on rules formulated by humans. Although they have a certain degree of interpretability, they often cannot cover all scenarios under complex urban conditions and naturally exhibit behaviors similar to those of human drivers. This will cause autonomous driving vehicles to be unable to naturally integrate into the human traffic environment, limiting the intelligence level of autonomous driving vehicles to a relatively low range.

[0003] Reinforcement learning (RL) shows the potential to exceed traditional control frameworks by autonomously exploring the dynamic characteristics of the environment. Deep reinforcement learning (DRL) algorithms can learn complex driving strategies from high-dimensional perceptual inputs without an explicit model. However, traditional DRL methods face many problems such as low sample efficiency and slow policy convergence speed in engineering implementation.

[0004] In recent years, some studies have tried different neural network architectures to introduce prior knowledge into the autonomous driving system to solve the problem of slow learning speed of agents. However, these results are still limited to the theoretical and experimental stages and have not been widely applied in the actual urban traffic environment.

[0005] Existing solutions generally have some problems when integrating prior knowledge. First, when existing methods use human-guided data, they often only use human actions for reward shaping and ignore the deeper implicit preferences in human data. This reward-shaping-based guidance method is prone to reward hacking phenomena, resulting in limited performance of the end-to-end network. Second, although some existing methods use preference reinforcement learning methods to shape the reward function using human behavior, because they are completely offline learning based on limited human operation data, they will inevitably lose a certain degree of autonomous exploration ability, resulting in problems such as poor task generalization and slow training speed. Summary of the Invention

[0006] The object of the present invention is an end-to-end autonomous driving decision-making method integrating prior knowledge, which reduces the exploration space of the end-to-end decision-making model in the initial stage of training through a control guidance model based on prior knowledge, and introduces the experience of manual driving mode during the reinforcement learning training process, which can improve the training speed while enhancing the anthropomorphic degree of end-to-end autonomous driving decision-making.

[0007] The technical solution provided by the present invention is as follows:

[0008] An end-to-end autonomous driving decision-making method integrating prior knowledge, comprising:

[0009] Establish a vehicle control and guidance model based on a two-degree-of-freedom vehicle model; in a simulation environment, obtain vehicle state information and traffic environment information where the vehicle is located in real time as state variables;

[0010] Use the vehicle control and guidance model to control the vehicle to perform preset trajectory tracking, and use the control variables output by the vehicle control and guidance model as the first action variables; and use the state variables and the corresponding first action variables as prior experience;

[0011] Construct a reinforcement learning network model, and use the prior experience to train the reinforcement learning network model to obtain a preliminarily trained reinforcement learning network model;

[0012] In a simulation environment, use the preliminarily trained reinforcement learning network model to control the vehicle to drive, and switch to the manual driving mode when it is determined that there is a collision risk; store the output control variables of the preliminarily trained reinforcement learning network model and the control variables output by the manual driving mode as the second action variables;

[0013] Use the state variables and their corresponding second action variables to iteratively optimize the preliminarily trained reinforcement learning network model to obtain an autonomous driving decision-making model;

[0014] During actual vehicle driving, obtain vehicle state information and traffic environment information where the vehicle is located in real time as state variables, input them into the autonomous driving decision-making model, and the autonomous driving decision-making model outputs a vehicle control strategy to achieve autonomous driving.

[0015] Preferably, the end-to-end autonomous driving decision-making method integrating prior knowledge further comprises:

[0016] Construct a VC-BEVFormer network model, which includes: a BEV feature extractor, a BEV encoder, a vehicle state encoder, and a cross-modal attention mechanism network;

[0017] Collect traffic environment pictures of the vehicle through an in-vehicle camera, extract the traffic environment pictures through the BEV feature extractor, and then input them into the BEV encoder for encoding to obtain BEV encoded segments;

[0018] Input the vehicle state information into the vehicle state encoder for encoding to obtain vehicle state encoded segments;

[0019] Fuse the BEV encoded segments and the vehicle state encoded segments through the cross-modal attention mechanism network to obtain state variables.

[0020] Preferably, the BEV feature extractor consists of a ResNet-50 module and a feature pyramid module.

[0021] Preferably, the reward function adopted when iteratively optimizing the preliminarily trained reinforcement learning network model is:

[0022] ;

[0023] where is the lane reward, is the lane keeping reward, is the speed reward, is the lateral stability reward.

[0024] Preferably, the calculation formula for the lane keeping reward is:

[0025] ;

[0026] In the formula, is the lane keeping reward coefficient; is the lane keeping cost coefficient; is the lateral offset distance of the vehicle relative to the center line of the lane; is the lane width.

[0027] Preferably, the calculation formula for the speed reward is:

[0028] ;

[0029] where is the speed reward coefficient; is the current speed of the vehicle; is the set target speed.

[0030] Preferably, the calculation formula for the lateral stability reward is:

[0031] ;

[0032] where and are weight coefficients; is the lateral speed, is the unit speed; is the lateral acceleration, is the unit acceleration.

[0033] Preferably, the loss function adopted when iteratively optimizing the policy network in the preliminarily trained reinforcement learning network model is:

[0034] ;

[0035] Among them, ;

[0036] In the formula, represents the parameter to be optimized of the policy network in the reinforcement learning network model; represents the th iteration of the parameters of the policy network; and are the state variable and the action variable respectively; represents the subsequent overall expectation for the state and obeys the state probability distribution under the policy ; ; represents the subsequent overall expectation for the action and obeys the action probability distribution of the policy at the state ; represents the advantage function of the agent taking the action under the policy at the state ; represents the degree of manual driving learning; represents the manual driving preference signal of the sample; represents the look-ahead value evaluation coefficient, represents the look-ahead value evaluation function, represents the difference in the preference data before and after policy optimization, represents the mean of the divergence of the current policy and the previous iteration policy in all states , represents the sigmoid function.

[0037] The beneficial effects of the present invention are as follows:

[0038] (1) The present invention introduces rule-based prior knowledge to solve the training speed problem. By constructing a control guidance model based on prior knowledge, the action data generated by the control theory method is incorporated into the learning scope of the decision model, compressing the meaningless exploration behavior of the end-to-end decision model in the initial stage of training and accelerating the learning speed of the end-to-end decision model.

[0039] (2) The present invention introduces a reinforcement learning scheme guided by manual driving to solve the problems of training speed and anthropomorphic degree. Based on the constructed P3O reinforcement learning algorithm, the end-to-end decision-making model can automatically analyze the intentions of human driving behaviors and convert the operation data of manual driving into manual driving (human behavior) preference data. By using the proposed P3O policy gradient method to improve the policy network, it can not only make full use of valuable human data to improve the performance and anthropomorphic degree of the model, but also accelerate the convergence speed of the end-to-end decision-making model.

[0040] (3) The present invention performs cross-modal feature fusion on vehicle BEV information and vehicle state information to extract the state features of the environment, and represents the vehicle state information of the environmental picture in a more abstract and refined form. This not only reduces the memory occupancy during the learning of the end-to-end decision-making model, but also accelerates the learning rate and learning effect of the end-to-end decision-making model. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is a schematic diagram of the end-to-end autonomous driving decision-making method integrating prior knowledge according to the present invention.

[0042] Figure 2 It is a schematic diagram of the prospective value estimation according to the present invention.

[0043] Figure 3 It is a comparison diagram between the VC-BEVFormer and the baseline network according to the present invention.

[0044] Figure 4 It is a comparison diagram of different reinforcement learning.

[0045] Among them, Figure 4 (a) in it is a comparison diagram of the collision rates of different reinforcement learning algorithms.

[0046] Figure 4 (b) in it is a comparison diagram of the lateral deviations of different reinforcement learning algorithms.

[0047] Figure 4 (c) in it is a comparison diagram of the longitudinal speeds of different reinforcement learning algorithms.

[0048] Figure 4 (d) in it is a comparison diagram of the longitudinal displacements of different reinforcement learning algorithms.

[0049] Figure 5 It is a schematic diagram of the waiting state of the agent (vehicle).

[0050] Among them, Figure 5 (a) in it is a schematic diagram of the agent (vehicle) decelerating and waiting before entering the intersection.

[0051] Figure 5Figure (b) is a schematic diagram of the agent (vehicle) waiting for oncoming vehicles to pass through the intersection.

[0052] Figure 5 Figure (c) is a schematic diagram of the agent (vehicle) starting to turn left after waiting for other vehicles to pass by.

[0053] Figure 5 Figure (d) is a schematic diagram of the agent (vehicle) driving into the target lane after waiting for other vehicles to pass by.

[0054] Figure 6 It is a schematic diagram of the driving state of the agent (vehicle).

[0055] Among them, Figure 6 Figure (a) is a schematic diagram of the agent (vehicle) accelerating into the intersection.

[0056] Figure 6 Figure (b) is a schematic diagram of the agent (vehicle) accelerating through the intersection when other vehicles drive into the intersection.

[0057] Figure 6 Figure (c) is a schematic diagram of the agent (vehicle) starting to turn left when other vehicles drive into the intersection.

[0058] Figure 6 Figure (d) is a schematic diagram of the agent (vehicle) driving into the target lane. Detailed implementation method

[0059] The following further elaborates on the present invention in conjunction with the attached drawings, enabling those skilled in the art to implement it with reference to the text of the specification.

[0060] As Figure 1 shown, the present invention discloses an end-to-end autonomous driving decision-making method integrating prior knowledge, which reduces the exploration space of the end-to-end decision-making model in the initial stage of training by constructing a control guidance model based on prior knowledge; simplifies the expression of the complex picture state space through VC-BEVFormer; and uses a reinforcement learning training method to generate human preference behavior data using manual driving data and improves the model using the P3O strategy to quickly realize the training of an anthropomorphic decision-making model.

[0061] The autonomous driving decision-making system based on the present invention adopts an end-to-end paradigm. By receiving information about the environment from the perception system, it outputs the control amount of the host vehicle and sends it to the execution system, enabling the vehicle to achieve end-to-end autonomous driving. Among them, the control amount of the host vehicle includes the opening of the throttle or brake pedal and the steering angle of the front wheels.

[0062] The specific implementation process of the present invention is as follows.

[0063] 1. Construct a control guidance model based on prior knowledge and collect control guidance experience data

[0064] In this embodiment, the control guidance model based on prior knowledge adopts a two-degree-of-freedom model of the vehicle:

[0065] ;

[0066] wherein, represents the sideslip angle of the vehicle; represents the yaw rate of the vehicle; and respectively represent the change rates of the sideslip angle and the yaw rate; and represent the front and rear wheel sideslip stiffnesses; represents the moment of inertia of the vehicle about the axis; and respectively represent the distances from the front and rear axles of the vehicle to the center of mass; represents the longitudinal speed of the vehicle; represents the mass of the vehicle; represents the front wheel steering angle.

[0067] Subsequently, an error response model is established. Considering the lateral error between the vehicle and the reference trajectory and the heading angle error of the reference trajectory, a vehicle path tracking error response model is established:

[0068] ;

[0069] where the state variables and respectively represent the lateral error and the heading angle error, and respectively represent the change rates of the lateral error and the heading angle error, represents the acceleration of the change of the lateral error, represents the acceleration of the change of the heading angle error; represents the front wheel steering angle, represents the change rate of the reference trajectory.

[0070] The above formula is simplified to:

[0071] ;

[0072] A cost function is introduced as an index to measure the control performance. The cost function is defined as:

[0073] ;

[0074] wherein, is a positive semi-definite matrix, indicating the penalty degree for the system state error; is a positive definite matrix, representing the penalty on the control input. By solving the minimum value of the above cost function, the optimal negative feedback gain matrix can be obtained , and its calculation formula is:

[0075] ;

[0076] where the matrix is obtained by solving the Riccati equation. According to the optimal control theory, the Riccati equation is:

[0077] ;

[0078] Using only a linear feedback controller is not sufficient to eliminate the error caused by road curvature changes. Therefore, a feedforward control term is introduced to compensate for the error caused by road curvature changes.

[0079] The feedforward term can be calculated by solving the following equation:

[0080]

[0081] Finally, the output expression of the control guidance model based on prior knowledge is:

[0082] ;

[0083] After constructing the control guidance model based on prior knowledge, it can be used to generate prior experience for model training in a simulation environment. The significance of prior experience is that it can greatly reduce the meaningless exploration of the agent in the initial stage of training. By introducing prior experience, the agent can quickly learn basic driving behaviors and lay a foundation for subsequent further learning.

[0084] II. Establish an agent environment perception network based on the VC-BEVFormer network to extract environmental features

[0085] The original image information received by the vehicle from the on-vehicle camera has a relatively high resolution. Directly treating it as a sequence input to the Transformer network incurs too high a learning cost. Therefore, a BEV feature extractor is designed to reduce the image resolution and retain high-level visual semantics while extracting image features. In this embodiment, the BEV feature extractor consists of a pre-trained ResNet-50 and a Feature Pyramid Networks (FPN), where all parameters are frozen and no backpropagation is performed. Its main function is to utilize the rich visual features learned by ResNet-50 on a large-scale dataset, and then further map and reduce the dimensionality of the extracted high-dimensional features through FPN, thereby generating a feature representation matrix with strong expressive ability. The output of the BEV feature extractor is the feature representation matrix , .

[0086] Among them, B represents the batch size; represents the height, represents the width, represents the number of RGB channels;

[0087] After obtaining the feature representation matrix , it is sent to the BEV encoder for further vectorization and Transformer encoding. Its operation logic first performs flattening and transposing, converting from shape to whose shape is :

[0088] ;

[0089] Among them, , represents the length along this dimension after flattening the width and height dimensions into one dimension. The Transformer module requires that each token (segment) of the input has the same dimension. Through linear projection, features from different sources can be mapped to the same embedding dimension, enabling features from images and states to be fused and compared in the same space and improving computational efficiency. In this embodiment, a pre-projection is performed on it, mapping the number of channels of each batch from to the unified embedding dimension . In this embodiment, . Applying a learnable mapping to , the representation in the embedding dimension is obtained:

[0090] ;

[0091] The Transformer module itself does not have the ability to capture the order or spatial position of each token in the sequence, because the self-attention mechanism calculates all positions of the input sequence in an unordered manner. To enable the model to understand the relative position of each token in the input sequence or the spatial position in the image, this embodiment introduces a CLS token (Classification token) and positional embeddings. The CLS token first appeared in language models such as BERT and is used to aggregate the global information of the entire input sequence. In this embodiment, the CLS token of the BEV information is denoted as , whose shape is , and after expanding the batch dimension, the shape becomes . Subsequently, it is concatenated to to obtain the representation after concatenating the CLS token:

[0092] ;

[0093] where the concatenation is performed along the token dimension, i.e., the second dimension;

[0094] Subsequently, a learnable positional encoding matrix is introduced, is a learnable encoding with a shape of , where after expanding the batch dimension has a shape of . Subsequently, it is added to to obtain the representation after embedding the positional encoding:

[0095] ;

[0096] where is the positional encoding matrix, and adding it element-wise to gives .

[0097] Finally, Transformer encoding is performed, but it requires a shape of , so the transpose symbol is defined:

[0098] ;

[0099] Then Transformer encoding is performed:

[0100] ;

[0101] Finally, the representation of the BEV information after Transformer encoding is obtained , that is, BEV tokens, whose shape is . Therefore, the BEV encoder in this embodiment receives a BEV feature map with the shape of and outputs as BEV tokens.

[0102] Subsequently, the vehicle state also needs to be encoded, which is responsible for encoding the vehicle state other than the image into the network and converting it into a token representation with the same dimension as the BEV tokens. The original vehicle state has the shape of . Add a dimension to it:

[0103] ;

[0104] Among them, is the representation after adding a dimension to the original state, B is the batch size, is the vehicle state parameter, .

[0105] Apply a learnable mapping to to obtain the state representation in the embedding dimension:

[0106] ;

[0107] Similar to the processing in the BEV encoder, there is also a relative position encoding matrix and a CLS token in the vehicle state encoder, where the relative position encoding is , with the shape of . After batch expansion, the shape is . The state representation after embedding position encoding is:

[0108] ;

[0109] To capture the global information of the entire vehicle state, a CLS token is added at the front of the sequence, denoted as , with the shape of . After expanding it to dimensions, the shape is . Subsequently, it is concatenated with along the second dimension to obtain the state representation after CLS concatenation:

[0110] ;

[0111] Finally, perform Transformer encoding for further information interaction and feature extraction. Each token calculates self-attention with other tokens, enabling the representation of each token to capture the context information of the entire vehicle state and extract the final vehicle state.

[0112] ;

[0113] Among them, is the output after Transformer encoding, that is, the vehicle state tokens.

[0114] After obtaining the BEV tokens and vehicle state tokens, in order to fuse the information from two different sources, a cross-modal attention network is used to achieve information fusion between images and scalars. First, at the token dimension, that is, the second dimension, the two are concatenated to form a new sequence :

[0115] ;

[0116] Subsequently, perform Transformer encoding to obtain:

[0117] ;

[0118] Directly extract the CLS token in the fused , and the BEV features after cross-modal fusion can be obtained , and the vehicle state features after cross-modal fusion . The two will be used as the final state representations for the value network and policy network of subsequent reinforcement learning.

[0119] III. Generate prior experience using the control guidance model and train the agent based on the P3O reinforcement learning algorithm.

[0120] In this embodiment, Carla is used as the simulation platform, and the map used is Town 05. This map provides the road topology, and each road consists of two-way four lanes. The selected scenario is an unprotected left turn, which includes straight lane changes, curves, and lateral straight vehicle interactions. The road number where the agent is located is 47, and the target road number is 0. During simulation, the vehicle will be randomly generated at different positions on lanes -1 or -2 on Road 47. The task objective is to move to Lane -1, pass through the intersection, and complete a left turn to Lane 1 on Road 0. The selected vehicle model is Tesla Model 3, and the physical parameters of the vehicle are obtained through the Carla python API. The Carla simulation environment can record the situation of the agent on the road in real time by outputting various state variables.

[0121] Using the control and guidance model and the environmental feature extraction model established previously, the acquisition of prior experience can be carried out. In this embodiment, for the working conditions of lane change and unprotected left turn, the trajectories to be followed are preset in the simulation environment, and the control and guidance model described in Step 1 is used for trajectory tracking. The agent collects prior experience in the Carla simulation environment by adopting the action output of the control and guidance model. These experiences will be used in the subsequent reinforcement learning training, aiming to reduce the initial exploration space of the agent and accelerate the learning of the policy network.

[0122] However, since these experiences are generated based on rules and the quality of the experiences has a certain upper limit, in order to enable the agent to have a faster learning speed and better performance, we cannot rely solely on these experiences. We also need to further use high-level data to guide the agent. The present invention introduces a human-in-the-loop reinforcement learning method to solve the problem of the source of high-level guiding data. The human-in-the-loop reinforcement learning method is divided into various forms, including human demonstration, human intervention, and human feedback, etc. The method of human demonstration is to use multiple complete demonstration operations performed by human experts before training; human intervention means that humans correct the agent's wrong behaviors in real time during the agent's training and take over the agent's operations in a timely manner through human operations; the method of human feedback is generally that after the agent completes a complete action, a human expert scores its performance. Essentially, it is to use human experts to generate reward signals for reward shaping.

[0123] This embodiment adopts a human-in-the-loop reinforcement learning method based on human intervention. Existing human intervention techniques usually impose a penalty on the state of the agent at the human intervention point to urge the agent to avoid going to that state in the future; or add additional rewards when estimating the action value of the human intervention action to encourage the agent to adopt actions similar to humans. However, the biggest problem with the existing technical methods is that guiding the agent's behavior through additional rewards will cause the reward hacking phenomenon, that is, the agent will generate behaviors that humans cannot understand in order to maximize the rewards, which reduces the performance of the model. Different from existing research, the P3O reinforcement learning algorithm proposed by the present invention analyzes the preference of human intervention behaviors and introduces the prospect theory in the field of economics to help the agent understand the preference degree of humans for different behaviors, so that the policy model can align with human intentions faster and better.

[0124] The state space of this embodiment includes the picture information directly collected by the vehicle from the built-in virtual camera of the simulator, and the vehicle position information obtained by the vehicle from sensors such as a combined inertial navigation system, that is, coordinates and coordinates, as well as seven state variables including the vehicle yaw angle, the lateral deviation from the reference road, the longitudinal speed, the longitudinal acceleration, and the yaw rate.

[0125] When it is judged that the behavior of the intelligent agent will significantly cause unsafe results such as collisions or large deviations from the lane, the human driver takes over the control of the intelligent agent. When there is no keyboard input for three consecutive seconds after the takeover, it is regarded as the completion of an intervention, and the intelligent agent resumes control.

[0126] In this embodiment, by setting a judgment module, it is judged whether the behavior of the intelligent agent will significantly cause unsafe results such as collisions or large deviations from the lane, as well as behaviors that do not conform to some traffic rules. If it is judged that the behavior of the intelligent agent will cause adverse results, after the judgment unit issues a prompt to switch to the manual driving mode, the human driver takes over the control of the intelligent agent.

[0127] The judgment module includes two major indicators: the speed anomaly indicator and the collision indicator.

[0128] The speed anomaly indicator judges whether the control right should be switched by analyzing the speed of the intelligent agent. Specifically, when the speed of the intelligent agent is less than 1 m / s and lasts for more than three seconds, it is considered that the speed of the intelligent agent is too low and the control right should be switched; when the speed of the intelligent agent exceeds 16.7 m / s (60 km / h), it is considered that the intelligent agent is speeding and the human driver should immediately take over the control right.

[0129] The collision indicator judges whether the control right should be switched by predicting whether the trajectory of the intelligent agent will collide with the environment within two seconds. For the collision prediction with the road environment, a vehicle kinematic model is used to judge whether the intelligent agent will drive out of the road within the next 2 seconds according to the current position, longitudinal speed, lateral speed, and yaw rate of the vehicle. If it is judged that the intelligent agent is about to drive out of the road, the human driver takes over the control right. For the collision with other vehicles at the crossroads, the collision is predicted by calculating the time for the ego vehicle and the other vehicle to reach the conflict point. Since the behavior model of the other vehicle can be set independently in Carla, and the reference lane line of the ego vehicle when passing through the crossroads is also set in advance before the simulation, the intersection point of the reference trajectories of the two is defined as the conflict point. Calculate the time for the intelligent agent to reach the conflict point along the reference line of the ego vehicle at the current vehicle speed. If > 2s, no control right switch is made. If , further calculate the time when the rear of the ego vehicle completely passes through the conflict point , and assume that the other vehicle moves at a constant speed, calculate the time when the front of the other vehicle passes through the conflict point , and the time when the rear of the other vehicle completely passes through the conflict point . If , or , it is determined that a collision will occur within 2 seconds, and the human driver takes over the control right.

[0130] In other embodiments, a human driver can also directly determine whether manual driving intervention is required through the operation of the simulation interface, and directly take over vehicle control when it is determined that manual driving intervention is required.

[0131] When switching to the manual mode, the human driver inputs intervention actions through keyboard operations, controls the longitudinal acceleration with the W / S keys, and controls the steering with the A / D keys. The data collection frequency is the same as the strategy frequency, which is 5 Hz.

[0132] In this embodiment, a number of reward functions are designed. To encourage the agent to stay in the center of the lane, a lane-keeping reward is designed, and the calculation formula is:

[0133] ;

[0134] where is the lane-keeping reward; is the lane-keeping reward coefficient; is the lane-keeping cost coefficient; is the lateral offset distance of the vehicle relative to the center line of the lane; is the lane width. To enable it to run not strictly along the center line of the lane, a smaller is set in this embodiment.

[0135] A speed reward is designed. The purpose is to impose a penalty when the vehicle speed is lower than the target speed to encourage the agent to accelerate to the target speed as much as possible, and to give a reward when the vehicle speed reaches or exceeds the target speed to encourage the agent to maintain or try to exceed the target speed. Specifically:

[0136] ;

[0137] where represents the speed reward; is the speed reward coefficient; is the current speed of the vehicle; is the set target speed.

[0138] Considering the lateral stability of the vehicle, a speed reward and an acceleration reward for lateral movement are introduced:

[0139] ;

[0140] where and are weight coefficients; is the lateral speed, is the unit speed, = 1 m / s; is the lateral acceleration, is the unit acceleration, = 1 m / s 2 。

[0141] Consider whether the vehicle is in the lane and introduce a lane reward term:

[0142] ;

[0143] Among them, is the lane reward, and its condition is being on the road and others.

[0144] The final reward function is:

[0145] ;

[0146] The P3O (Prospect Proximal Policy Optimization) reinforcement learning algorithm proposed by the present invention includes a policy network, a value network, and a P3O policy gradient algorithm. In addition, in order to make full use of human experience and different from the experience replay pool of traditional reinforcement learning methods, the P3O reinforcement learning method described in the present invention also includes a human preference experience pool and a temporary experience pool 。

[0147] At the end of each simulation episode, the environment will generate a quadruple experience for that episode 。Since there is human intervention in this technology, in order to distinguish the action source of the experience, an additional element is required to mark whether the operation comes from a human. Therefore, the experience containing the human intervention signal, that is, the intervention experience, is defined as the following five-tuple:

[0148] ;

[0149] Among them, represents the current state, represents the current action, represents the next state after the state transition, represents the immediate reward variable, represents the human intervention signal.

[0150] The human intervention signal is a boolean quantity, and its definition is:

[0151] ;

[0152] A temporary experience pool composed of the intervention experiences of each simulation episode is called the temporary experience pool , which is used to temporarily store and process data including human operations and the agent's own operations. By modifying the simulation environment, the Carla simulation environment can identify whether there is human intervention by real-time recognition of keyboard input signals, and then generate five-tuple intervention experience on the basis of general quadruple experience.

[0153] After obtaining the intervention experience, the existing technology often directly uses the intervention experience for reward shaping. In order to further utilize the human intervention experience, this technology constructs a human preference experience pool .

[0154] The human preference experience pool consists of preference experiences and the definition of preference experience is:

[0155] ;

[0156] Among them, represents the current state, represents the current action, represents the next state after the state transition, represents the immediate reward, represents the preference signal.

[0157] The definition of the preference signal is

[0158] ;

[0159] Among them, human aversion (experience) refers to the experience generated by the agent (reinforcement learning model) at the moment of switching to the manual mode, human preference (experience) refers to the experience generated by the manual driving mode, and human indifference (experience) refers to the experience generated by the reinforcement learning model.

[0160] The human preference experience pool can be processed and automatically generated according to the following steps:

[0161] Step 1, for a scene of the simulation, after the simulation ends, process the quadruple experience, generate the intervention experience, and form a temporary experience pool ;

[0162] Step 2, traverse , for the intervention experience among them , add the experience to the human preference experience pool ; for all the intervention experience among them , add the experience to the human preference experience pool , and at the same time generate a virtual action using the policy network; for the human preference experience pool Add experience 。

[0163] This means that the operation data generated during human intervention is human preference data, which encourages the agent to learn; the actions that the agent attempts to generate during human intervention are human-averse data, which encourages the agent to avoid similar behaviors; and when there is no human intervention, the agent's behavior is neither encouraged nor avoided, and the policy network updates the policy according to the maximized reward.

[0164] Preference experience can be a triple or a quintuple, which is determined by the source of the preference experience. The quintuple preference experience represents the experience that humans are indifferent to or prefer, because these experiences all come from the behavior of humans during intervention and the behavior of the agent without human intervention, and these actions actually appear in the simulation training process; while the human-averse data is virtually generated, because humans interrupt and replace the behavior of the agent through intervention, and the actions replaced by the agent are not applied in the simulation, so there will be no state transition and reward signal, so it is a triple.

[0165] Traditional reinforcement learning methods will randomly initialize the parameters of the policy network and then start training, continuously explore the environment and collect experience. The experience replay pool starts to accumulate from 0. When the collected experience reaches the minimum value of the preset update batch, the policy improvement starts. Different from the traditional technical solution, after constructing a control guidance model based on prior knowledge before the training starts, this model can be used to generate prior experience and store it in the human preference experience pool. Since there is no human intervention in the process of generating actions by the control guidance model based on prior knowledge, the preference signals contained in the generated prior experience should all be 0.

[0166] With the basis of these prior experiences, the agent can start learning the prior experiences in the experience pool at the beginning, so that the policy converges quickly. At the beginning of the training, the human preference experience pool only contains the experiences generated by the control guidance model. As the training progresses, the experiences generated by the agent's own policy network and the experiences generated by human intervention will be gradually added to the human preference experience pool. The human preference experience pool is a queue structure. When the number of experiences in it reaches the preset maximum value, the earliest experience is forgotten and new experiences are accepted.

[0167] After collecting the human preference experience, it can be used to guide the update of the policy network and the value network. The input of the policy network is the environmental state quantity, and the output is the action. That is to say, this network generates actions by receiving the observed quantity of the environment in real time, so it is named the policy network. The architecture of the policy network in this embodiment is the same as that of the policy network in the traditional reinforcement learning continuous action space. In the policy network of this embodiment Receives the state output by the cross-modal attention network and After that, they are spliced to form a splicing matrix :

[0168] ;

[0169] Then, it is mapped to 64 dimensions through a small multi-layer perceptron (MLP), and the mapped matrix is denoted as :

[0170] ;

[0171] Finally, the mean of the normal distribution followed by the continuous action is obtained through two sets of linear layers and the log standard deviation . At each time step, an action is output after sampling from this normal distribution

[0172] Since the collected preference experience has preference signals, how to quantify the preference degree of the preference experience and use it for policy update is the key issue and also the core of this technology. The P3O policy gradient algorithm introduces prospect theory. Prospect theory is an important basic theory in the field of behavioral economics, which believes that humans' evaluation of results is not absolutely rational, but they make judgments after psychologically evaluating their prospect value Figure 2 shows the image of the prospect value evaluation function. Based on this, prospect theory further explains four major effects

[0173] Certainty effect: When in a gain state, most people are risk-averse; Reflection effect: When in a loss state, most people are risk-seeking; Loss aversion: Most people are more sensitive to losses than to gains; Reference dependence: Most people's judgment of prospect value is often determined by the reference point

[0174] Based on prospect theory, the P3O policy gradient algorithm proposes the following loss function

[0175] ;

[0176] ;

[0177] ;

[0178] ;

[0179] where represents the parameters to be optimized of the current policy network represents the advantage function represents the th iteration of the parameters of the policy network It represents the learning degree of human behavior and is a hyperparameter that determines the degree to which the policy network learns human behavior; It represents the preference signal of the sample;

[0180] The loss function of the P3O policy gradient algorithm It is divided into two parts. The first part has the same core idea as existing technologies such as the TRPO policy gradient algorithm and the PPO policy gradient algorithm. That is, by maximizing the advantage function, it is ensured that each step of the policy improvement is monotonic. Among them:

[0181] represents the advantage function, which is used to evaluate the revenue value brought by a specific action. Compared with the average revenue obtained by taking actions according to the policy it shows how much more revenue is obtained. Since estimating the state value is required to calculate the advantage function, a value network is also needed .

[0182] The value network of this embodiment is similar to the policy network and uses a multi-layer perceptron. Its input is the state quantity, and the output is the value of this state under the established policy. The parameter update method of the value network adopts the temporal difference method and uses the following loss function for gradient descent:

[0183] ;

[0184] After the state value estimation of the policy is realized, this embodiment applies the Generalized Advantage Estimation method (GAE) widely adopted in the academic circle to evaluate the advantage function using the temporal difference idea. Specifically, when the agent is in the state and takes the action the advantage can be calculated according to the following formula:

[0185] ;

[0186] ;

[0187] Among them, is an additional hyperparameter introduced in GAE, and its range is between [0, 1], is the discount factor, is the temporal difference error.

[0188] The loss function of the P3O policy gradient algorithm The second half of is the human preference learning part, where:

[0189] represents the in all states between the current policy and the previous iteration policyThe mean of divergence measures the average difference between the new and old strategies. According to the reference dependence effect, the average difference between the new and old strategies is used as the reference benchmark for prospect value evaluation.

[0190] represents the actual difference between the new and old strategies at the preference data. Based on the reference dependence effect, and are subtracted to represent the deviation between the result and the expectation. Based on this, the prospect value is further evaluated. Among them, when the strategy is improved, samples are drawn from the human preference experience pool, and a series of values such as outcome are calculated using the drawn samples for calculating the policy gradient. Using the drawn samples is what is called "calculating at the preference data".

[0191] represents the prospect value evaluation coefficient, indicating the degree of human liking or disliking for the preference data. It is an adjustable non - negative hyperparameter, and its value is taken according to the preference signal of the sampled preference data samples during policy update.

[0192] ;

[0193] The larger it is, the more the agent imitates the data liked by humans during policy update; The larger it is, the more the agent avoids the data disliked by humans during policy update.

[0194] represents the sigmoid function, represents the prospect value evaluation function. Applying the sigmoid function as the prospect value evaluation function conforms to the certainty effect and reflection effect of prospect theory. quantifies the size of the policy gradient when using preference data for policy improvement and is a bridge connecting the policy and human preferences.

[0195] It should be noted that the hyperparameters and jointly determine the weights of maximizing the reward and maximizing human preferences when the agent performs policy update. When and are very small, the policy gradient is only oriented towards maximizing the reward, and the policy gradient of the preference term is no different from the traditional reinforcement learning method; when and are very large, the policy gradient only maximizes human preferences and completely ignores the reward. Therefore, in order to retain the exploration ability of the agent and its generalization in multiple scenarios, the existence of the gradient of maximizing the reward cannot be completely ignored, and should not be taken to be very large. In this embodiment, , , .

[0196] Through the above method, the described P3O reinforcement learning algorithm can be implemented, including the construction of a temporary experience pool and a human preference experience pool, as well as the construction and improvement of a value network and a policy network. Then, the trained agent (reinforcement learning policy network model) can be used for real vehicle autonomous driving control. When the agent is iterating in policy learning, it can increase the generation probability of human preference behaviors and decrease the generation probability of human aversion behaviors on the premise of the overall change range of the control policy, so as to realize the understanding of human preference behaviors and ensure that the end-to-end decision-making network of the trained intelligent vehicle can make anthropomorphic decisions at each time step according to the specific traffic conditions of the city, achieving the rapid training of the end-to-end decision-making network of autonomous driving vehicles.

[0197] To illustrate the superiority of the method proposed by the present invention, an ablation experiment was conducted on the end-to-end reinforcement learning decision-making method that combines prior knowledge and human behavior proposed by the present invention in the Carla simulation environment, verifying the roles of each module and comparing it with the open-source baseline method. The results are as Figures 3 - 4 shown.

[0198] According to Figure 3 it can be found that, with the RL algorithm fixed, the VC-BEVFormer network proposed by the present invention reaches the level that the baseline algorithm reaches only after 2000 episodes in less than 1000 episodes, which indicates that the representation ability based on the VC-BEVFormer network accelerates the learning speed of the RL algorithm; in addition, after the decision-making network training is stable, the performance of the decision-making model based on the VC-BEVFormer network is also better than other baseline methods, which indicates that the VC-BEVFormer network can improve the performance of the end-to-end decision-making network of autonomous driving vehicles.

[0199] According to Figure 4It can be found that, on the premise of fixing the VC-BEVFormer network, the scheme that simultaneously applies the control guidance model based on prior knowledge and the P3O reinforcement learning algorithm (LQR+P3O), compared with the scheme that only applies the control guidance model based on prior knowledge and the traditional reinforcement learning method SAC (LQR+SAC), the scheme that only adopts the P3O reinforcement learning algorithm (P3O), and the scheme that only adopts the traditional reinforcement learning method SAC (SAC), not only has the lowest collision rate, but also travels the farthest in the forward direction of the road under the premise of the lowest average speed. This indicates that the agent has quickly mastered the driving skills of safely interacting with environmental vehicles along the correct road at an unprotected left-turn intersection. Equally importantly, the agent of the LQR+P3O scheme reached the collision rate level after the baseline method was trained for 60,000 episodes with only about 20,000 episodes of training time, which indicates that this technical scheme can improve the training speed of the end-to-end decision-making network for autonomous driving. In addition, due to the limitations of the human intervention device in terms of the keyboard's control performance, the human path following level is also relatively limited, and the lateral deviation of the trained agent is relatively large. Compared with human intervention, the control guidance model based on prior knowledge shows excellent lateral deviation stability, which can help the agent quickly learn basic lane-keeping skills. However, the control guidance model based on prior knowledge lacks flexibility in terms of longitudinal speed and some avoidance behaviors, and cannot generate humanized interaction actions with other vehicles. It only accelerates forward to maximize the reward. The existence of human intervention has generated high-quality interaction experience and guided the agent to learn various interaction behaviors. Figure 5 and Figure 6 shows the specific performance of the trained agent in the simulated road environment. From the agent's trajectory, it can be found that the agent can not only successfully complete the passing task, but also exhibits various anthropomorphic behavior patterns, such as the behavior pattern of decelerating in advance before the intersection and accelerating through the intersection after waiting for other vehicles to pass through the intersection, as shown in Figure 5 ; or the behavior pattern of accelerating forward when other vehicles are far from the intersection and quickly passing through the intersection before other vehicles enter the crossroads, as shown in Figure 6 . In addition, by using human intervention data, the agent has successfully learned the action of "slightly avoiding" the vehicle on the left to the right side of the lane when passing through the intersection, as shown in Figure 5 (b), which indicates that the strategy trained based on human intervention is more flexible and delicate and has a higher degree of humanization.

[0200] An end-to-end reinforcement learning decision-making method that integrates prior knowledge and human behavior reduces the exploration space of the end-to-end decision-making model in the initial stage of training by constructing a control guidance model based on prior knowledge; simplifies the representation of the complex image state space by designing the VC-BEVFormer network; and uses the P3O reinforcement learning method to generate preference experience from human operation data and rapidly improve the model using the preference experience. This technology helps to improve the training speed and anthropomorphic degree of the end-to-end decision-making model for autonomous vehicles.

[0201] Although the embodiments of the present invention have been disclosed as above, they are not limited to the applications listed in the specification and embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the field, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details and the illustrated and described examples here.

Claims

1. An end-to-end autonomous driving decision-making method integrating prior knowledge, characterized in that, Including: Establish a vehicle control and guidance model based on a two-degree-of-freedom vehicle model; in a simulation environment, obtain vehicle state information and traffic environment information where the vehicle is located in real time as state variables; Use the vehicle control and guidance model to control the vehicle to perform preset trajectory tracking, and use the control variables output by the vehicle control and guidance model as the first action variables; and use the state variables and the corresponding first action variables as prior experience; Construct a reinforcement learning network model, and use the prior experience to train the reinforcement learning network model to obtain a preliminarily trained reinforcement learning network model; In a simulation environment, use the preliminarily trained reinforcement learning network model to control the vehicle to drive, and switch to the manual driving mode when it is determined that there is a collision risk; store the output control variables of the preliminarily trained reinforcement learning network model and the control variables output by the manual driving mode as the second action variables; Use the state variables and their corresponding second action variables to iteratively optimize the preliminarily trained reinforcement learning network model to obtain an autonomous driving decision-making model; During actual vehicle driving, obtain vehicle state information and traffic environment information where the vehicle is located in real time as state variables, input them into the autonomous driving decision-making model, and the autonomous driving decision-making model outputs a vehicle control strategy to achieve autonomous driving; The loss function used when iteratively optimizing the policy network in the preliminarily trained reinforcement learning network model is: where v(s,a) = σ[p[Outcome(s,a)-z ref ; Where, θ represents the parameter to be optimized in the policy network of the reinforcement learning network model; θ k represents the parameter of the policy network at the k-th iteration; s and a are the state variable and the action variable respectively; represents the subsequent overall expectation with respect to the state s, and s follows the state probability distribution under the policy represents the subsequent overall expectation with respect to the action a, and a follows the action probability distribution at the state s under the policy represents the advantage function of the agent taking the action a at the state s under the policy ; ω represents the degree of manual driving learning; p represents the manual driving preference signal of the sample; ξ represents the prospective value evaluation coefficient, v(s,a) represents the prospective value evaluation function, Outcome(s,a) represents the difference at the preference data before and after policy optimization, z ref represents the mean of the KL divergence between the current policy and the previous iteration policy in all states, and σ represents the sigmoid function.

2. The end-to-end autonomous driving decision-making method integrating prior knowledge according to claim 1, wherein Also including: Construct a VC-BEVFormer network model, which includes: a BEV feature extractor, a BEV encoder, a vehicle state encoder, and a cross-modal attention mechanism network; Collect traffic environment pictures where the vehicle is located through an on-vehicle camera. After extracting the traffic environment pictures through the BEV feature extractor, input them into the BEV encoder for encoding to obtain BEV encoding segments; Input the vehicle state information into the vehicle state encoder for encoding to obtain vehicle state encoding segments; Fuse the BEV encoding segments and the vehicle state encoding segments through the cross-modal attention mechanism network to obtain state variables.

3. The end-to-end autonomous driving decision-making method integrating prior knowledge according to claim 2, wherein, The BEV feature extractor is composed of a ResNet-50 module and a feature pyramid module.

4. The end-to-end autonomous driving decision-making method integrating prior knowledge according to any one of claims 1-3, characterized in that, The reward function used when iteratively optimizing the preliminarily trained reinforcement learning network model is: r=r road ·(r lane +r speed +r lateral ); Among them, r road is the lane reward, r lane is the lane keeping reward, r speed is the speed reward, r lateral is the lateral stability reward.

5. The end-to-end autonomous driving decision-making method integrating prior knowledge according to claim 4, wherein, The calculation formula for the lane keeping reward is: In the formula, C1 is the lane keeping reward coefficient; C2 is the lane keeping cost coefficient; l is the lateral offset distance of the vehicle relative to the center line of the lane; d is the lane width.

6. The end-to-end autonomous driving decision-making method integrating prior knowledge according to claim 5, wherein, The calculation formula for the speed reward is: Among them, C3 is the speed reward coefficient; v is the current speed of the vehicle; v target is the set target speed.

7. The end-to-end autonomous driving decision-making method integrating prior knowledge according to claim 6, wherein The calculation formula for the lateral stability reward is: Among them, C4 and C5 are weight coefficients; is the lateral velocity, is the unit velocity; is the lateral acceleration, is the unit acceleration.

Citation Information

Patent Citations

  • Automobile limit driving control method and system self-adaptive to unexpected external environment change

    CN116605242A

  • Knowledge and data fusion driven cloud control type networked vehicle cooperative cruise control method

    CN116853273A

  • Parking assisting method and device, equipment and storage medium

    CN118722598A

  • Knowledge-enhanced reinforcement learning vehicle decision control method and system

    CN118928468A

  • End-to-end automatic driving control system and device based on human preference reinforcement learning

    CN119018181A