Priori knowledge fused end-to-end automatic driving decision-making method

By constructing a control guidance model based on prior knowledge in the decision-making method of autonomous driving and introducing manual driving modes, the problems of low sample efficiency and slow strategy convergence in the existing technology are solved, and more efficient training and more anthropomorphic decisions are achieved.

CN119975417AActive Publication Date: 2025-05-13CHANGSHA AUTOMOBILE INNOVATION RES INST

Patent Information

Application Number
CN202510483417.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-13
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

When integrating prior knowledge, existing autonomous driving decision-making methods have problems such as low sample efficiency, slow strategy convergence speed and poor task generalization, and it is difficult to naturally integrate into the human traffic environment.

Method used

By building a control-guided model based on prior knowledge, the exploration space of the end-to-end decision model in the early stage of training is reduced, and the experience of manual driving mode is introduced in the reinforcement learning training process, improving the training speed and anthropomorphism.

Benefits of technology

It realizes the degree of personification of end-to-end autonomous driving decisions while improving training speed, quickly converge strategies and improve model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119975417A_ABST
    Figure CN119975417A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of automatic driving, and discloses a priori knowledge fused end-to-end automatic driving decision-making method, which comprises the following steps of: establishing a vehicle control guide model; in the simulation environment, vehicle control is conducted through the vehicle control guiding model, and a control variable output by the vehicle control guiding model serves as a first action variable; taking the state variable and the corresponding first action variable group as prior experience; training the reinforcement learning network model by using prior experience to obtain a preliminarily trained reinforcement learning network model; using the preliminarily trained reinforcement learning network model to control the vehicle to run, and switching to a manual driving mode when the collision risk is judged to occur; storing an output control variable of the preliminarily trained reinforcement learning network model and a control variable output by the manual driving mode as a second action variable; and performing iterative optimization on the preliminarily trained reinforcement learning network model by using the state variable and the second action variable corresponding to the state variable to obtain an automatic driving decision model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of autonomous driving, and in particular relates to an end-to-end autonomous driving decision-making method integrating prior knowledge. Background Art

[0002] Today, the technology in the field of autonomous driving is booming, but the implementation of autonomous driving vehicle algorithms in urban conditions still faces many challenges. Among them, the problem of intelligent decision-making technology is the focus of many problems and has received widespread attention from the academic community. Traditional decision-making planning and control methods rely on rules formulated by humans. Although they have certain interpretability, they often cannot cover all scenarios under complex urban conditions and naturally present behaviors similar to those of human drivers. This will result in the inability of autonomous vehicles to naturally integrate into the human traffic environment, limiting the intelligence level of autonomous vehicles to a lower range.

[0003] Reinforcement learning (RL) has shown the potential to surpass traditional control frameworks by autonomously exploring the dynamic characteristics of the environment. Deep reinforcement learning (DRL) algorithms can learn complex driving strategies from high-dimensional sensory inputs without explicit models. However, traditional DRL methods face many problems in engineering implementation, such as low sample efficiency and slow strategy convergence.

[0004] In recent years, some studies have tried different neural network architectures to introduce prior knowledge into autonomous driving systems to solve the problem of slow learning speed of intelligent agents. However, these results are still limited to the theoretical and experimental stages and have not yet been widely used in actual urban traffic environments.

[0005] There are some common problems in the integration of prior knowledge in existing solutions. First, when using human guidance data, existing methods often only use human actions to shape rewards, ignoring the deeper implicit preferences of human data. This guidance method that relies on reward shaping is very likely to cause reward hacking, resulting in limited performance of end-to-end networks; secondly, although some existing methods use preference reinforcement learning methods to shape reward functions using human behavior, they are completely based on offline learning with limited human operation data, and will inevitably lose a certain degree of autonomous exploration capabilities, resulting in poor task generalization and slow training speed. Summary of the invention

[0006] The purpose of the present invention is to provide an end-to-end autonomous driving decision-making method that integrates prior knowledge. It reduces the exploration space of the end-to-end decision-making model in the early stage of training through a control guidance model based on prior knowledge, and introduces the experience of manual driving mode in the reinforcement learning training process. It can improve the training speed while improving the degree of humanization of end-to-end autonomous driving decisions.

[0007] The technical solution provided by the present invention is: An end-to-end autonomous driving decision-making method integrating prior knowledge, including: A vehicle control and guidance model is established based on the vehicle two-degree-of-freedom model; in a simulation environment, the vehicle status information and the vehicle's traffic environment information are obtained in real time as state variables; The vehicle control guidance model is used to control the vehicle to track a preset trajectory, and the control variable output by the vehicle control guidance model is used as a first action variable; and the state variable and the corresponding first action variable are used as prior experience; Constructing a reinforcement learning network model, and using the prior experience to train the reinforcement learning network model to obtain a preliminarily trained reinforcement learning network model; In a simulation environment, the vehicle is controlled by using the initially trained reinforcement learning network model, and the vehicle is switched to a manual driving mode when a collision risk is determined; the output control variable of the initially trained reinforcement learning network model and the control variable output by the manual driving mode are stored as a second action variable; The state variables and their corresponding second action variables are used to iteratively optimize the initially trained reinforcement learning network model to obtain an autonomous driving decision model; When driving a real vehicle, vehicle status information and traffic environment information of the vehicle are obtained in real time as state variables, which are input into the automatic driving decision model. The automatic driving decision model outputs a vehicle control strategy to realize automatic driving.

[0008] Preferably, the end-to-end autonomous driving decision-making method integrating prior knowledge further includes: Construct the VC-BEVFormer network model, which includes: BEV feature extractor, BEV encoder, vehicle state encoder and cross-modal attention mechanism network; The vehicle-mounted camera collects a picture of the traffic environment in which the vehicle is located, extracts the picture of the traffic environment through a BEV feature extractor, and then inputs the picture into a BEV encoder for encoding to obtain a BEV encoding segment; Inputting the vehicle state information into a vehicle state encoder for encoding to obtain a vehicle state encoding segment; The BEV encoding segment and the vehicle state encoding segment are fused through a cross-modal attention mechanism network to obtain a state variable.

[0009] Preferably, the BEV feature extractor consists of a ResNet-50 module and a feature pyramid module.

[0010] Preferably, the reward function used when iteratively optimizing the initially trained reinforcement learning network model is: ; in, For lane rewards, Rewards for lane keeping, For speed reward, Bonus for lateral stability.

[0011] Preferably, the calculation formula of the lane keeping reward is: ; In the formula, is the lane keeping reward factor; is the lane keeping cost coefficient; is the lateral offset distance of the vehicle relative to the centerline of the lane; is the lane width.

[0012] Preferably, the calculation formula of the speed reward is: ; in, is the speed bonus factor; is the current speed of the vehicle; is the set target speed.

[0013] Preferably, the calculation formula of the lateral stability reward is: ; in, and is the weight coefficient; is the lateral velocity, is the unit speed; is the lateral acceleration, is the unit acceleration.

[0014] Preferably, the loss function used when iteratively optimizing the policy network in the initially trained reinforcement learning network model is: ; in, ; In the formula, Represents the parameters to be optimized of the policy network in the reinforcement learning network model; Indicates The parameters of the policy network at this iteration; and They are state variables and action variables respectively; Indicates the subsequent overall status Seek expectations, and Obey the strategy The state probability distribution under ; Indicates the subsequent overall action Seek expectations, and Submission to strategy In Status The action probability distribution at ; Representative in Strategy In this case, the agent is in state Take action Advantage function of Indicates the manual driving learning level; Manual driving preference signals of representative samples; represents the prospect value assessment coefficient, represents the prospect value assessment function, represents the difference in preference data before and after strategy optimization, Indicates the difference between the current strategy and the previous iteration strategy in all states The mean of the divergence, Represents the sigmoid function.

[0015] The beneficial effects of the present invention are: (1) The present invention introduces rule-based prior knowledge to solve the problem of training speed. By constructing a control guidance model based on prior knowledge, the action data generated by the control theory method is incorporated into the learning scope of the decision model, which compresses the meaningless exploration behavior of the end-to-end decision model in the early stage of training and accelerates the learning speed of the end-to-end decision model.

[0016] (2) This invention introduces a manual driving guided reinforcement learning solution to solve the problems of training speed and anthropomorphism. Based on the constructed P3O reinforcement learning algorithm, the end-to-end decision model can automatically analyze the intention of human driving behavior and convert the manual driving operation data into manual driving (human behavior) preference data. The proposed P3O policy gradient method is used to improve the policy network, which can not only make full use of valuable human data to improve the performance and anthropomorphism of the model, but also accelerate the convergence speed of the end-to-end decision model.

[0017] (3) The present invention performs cross-modal feature fusion of vehicle BEV information and vehicle status information, extracts the state features of the environment, and concisely represents the vehicle status information of the environment image in a more abstract form. This not only reduces the memory usage during learning of the end-to-end decision model, but also speeds up the learning rate and learning effect of the end-to-end decision model. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is a schematic diagram of the end-to-end autonomous driving decision-making method integrating prior knowledge according to the present invention.

[0019] Figure 2This is a schematic diagram of the prospect value estimation described in the present invention.

[0020] Figure 3 A comparison diagram of the VC-BEVFormer described in the present invention and the baseline network.

[0021] Figure 4 A comparison chart of different reinforcement learning methods.

[0022] in, Figure 4 (a) is a comparison chart of collision rates of different reinforcement learning algorithms.

[0023] Figure 4 (b) is a comparison chart of lateral deviations of different reinforcement learning algorithms.

[0024] Figure 4 (c) is a comparison chart of the longitudinal speed of different reinforcement learning algorithms.

[0025] Figure 4 (d) is a comparison chart of longitudinal displacement of different reinforcement learning algorithms.

[0026] Figure 5 Schematic diagram of the waiting state of the intelligent agent (vehicle).

[0027] in, Figure 5 (a) is a schematic diagram of the intelligent agent (vehicle) slowing down and waiting before entering the intersection.

[0028] Figure 5 (b) is a schematic diagram of the intelligent agent (vehicle) waiting for a lateral vehicle to pass through the intersection.

[0029] Figure 5 (c) is a schematic diagram of the intelligent agent (vehicle) waiting for other cars to pass and then starting to turn left.

[0030] Figure 5 (d) is a schematic diagram of the intelligent agent (vehicle) waiting for other vehicles to pass and then entering the target lane.

[0031] Figure 6 Schematic diagram of the driving state of the intelligent agent (vehicle).

[0032] in, Figure 6 (a) is a schematic diagram of an intelligent agent (vehicle) accelerating into an intersection.

[0033] Figure 6 (b) is a schematic diagram of the intelligent agent (vehicle) accelerating through the intersection when another car enters the intersection.

[0034] Figure 6 (c) is a schematic diagram of the agent (vehicle) starting to turn left when the other car enters the intersection.

[0035] Figure 6 (d) is a schematic diagram of the intelligent agent (vehicle) entering the target lane. DETAILED DESCRIPTION

[0036] The present invention is further described in detail below in conjunction with the accompanying drawings so that those skilled in the art can implement the invention with reference to the description.

[0037] like Figure 1 As shown, the present invention discloses an end-to-end autonomous driving decision-making method integrating prior knowledge, which reduces the exploration space of the end-to-end decision-making model in the early stage of training by constructing a control guidance model based on prior knowledge; simplifies the representation of complex image state space by VC-BEVFormer; generates human preference behavior data by using manual driving data through reinforcement learning training method, and uses P3O strategy improvement algorithm to improve the model, so as to quickly realize the training of an anthropomorphic decision-making model.

[0038] The autonomous driving decision system based on the present invention adopts an end-to-end paradigm, which receives information about the environment from the perception system, outputs the control amount of the vehicle, and sends it to the execution system, so that the vehicle can achieve end-to-end autonomous driving. The control amount of the vehicle includes the throttle or brake pedal opening and the front wheel angle.

[0039] The specific implementation process of the present invention is as follows.

[0040] 1. Build a control guidance model based on prior knowledge and collect control guidance experience data In this embodiment, the control guidance model based on prior knowledge adopts a two-degree-of-freedom model of the vehicle: ; in, Indicates the side slip angle of the vehicle; represents the yaw rate of the vehicle; and They represent the rates of change of the sideslip angle and yaw rate respectively; and Indicates the cornering stiffness of the front and rear wheels; Table car around The moment of inertia of the shaft; and They represent the distances from the front and rear axles of the vehicle to the center of mass respectively; Indicates the longitudinal velocity of the vehicle; Indicates the mass of the vehicle; Indicates the front wheel turning angle.

[0041] Then, an error response model is established. Considering the lateral error between the car and the reference trajectory and the heading angle error of the reference trajectory, the car path tracking error response model is established: ; Among them, the state variable and Represent the lateral error and heading angle error respectively, and They represent the lateral error change rate and the heading error change rate respectively, represents the lateral error change acceleration, Represents the acceleration of the heading angle error change; represents the front wheel turning angle, represents the rate of change of the reference trajectory.

[0042] Simplify the above formula to: ; Introducing the cost function , as an indicator to measure control performance. The cost function is defined as: ;

[0043] in, is a semi-positive definite matrix, which indicates the degree of penalty for the system state error; is a positive definite matrix, which represents the penalty for the control input. By solving the minimum value of the above cost function, the optimal negative feedback gain matrix can be obtained: , and its calculation formula is: ; Among them, the matrix It is obtained by solving the Riccati equation. According to the optimal control theory, the Riccati equation is: ; Using only a linear feedback controller is not enough to eliminate the error caused by the change in road curvature. Therefore, a feedforward control term is introduced , compensating for the error caused by changes in road curvature.

[0044] feedforward term The calculation of can be obtained by solving the following equation:

[0045] Finally, the output expression of the control guidance model based on prior knowledge is: ; After constructing a control guidance model based on prior knowledge, it can be used in a simulation environment to generate prior experience for model training. The significance of prior experience is that it can greatly reduce the meaningless exploration of the agent in the early stage of training. By introducing prior experience, the agent can quickly learn basic driving behaviors, laying the foundation for further learning.

[0046] 2. Establish an intelligent agent environment perception network based on the VC-BEVFormer network to extract environmental features The original image information received by the vehicle from the on-board camera has a large resolution. It is too costly to directly regard it as a sequence input Transformer network learning cost. Therefore, the BEV feature extractor is designed to reduce the image resolution and extract image features while retaining high-level visual semantics. In this embodiment, the BEV feature extractor consists of a pre-trained ResNet-50 and a feature pyramid (FPN), in which all parameters are frozen and no backpropagation is performed. Its main function is to use the rich visual features learned by ResNet-50 on large-scale data sets, and then further map and reduce the extracted high-dimensional features through FPN, thereby generating a feature representation matrix with strong expressive power. The output of the BEV feature extractor is the feature representation matrix , .

[0047] in, B Represents the batch size; Represents height, Represents the width, Represents the number of RGB channels; Get the feature representation matrix After that, it is sent to the BEV encoder for further vectorization and Transformer encoding. Its operation logic first flattens and transposes from The shape is converted to Its shape is : ; in, , represents the length along the dimension after the width and height dimensions are flattened into one dimension. The Transformer module requires that each input token (fragment) has the same dimension. Through linear projection, features from different sources can be mapped to the same embedding dimension, so that features from images and states can be fused and compared in the same space, and the computational efficiency is improved. This embodiment performs a pre-projection and increases the number of channels in each batch from Mapping to a unified embedding dimension In this embodiment, .right Apply a learnable mapping , and get the representation in the embedding dimension : ; The Transformer module itself does not have the ability to capture the order or spatial position of each token in the sequence, because the self-attention mechanism performs disordered calculations on all positions in the input sequence. In order to let the model understand the relative position of each token in the input sequence or the spatial position in the image, this embodiment introduces CLS token (Classification token) and position embedding. CLS token originally appeared in language models such as BERT and is used to gather global information of the entire input sequence. In this embodiment, the CLS token of BEV information is recorded as , whose shape is , after batch dimension expansion, the shape is , which is then spliced ​​into After that, we get the representation after splicing CLStoken : ; Among them, it means splicing along the dimension of token, that is, the second dimension; Then we introduce a learnable position encoding matrix , is a shape of A learnable encoding, where After batch dimension expansion, the shape is . Then, compare it with Add together to get the representation after embedding position encoding : ; in, is the position encoding matrix, and Adding each element together gives .

[0048] Finally, Transformer encoding is performed, but its required shape is , so define the transposition operator : ; Then perform Transformer encoding: ; Finally, we get the representation of BEV information after Transformer encoding. , i.e. BEV tokens, whose shape is Therefore, the BEV encoder in this embodiment receives the shape BEV characteristic map, output For BEVtokens.

[0049] The vehicle state needs to be encoded later, which is responsible for encoding the vehicle state other than the image into the network and converting it into a token representation of the same dimension as the BEV tokens. Shape , adding a dimension to it: ; in, The representation after adding a dimension to the original state, B is the batch size, is the vehicle state parameter, .

[0050] right Apply a learnable mapping , and get the state representation under the embedding dimension : ; Similar to the processing in the BEV encoder, the vehicle state encoder also has a relative position encoding matrix and CLStoken, where the relative position encoding is , the shape is , after batch expansion, the shape is The state representation after embedding position encoding for: ; In order to capture the global information of the entire vehicle state, a CLS token is added at the beginning of the sequence, denoted as , whose shape is , expanding it to After dimensioning, the shape is . Then with Splicing along the second dimension, we get the state representation after CLS splicing for: ; Finally, Transformer encoding is performed for further information interaction and feature extraction. Each token will calculate self-attention with other tokens, so that the representation of each token can capture the contextual information of the entire vehicle state and extract the final vehicle state.

[0051] ; in, is the output of Transformer encoding, i.e., vehicle status tokens.

[0052] After obtaining BEVtokens and vehicle status tokens, in order to fuse the information from two different sources, a cross-modal attention network is used to achieve information fusion between images and scalars. First, the two are concatenated in the token dimension, that is, the second dimension, to form a new sequence : ; Then Transformer encoding is performed to obtain: ; Direct extraction after fusion The CLS token in the image can be used to obtain the BEV features after cross-modal fusion. , and the vehicle state features after cross-modal fusion , both will be used as the final state representation in the value network and strategy network of subsequent reinforcement learning.

[0053] 3. Use the control guidance model to generate prior experience and train the intelligent agent based on the P3O reinforcement learning algorithm.

[0054] In this embodiment, Carla is used as a simulation platform, and the map used is Town 05, which provides a road topology. Each road consists of two-way four lanes. The selected scene is an unprotected left turn, which includes straight lane changes, curves, and lateral straight vehicle interactions. The road where the agent is located is numbered 47, and the target road is numbered 0. During the simulation, the vehicle will be randomly generated at different positions on lanes -1 or -2 on Road 47. The task goal is to move to lane -1 and pass the intersection to complete the left turn to lane 1 of Road 0. The selected vehicle model is Tesla Model 3, and the physical parameters of the vehicle are obtained through Carla python API. The Carla simulation environment can record the situation of the agent on the road in real time by outputting a variety of state quantities.

[0055] By using the control guidance model and the environmental feature extraction model established previously, prior experience can be collected. In this embodiment, for the conditions of lane change and unprotected left turn, the trajectory to be followed is preset in the simulation environment, and the control guidance model described in step 1 is used to track the trajectory. The intelligent agent collects prior experience in the Carla simulation environment by using the action output of the control guidance model. These experiences will be used in the subsequent reinforcement learning training, in order to reduce the initial exploration space of the intelligent agent and accelerate the learning of the policy network.

[0056] However, since these experiences are generated based on rules, the quality of the experience has a certain upper limit. Therefore, in order for the intelligent agent to have a faster learning speed and better performance, it is not enough to rely solely on these experiences. It is necessary to further use high-level data to guide the intelligent agent. The present invention introduces a human-in-the-loop reinforcement learning method to solve the problem of the source of high-level guidance data. Human-in-the-loop reinforcement learning methods are divided into many forms, including human demonstration, human intervention, human feedback, etc. The human demonstration method is to use human experts to perform multiple complete demonstration operations before training; human intervention refers to humans correcting the wrong behavior of the intelligent agent in real time during the training of the intelligent agent, and taking over the operation of the intelligent agent in time through human operation; the human feedback method is generally after the intelligent agent performs a complete scene of action, the human expert scores its performance, which is essentially using human experts to generate reward signals for reward shaping.

[0057] This embodiment adopts a human-in-the-loop reinforcement learning method based on human intervention. Existing human intervention technology usually imposes a penalty on the state of the agent at the point of human intervention to urge the agent to avoid going to that state in the future; or adds additional rewards when estimating the action value of the human intervention action to encourage the agent to adopt human-like actions. However, the biggest problem with the existing technical methods is that guiding the behavior of the agent through additional rewards will cause the reward hacking phenomenon, that is, the agent will produce behaviors that humans cannot understand in order to maximize the reward, which reduces the performance of the model. Different from existing studies, the P3O reinforcement learning algorithm proposed in the present invention analyzes the preference of human intervention behavior and introduces the prospect theory in the field of economics to help the agent understand the degree of human preference for different behaviors, so that the strategy model can align human intentions faster and better.

[0058] The state space of this embodiment includes image information directly collected by the vehicle from the built-in virtual camera of the simulator, and vehicle position information obtained by the vehicle from sensors such as the combined inertial navigation, that is, coordinates and coordinates, as well as seven state variables: vehicle yaw angle, lateral deviation from the reference road, longitudinal speed, longitudinal acceleration, and yaw angular velocity.

[0059] When it is judged that the behavior of the intelligent agent will significantly cause unsafe consequences such as collision or large lane deviation, the human driver will take over the control of the intelligent agent. If there is no keyboard input for three consecutive seconds after the takeover, it is considered that the intervention is completed and the intelligent agent resumes control.

[0060] In this embodiment, a judgment module is set to judge whether the behavior of the agent will significantly cause unsafe consequences such as collision or lane deviation, as well as behaviors that do not comply with some traffic regulations. If it is judged that the behavior of the agent will cause adverse consequences, the judgment unit issues a manual driving mode switching prompt, and the human driver takes over the control of the agent.

[0061] The judgment module includes two major indicators: abnormal speed indicator and collision indicator.

[0062] The speed anomaly indicator determines whether control should be switched by analyzing the speed of the intelligent body. Specifically, when the speed of the intelligent body is less than 1m / s and lasts for more than three seconds, it is considered that the speed of the intelligent body is too low and the control should be switched; when the speed of the intelligent body exceeds 16.7m / s (60km / h), it is considered that the intelligent body is speeding and the human driver should take over the control immediately.

[0063] The collision indicator determines whether control should be switched by predicting whether the trajectory of the agent will collide with the environment within two seconds. For collision prediction with the road environment, the vehicle kinematic model is used to determine whether the agent will run out of the road within the next two seconds based on the vehicle's current position, longitudinal speed, lateral speed, and yaw angular velocity. If it is determined that the agent is about to run out of the road, the human driver takes over control. For collisions with other vehicles at intersections, collision prediction is performed by calculating the time it takes for the vehicle and the other vehicle to reach the conflict point. Since the behavior model of the other vehicle can be set autonomously in Carla, and the reference lane line of the vehicle when passing through the intersection is also set in advance before the simulation, the intersection of the two reference trajectories is defined as the conflict point. Calculate the time it takes for the agent to reach the conflict point along the reference line of the vehicle at the current vehicle speed. .like >2s, no control switching is performed. , further calculate the time it takes for the rear end of the vehicle to completely pass the conflict point , and assuming that the other car moves at a constant speed, calculate the time it takes for the front of the other car to pass the conflict point , and the time it takes for the rear end of his car to completely pass the conflict point .like ,or , it is determined that a collision is about to occur within 2 seconds, and the human driver takes over control.

[0064] In other embodiments, a human driver may also directly determine whether manual driving intervention is required through the operation of a simulation interface, and directly take over vehicle control when it is determined that manual driving intervention is required.

[0065] When switching to manual mode, the human driver intervenes by inputting action through keyboard operation, controlling longitudinal acceleration through the W / S key and steering through the A / D key. The data acquisition frequency is the same as the strategy frequency, which is 5 Hz.

[0066] This embodiment designs a number of reward functions. In order to encourage the agent to stay in the center of the lane, a lane keeping reward is designed, and the calculation formula is: ; in, It’s a lane keeping reward; is the lane keeping reward factor; is the lane keeping cost coefficient; is the lateral offset distance of the vehicle relative to the centerline of the lane; is the lane width. In order to allow the vehicle to run without strictly following the lane centerline, this embodiment sets a smaller .

[0067] The speed reward is designed to impose a penalty when the vehicle speed is lower than the target speed, motivating the agent to accelerate to the target speed as much as possible. At the same time, a reward is given when the vehicle speed reaches or exceeds the target speed, prompting the agent to maintain or try to exceed the target speed. The details are as follows: ; in, Indicates speed reward; is the speed bonus factor; is the current speed of the vehicle; is the set target speed.

[0068] Taking into account the lateral stability of the vehicle, the velocity bonus and acceleration bonus for lateral motion are introduced: ; in, and is the weight coefficient; is the lateral velocity, is the unit speed, =1m / s; is the lateral acceleration, is the unit acceleration, =1m / s 2 .

[0069] Considering whether the vehicle is in the lane, a lane reward term is introduced: ; in, For lane bonus, the conditions are on the road and others.

[0070] The final reward function is: ; The P3O (Prospect Proximal Policy Optimization) reinforcement learning algorithm proposed in the present invention includes a policy network, a value network, and a P3O policy gradient algorithm. In addition, in order to make full use of human experience, the P3O reinforcement learning method described in the present invention also includes a human preference experience pool, which is different from the experience playback pool of traditional reinforcement learning methods. With temporary experience pool .

[0071] At the end of each simulation episode, the environment generates a quadruple of experience points for that episode. Since this technology involves human intervention, in order to distinguish the source of the action of the experience, an additional element is required to mark whether the operation comes from humans. Therefore, the experience that contains human intervention signals, namely intervention experience, is defined as the following five-tuple: ; in, Represents the current state, Represents the current action, Represents the next state after state transfer, represents the immediate reward variable, Represents a human intervention signal.

[0072] The human intervention signal is a Boolean variable defined as: ; Intervention experience simulated by each scene A temporary experience pool is formed, called a temporary experience pool , which is used to temporarily store and process data including human operations and the agent's own operations. By modifying the simulation environment, the Carla simulation environment can identify whether there is human intervention by real-time recognition of keyboard input signals, and generate five-tuple intervention experience based on general four-tuple experience.

[0073] After obtaining intervention experience, existing technologies often directly use intervention experience to shape rewards. In order to further utilize human intervention experience, this technology constructs a human preference experience pool. .

[0074] The human preference experience pool consists of preference experience Composition, preference experience is defined as: ; in, Represents the current state, Represents the current action, Represents the next state after state transfer, Represents instant rewards, Represents a preference signal.

[0075] Preference Signals The definition is ; Among them, human aversion (experience) refers to the experience generated by the agent (reinforcement learning model) when switching to manual mode, human preference (experience) refers to the experience generated by the manual driving mode, and human indifference (experience) refers to the experience generated by the reinforcement learning model.

[0076] The human preference experience pool can be processed and automatically generated according to the following steps: Step 1: For a simulation scene, after the simulation is finished, process the quadruple experience, generate intervention experience, and form a temporary experience pool. ; Step 2, traverse , for which intervention experience, adding to the human preference experience pool Add experience ; For all of them intervention experience, adding to the human preference experience pool Add experience , and use the policy network to generate virtual actions ; is the human preference experience pool Adding Experience .

[0077] This means that the operation data generated when humans intervene are human preference data, which encourages the agent to learn; the actions that the agent attempts to generate when humans intervene are human aversion data, which encourages the agent to avoid similar behaviors; and when humans do not intervene, the agent's behavior is neither encouraged nor avoided, and the policy network updates the strategy according to the maximum reward.

[0078] Preference experience may be a triple or a quintuple, which is determined by the source of the preference experience. Quintuple preference experience represents human indifference or preference experience, because these experiences come from human intervention behavior and the behavior of the intelligent agent without human intervention, and these actions actually appear in the simulation training process; while human disgust data is virtually generated, because humans interrupt and replace the behavior of the intelligent agent through intervention, and the replaced action of the intelligent agent is not used in the simulation, so no state transfer and reward signal will be generated, so it is a triple.

[0079] Traditional reinforcement learning methods will randomly initialize the policy network parameters and then start training, continuously explore the environment and collect experience. The experience replay pool starts to accumulate from 0. When the collected experience reaches the minimum value of the preset update batch, the policy improvement will begin. This technology is different from traditional technical solutions. After building a control guidance model based on prior knowledge before training begins, the model can be used to generate prior experience and store it in the human preference experience pool. Since there is no human intervention in the process of generating actions based on the control guidance model based on prior knowledge, the preference signals contained in the generated prior experience should all be 0.

[0080] With these prior experiences as a foundation, the agent can learn the prior experiences in the experience pool from the beginning, so that the strategy converges quickly. At the beginning of training, the human preference experience pool only contains the experience generated by the control guidance model. As training progresses, the human preference experience pool will gradually add the experience generated by the agent's own strategy network and the experience generated by human intervention. The human preference experience pool is a queue structure. When the number of experiences in it reaches the preset maximum value, the earliest experience is forgotten and new experience is accepted.

[0081] After collecting human preference experience, it can be used to guide the update of the policy network and the value network. The input of the policy network is the state of the environment, and the output is the action. In other words, the network generates actions by receiving the observation of the environment in real time, hence the name policy network. The policy network of this embodiment has the same architecture as the policy network in the traditional reinforcement learning continuous action space. The state that receives the output of the cross-modal attention network and After that, they are spliced ​​to form a splicing matrix : ; Then it is mapped to 64 dimensions through a small multi-layer perceptron (MLP), and the mapped matrix is ​​recorded as : ; Finally, two sets of linear layers are used to obtain the mean of the normal distribution of continuous actions. and log standard deviation At each time step, the normal distribution is sampled and the output action is .

[0082] Since the collected preference experience has preference signals, how to quantify the preference degree of preference experience and use it for strategy update is the key to the problem and the core of this technology. The P3O policy gradient algorithm introduces the prospect theory. The prospect theory is an important basic theory in the field of behavioral economics. It believes that human evaluation of results is not absolutely rational, but makes judgments after psychologically evaluating its prospect value. Figure 2 The graph of the prospect value assessment function is shown. Based on this, the prospect theory further explains four major effects: Certainty effect: when in a state of gain, most people are risk-averse; reflection effect: when in a state of loss, most people are risk-seekers; loss aversion: most people are more sensitive to losses than gains; reference dependence: most people's judgment of prospective value is often determined by the reference point.

[0083] Based on the prospect theory, the P3O policy gradient algorithm proposes the following loss function: ; ; ; ; in, Indicates the parameters to be optimized for the current policy network; represents the advantage function; Indicates The parameters of the policy network at this iteration; Indicates the degree of human behavior learning, which is a hyperparameter that determines the degree to which the policy network learns human behavior; represents the preference signal of the sample; Loss function of the P3O policy gradient algorithm It is divided into two parts. The first part has the same core idea as existing technologies, such as the TRPO policy gradient algorithm and the PPO policy gradient algorithm. They all ensure that the improvement of each step of the strategy is monotonic by maximizing the advantage function. Among them: Represents the advantage function, which is used to evaluate the benefit value brought by a specific action and the benefit value brought by following the strategy How much greater is the benefit compared to the average benefit of taking an action? Because calculating the advantage function requires estimating the state value, a value network is also needed .

[0084] The value network of this embodiment Similar to the policy network, a multi-layer perceptron is used, whose input is the state quantity and output is the value of the state under the given policy. The parameter update method of the value network adopts the temporal difference method and uses the following loss function for gradient descent: ; After realizing the state value estimation of the strategy, this embodiment applies the generalized advantage estimation method (GAE) widely used in academia to evaluate the advantage function using the idea of ​​temporal difference. Next action Advantages It can be calculated according to the following formula: ; ; in, is an additional hyperparameter introduced in GAE, ranging from [0,1]. is the discount factor, is the timing difference error.

[0085] Loss function of the P3O policy gradient algorithm The second half is the human preference learning part, where: Indicates the difference between the current strategy and the previous iteration strategy in all states The mean of the divergence measures the average difference between the new and old strategies. Serves as a reference benchmark for prospect value assessment.

[0086] represents the actual difference between the new and old strategies in the preference data. Based on the reference dependence effect, and The difference indicates the deviation between the result and the expectation, and the prospect value is further evaluated based on this. When the strategy is improved, samples will be drawn from the human preference experience pool, and the drawn samples will be used to calculate a series of values ​​such as the outcome, which are used to calculate the policy gradient. Using the drawn samples is the so-called "calculation at the preference data".

[0087] Represents the prospect value assessment coefficient, which indicates the degree of human liking or disliking of preference data. It is an adjustable non-negative hyperparameter. When the strategy is updated, it is based on the preference signal of the sampled preference data.

[0088] ; The larger it is, the more the agent imitates the data that humans like when updating its strategy. The larger it is, the more the agent avoids human-disliked data when updating its strategy.

[0089] represents the sigmoid function, Represents the prospect value evaluation function. Applying the sigmoid function as the prospect value evaluation function conforms to the deterministic effect and reflection effect of the prospect theory. It quantifies the size of the policy gradient when using preference data to improve the policy, and is a bridge between the policy and human preferences.

[0090] It should be noted that the hyperparameters and Together they determine the weight of pursuing maximization of rewards and the weight of pursuing maximization of human preferences when the agent updates its strategy. and Very small, the policy gradient is only guided by maximizing the reward, ignoring the policy gradient of the preference item, which is no different from the traditional reinforcement learning method; and When θ is large, the policy gradient only maximizes human preferences and completely ignores the reward. Therefore, in order to preserve the agent's exploration ability and generalization in multiple scenarios, the existence of the gradient that maximizes the reward cannot be completely ignored. and It is not advisable to take a large value. In this embodiment, , , .

[0091] Through the above method, the P3O reinforcement learning algorithm can be realized, including the construction of temporary experience pool, human preference experience pool, and the construction and improvement of value network and strategy network. After that, the trained intelligent agent (reinforcement learning strategy network model) can be used to control the automatic driving of the real vehicle. When the intelligent agent is iterating the strategy learning, it can increase the probability of generating human preference behavior and reduce the probability of generating human aversion behavior under the premise of controlling the overall change range of the strategy, so as to realize the understanding of human preference behavior and ensure that the trained intelligent car end-to-end decision network can make anthropomorphic decisions at each time step according to the specific traffic conditions of the city, realizing the rapid training of the end-to-end decision network of the autonomous driving car.

[0092] To illustrate the superiority of the method proposed in this paper, an ablation test was conducted on the end-to-end reinforcement learning decision-making method that integrates prior knowledge and human behavior proposed in this paper in the Carla simulation environment to verify the role of each module and compare it with the open source baseline method. The results are as follows: Figure 3-Figure 4 shown.

[0093] according to Figure 3It can be found that when the RL algorithm is fixed, the VC-BEVFormer network proposed in the present invention reaches the level of the baseline algorithm in 2000 episodes in less than 1000 episodes, which shows that the representation ability based on the VC-BEVFormer network accelerates the learning speed of the RL algorithm; in addition, after the decision network training is stable, the performance of the decision model based on the VC-BEVFormer network is also better than other baseline methods, which shows that the VC-BEVFormer network can improve the performance of the end-to-end decision network of autonomous driving vehicles.

[0094] according to Figure 4 It can be found that under the premise of fixing the VC-BEVFormer network, the scheme of applying the control guidance model based on prior knowledge and the P3O reinforcement learning algorithm (LQR+P3O) at the same time, compared with the scheme of only the control guidance model based on prior knowledge and the traditional reinforcement learning method SAC (LQR+SAC), the scheme of only using the P3O reinforcement learning algorithm (P3O), and the scheme of only using the traditional reinforcement learning method SAC (SAC), not only has the lowest collision rate, but also the longest direction along the road under the premise of the lowest average speed, which shows that the intelligent agent quickly mastered the driving skills of safely interacting with environmental vehicles along the correct road at unprotected left-turn intersections. Equally important, the LQR+P3O scheme intelligent agent only took about 20,000 episodes of training time to reach the collision rate level of the baseline method after 60,000 episodes of training, which shows that this technical solution can improve the training speed of the end-to-end decision network of autonomous driving. In addition, because the equipment with human intervention is limited by the control performance of the keyboard, the path following level of humans is also relatively limited, and the lateral deviation of the trained intelligent agent is large. Compared with human intervention, the control guidance model based on prior knowledge shows excellent lateral deviation stability, which can help the intelligent agent quickly learn basic lane keeping skills. However, the control guidance model based on prior knowledge lacks flexibility in longitudinal speed and some avoidance behaviors, and cannot produce humanized interactive actions with other vehicles. It only accelerates forward to maximize rewards. The existence of human intervention generates high-quality interactive experience and guides the intelligent agent to learn a variety of interactive behaviors. Figure 5 and Figure 6 The specific performance of the trained agent in the simulated road environment is shown. Through the trajectory of the agent, it can be found that the agent can not only successfully complete the traffic task, but also show a variety of anthropomorphic behavior patterns, such as slowing down in advance before the intersection, waiting for other cars to pass the intersection, and then speeding up to pass the intersection. Figure 5 or when the other car is far away from the intersection, choose to speed up and quickly pass the intersection before the other car enters the intersection, as shown in Figure 6In addition, by using human intervention data, the agent successfully learned to "slightly avoid" the vehicle on the left side of the lane when passing through the intersection, as shown in Figure 5 As shown in (b), this shows that the strategy based on human intervention training is more flexible and delicate, and has a higher degree of humanization.

[0095] The present invention discloses an end-to-end reinforcement learning decision-making method that integrates prior knowledge and human behavior. The method reduces the exploration space of the end-to-end decision-making model in the early stage of training by constructing a control guidance model based on prior knowledge; simplifies the representation of the state space of complex images by designing a VC-BEVFormer network; and generates preference experience using human operation data through a P3O reinforcement learning method, and uses the preference experience to quickly improve the model. The technology helps to improve the training speed and anthropomorphism of the end-to-end decision-making model of autonomous driving vehicles.

[0096] Although the embodiments of the present invention have been disclosed as above, they are not limited to the applications listed in the specification and the implementation modes, and they can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to the specific details and the illustrations shown and described herein.

Claims

1. An end-to-end autonomous driving decision-making method integrating prior knowledge, characterized in that: include: A vehicle control and guidance model is established based on the vehicle two-degree-of-freedom model; in a simulation environment, the vehicle status information and the vehicle's traffic environment information are obtained in real time as state variables; The vehicle control guidance model is used to control the vehicle to track a preset trajectory, and the control variable output by the vehicle control guidance model is used as a first action variable; and the state variable and the corresponding first action variable are used as prior experience; Constructing a reinforcement learning network model, and using the prior experience to train the reinforcement learning network model to obtain a preliminarily trained reinforcement learning network model; In a simulation environment, the vehicle is controlled by using the initially trained reinforcement learning network model, and the vehicle is switched to a manual driving mode when a collision risk is determined; the output control variable of the initially trained reinforcement learning network model and the control variable output by the manual driving mode are stored as a second action variable; The state variables and their corresponding second action variables are used to iteratively optimize the initially trained reinforcement learning network model to obtain an autonomous driving decision model; When driving a real vehicle, vehicle status information and traffic environment information of the vehicle are obtained in real time as state variables, which are input into the automatic driving decision model. The automatic driving decision model outputs a vehicle control strategy to realize automatic driving.

2. The end-to-end autonomous driving decision-making method integrating prior knowledge according to claim 1, characterized in that: Also includes: Construct the VC-BEVFormer network model, which includes: BEV feature extractor, BEV encoder, vehicle state encoder and cross-modal attention mechanism network; The vehicle-mounted camera collects a picture of the traffic environment in which the vehicle is located, extracts the picture of the traffic environment through a BEV feature extractor, and then inputs the picture into a BEV encoder for encoding to obtain a BEV encoding segment; Inputting the vehicle state information into a vehicle state encoder for encoding to obtain a vehicle state encoding segment; The BEV encoding segment and the vehicle state encoding segment are fused through a cross-modal attention mechanism network to obtain a state variable.

3. The end-to-end autonomous driving decision-making method integrating prior knowledge according to claim 2, characterized in that: The BEV feature extractor consists of a ResNet-50 module and a feature pyramid module.

4. The end-to-end autonomous driving decision-making method integrating prior knowledge according to any one of claims 1 to 3, characterized in that: The reward function used in iterative optimization of the initially trained reinforcement learning network model is: ; in, For lane rewards, Rewards for lane keeping, For speed reward, Bonus for lateral stability.

5. The end-to-end autonomous driving decision-making method integrating prior knowledge according to claim 4, characterized in that: The calculation formula of the lane keeping reward is: ; In the formula, is the lane keeping reward factor; is the lane keeping cost coefficient; is the lateral offset distance of the vehicle relative to the centerline of the lane; is the lane width.

6. The end-to-end autonomous driving decision-making method integrating prior knowledge according to claim 5, characterized in that: The calculation formula of the speed bonus is: ; in, is the speed bonus coefficient; is the current speed of the vehicle; is the set target speed.

7. The end-to-end autonomous driving decision-making method integrating prior knowledge according to claim 6, characterized in that: The calculation formula of the lateral stability bonus is: ; in, and is the weight coefficient; is the lateral velocity, is the unit speed; is the lateral acceleration, is the unit acceleration.

8. The end-to-end autonomous driving decision-making method integrating prior knowledge according to claim 7, characterized in that: The loss function used in iterative optimization of the policy network in the initially trained reinforcement learning network model is: ; in, ; In the formula, Represents the parameters to be optimized of the policy network in the reinforcement learning network model; Indicates The parameters of the policy network for this iteration; and They are state variables and action variables respectively; Indicates the subsequent overall status Seek expectations, and Obey the strategy The state probability distribution under ; Indicates the subsequent overall action Seek expectations, and Submission to strategy In Status The action probability distribution at ; Representative in Strategy In this case, the agent is in state Take action Advantage function of Indicates the manual driving learning level; Manual driving preference signals of representative samples; represents the prospect value assessment coefficient, represents the prospect value assessment function, represents the difference in preference data before and after strategy optimization, Indicates the difference between the current strategy and the previous iteration strategy in all states The mean of the divergence, Represents the sigmoid function.

Citation Information

Patent Citations

  • Automatic driving lane changing decision control method based on rule fusion reinforcement learning

    CN115257745A

  • Automobile limit driving control method and system self-adaptive to unexpected external environment change

    CN116605242A

  • Unmanned vehicle adaptive path planning method based on dynamic window method and near-end strategy

    CN116679719A

  • Knowledge and data fusion driven cloud control type networked vehicle cooperative cruise control method

    CN116853273A

  • Automatic driving model, training method and device and vehicle

    CN116881707A

Cited By

  • Autonomous driving reinforcement learning method and device fusing human driving prior data

    CN122596170A