An autonomous driving decision-making method with variable driving style based on imitation learning

Through the autonomous driving decision-making method based on variable driving styles based on imitation learning, data collection, basic model generators and driving style adjustments are used to solve the problem that autonomous driving vehicles are difficult to adapt to and change their driving style in complex traffic scenarios, and the adaptation of diversified driving styles and efficient utilization of resources are achieved.

CN117799637BActive Publication Date: 2025-05-09YANGTZE DELTA REGION INST (QUZHOU) UNIV OF ELECTRONIC SCI & TECH OF CHINA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311710703.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-13
Publication Date
2025-05-09
Estimated Expiration
2043-12-13

AI Technical Summary

Technical Problem

Existing autonomous driving decision-making technologies are difficult to adapt to and change driving styles in complex and changing traffic scenarios, and the existing technology requires retraining the original model to adjust the driving style, resulting in waste of computing, storage and time resources.

Method used

Using a variable driving style autonomous driving decision-making method based on imitation learning, incremental learning and rapid fine-tuning are achieved through data collection, basic model generators and driving style adjustments. This method uses the converter architecture and beta distribution, combined with the attention environment perception framework, dynamically adjusts the attention to features, and achieves a diverse driving style.

Benefits of technology

It realizes the adaptation of diversified driving styles of autonomous driving vehicles in complex traffic scenarios, reduces the consumption of computing, storage and time resources, and improves the adaptability and efficiency of autonomous driving decision algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117799637B_ABST
    Figure CN117799637B_ABST
Patent Text Reader

Abstract

The present invention discloses an automatic driving decision method with variable driving style based on imitation learning. The present invention proposes an imitation learning model that supports incremental learning. The basic imitation learning model can be trained with expert data of multiple driving styles and achieve performance that exceeds that of experts in imitating expert driving behaviors. Through incremental learning, the model can continuously improve performance based on continuously accumulated data without the need to retrain the entire model each time. This will enable automatic driving vehicles to update model parameters more promptly to adapt to new road scenarios and traffic conditions. The present invention can quickly fine-tune the basic automatic driving decision model based on imitation learning according to different driving style preferences, while reducing a large amount of computing, storage and time overhead. The present invention is based on an attention environment perception framework to enable the automatic driving decision algorithm to focus on more useful key information in changing traffic scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of automobile automatic driving, and in particular to an automatic driving decision method with a variable driving style based on imitation learning. Background Art

[0002] Existing research on autonomous driving decision-making can be roughly divided into two categories. They are rule-based autonomous driving decision-making algorithms and data-driven autonomous driving decision-making algorithms. Data-driven autonomous driving decision-making algorithms can also be subdivided into deep learning-based autonomous driving decision-making algorithms, imitation learning-based autonomous driving decision-making algorithms, and reinforcement learning-based autonomous driving decision-making algorithms. Rule-based autonomous driving decision-making algorithms are a traditional method that relies on predefined rules and logic to guide the decision-making process of autonomous driving vehicles. These rules are usually manually formulated to ensure that the vehicle behaves safely and legally in various traffic scenarios. Data-driven autonomous driving decision-making algorithms are a method that uses deep learning, imitation learning, or reinforcement learning techniques to learn driving decision-making patterns from a large amount of perception data to help autonomous vehicles make decisions.

[0003] Rule-based autonomous driving decision-making algorithm: The rule-based autonomous driving decision-making algorithm forms a simple speed decision-making strategy by manually designing rules. Its advantages lie in its interpretability and controllability. The setting of rules can ensure that the vehicle's behavior complies with regulations and safety standards. However, this approach also has some challenges, because traffic scenarios are complex and changeable, and rules in different situations may conflict with each other, requiring careful rule design and management. In the complex and changeable real-world traffic environment, rule-based autonomous driving decision-making algorithms rely on established rules and are far from sufficient to cope with various real-world tasks.

[0004] Data-driven autonomous driving decision-making algorithms: The advantages of data-driven autonomous driving decision-making algorithms are their adaptability and ability to cope with complex and changing traffic situations. Autonomous driving decision-making algorithms based on deep learning are trained by collecting a large amount of driving data to give decision signals to control the efficient, safe and comfortable driving of the intelligent agent. Autonomous driving decision-making algorithms based on reinforcement learning are designed to guide the intelligent agent to make decisions in complex and changing traffic by designing reward functions. With the emergence of deep learning, imitation learning methods have also received attention. Autonomous driving decision-making algorithms based on imitation learning use deep neural networks to approximate complex strategies and have shown promising results in the field of autonomous driving. Imitation learning, often referred to as learning from demonstrations, aims to enable intelligent agents to imitate the behavior of experts without explicit rewards. The idea is to bridge the gap between imitation supervised learning and reinforcement learning by leveraging expert demonstrations.

[0005] Existing research can well solve the maneuver decision-making problem in a single traffic scenario, but ignores the necessity of changing driving style in complex and changing traffic scenarios. Summary of the invention

[0006] The purpose of the present invention is to provide an automatic driving decision-making method with a variable driving style based on imitation learning.

[0007] To achieve the above object, the present invention is implemented according to the following technical solutions:

[0008] The present invention includes data collection, base model generator and driving style adjustment;

[0009] The data collection includes environmental data from different sensors and action distributions demonstrated by reinforcement learning experts, the action distributions including the distributions of acceleration and steering angles; the environmental data is encoded separately by the data acquisition module for the features of lanes, vehicles and traffic information, and then the collected features are integrated; the environmental data and action distributions are integrated to form a synthetic data set customized for offline training;

[0010] The basic model generator is divided into an encoder and a decoder. The encoder uses a transformer architecture to assign weights to environmental data. In the decoder, the imitation learning agent receives weighted environmental data, expert-driven waypoints and action distribution as input, and outputs an action distribution strategy for future time steps.

[0011] The driving style adjustment adjusts the action distribution of the reinforcement learning expert according to the basic model generator, and adjusts the α and β values ​​of the beta distribution output by the basic model generator to change the driving style of the reinforcement learning expert for subsequent fine-tuning of the imitation learning model designed by the present invention.

[0012] Further, the data collection comprises the following steps:

[0013] S11: The data from GNSS and IMU sensors are processed using a particle filter, followed by extraction of D u Characteristics of conventional vehicles ahead in range and D u Extract traffic signals and road signs from image data within the range;

[0014] S12: Coding of vehicles and pedestrians: The characteristics of vehicles and pedestrians are represented by F c and F p :

[0015] Δf t (C i , A) = (vel, θ, dis, θ vel )

[0016] Among them, vel, θ, dis, θvel Represents vehicle C i The relative speed, angle, longitudinal distance and speed angle between the autonomous driving vehicle A and the front distance D on the three lanes u The nine nearest conventional vehicles in the range cover the three key areas in front, namely the left front (C1, C2, C3), the front (C4, C5, C6) and the right front (C7, C8, C9). Similarly, there are nine pedestrians distributed in these three areas. When there are not enough vehicles in an area, a feature set (vel, θ, dis, θ vel ) are set to zero and these features are encoded as follows:

[0017] Enc A =Φ A (f A ;w A )

[0018] Enc C =Φ C (Δf t (C i , A); w C )

[0019] Enc P =Φ P (Δf t (P i , A); w P )

[0020] Among them, Φ A , Φ C and Φ P is the embedding layer; w A 、w C and w P is the corresponding parameter; Enc A 、Enc C and Enc P Encode the characteristics of the agent, vehicle, and pedestrian, and use these parameters to calculate the relative displacement at the subsequent time step t;

[0021] S13: Traffic rules encoding: Extracting traffic rules F R features, including traffic lights, stop lines, yield signs, speed limits, and zebra crossings, are represented by f tl = (status, dis) describes the detailed characteristics of the traffic light, where status represents the state of the traffic light tl, and dis represents the relative distance between tl and the autonomous driving vehicle A; the entire traffic rule F R Enc embedded in the network R In order to capture the relationship among the traffic rules features;

[0022] S14: Waypoint Encoding: The agent receives a set of waypoints arranged in order, which represent its forward route. At each time step t, the global routing algorithm provides 16 waypoints and embeds them as Enc W ;

[0023] S15: Use Beta distribution to model the output behavior of the agent, which includes two distributions: the output distribution of steering angle and the output distribution of acceleration. Fine-tune these action distributions to produce different driving styles. Then, use filters to remove expert experience with low completion rates to ensure high-quality input data for the reinforcement learning agent.

[0024] Furthermore, the encoder in the basic model generator has a perception layer and a state layer. By encoding the training data collected at time t, the perception layer recognizes roads and traffic signs and converts them into env_feature through the transformer architecture; the state consists of instructions, speed and target heading points and is processed into state_feature through a multi-layer perceptron; each transformer block combines a multi-head self-attention mechanism and a feedforward network (FFN), supplemented by feedforward connections and layer normalization; Enc A 、Enc C 、Enc P 、Enc R and Enc W The features are embedded into a 64-dimensional vector through the embedding layer E, and then these vectors are combined to form a matrix E, which is defined as follows:

[0025] E=[Enc A ,Enc c ,Enc P ,Enc R ,Enc W ]

[0026] The matrix E is linearly projected to form the key, query, and value of the Transformer architecture input, and the attention weight matrix A and attention output V″ are calculated:

[0027]

[0028] The matrix V" is transformed using a feed-forward network FFN(V"), and then a multi-layer perceptron (MLP) combines these features and measurements into an end-to-end decision feature:

[0029] action_feature=MLP(env_feature, state_feature)

[0030] To facilitate future decision making, waypoints are passed through a waypoint multilayer perceptron:

[0031] wp_feature = MLP(waypoint)

[0032] The decoder uses the features from the encoder to predict the future actions of the vehicle. Initially, it obtains three features from the encoder, namely waypoint features, action features, and environment features. The policy network processes the action features to obtain the current output action distribution α t and β t ; Using the action distribution α t and β t And the action feature action_feature input gated recurrent unit (GRU) can determine the action hidden feature h of the next time step t+1 ; The formula for this process is as follows:

[0033] h t+1 =GRU t (Concat(action_feature t , α t , β t ), h t )

[0034] The training of the gated recurrent unit (GRU) integrates the action distribution and waypoint features, and generates output to the output layer through a weighted multi-layer perceptron. The newly generated h t+1 Connect with wp_feature, then dot-multiply with env_feature to generate the action feature action_feature of the next time step t+1 :

[0035] weight_feature t+1 =Concat(h t+1 ,wp_feature)⊙env_feature

[0036] action_feature t+1 =MLP(Concat(weight_feature t+1 ,h t+1 ))

[0037] action_feature t+1 After passing through the policy network, the action distribution α at time t+1 is generated t+1 and β t+1 The imitation learning agent can roughly predict the action distribution output that should be taken in the next four steps, that is, (α t+1 , β t+1 ) to (α t+4 , β t+4)’s entire sequence of actions.

[0038] Furthermore, when the beta distribution is used to calculate the loss function in the basic model generator, the relative entropy (KL divergence) of the beta distribution is used as the negative loss component.

[0039]

[0040] Among them, Beta (α, β) represents the beta distribution represented by each prediction distribution parameter; relative entropy (KL divergence) is used to quantify the similarity between the prediction control distribution and the expert control distribution, and the beta distribution output by the expert is expressed as express;

[0041] Features are extracted from the policy and value networks of the reinforcement learning expert and integrated using the L2 loss:

[0042]

[0043] Where V and Represent the predicted value features and the value features obtained from the reinforcement learning expert, respectively. Similarly, F and Represent the predicted policy network features and the features obtained from the expert’s policy network experience, respectively. In the calculation process, the weighted average of the policy network features spanning the next four steps is used; W v and W f is a hyperparameter used to calculate the weighted loss;

[0044] The overall losses are as follows:

[0045] Loss=Loss action +Loss feature .

[0046] The beneficial effects of the present invention are:

[0047] The present invention is an automatic driving decision method with variable driving style based on imitation learning. Compared with the prior art, the present invention proposes an imitation learning model that supports incremental learning. The basic imitation learning model can be trained with expert data of multiple driving styles and achieve performance that exceeds that of experts in imitating expert driving behaviors. Through incremental learning, the model can continuously improve performance based on continuously accumulated data without retraining the entire model each time. This will enable automatic driving vehicles to update model parameters more timely to adapt to new road scenarios and traffic conditions. The present invention can quickly fine-tune the basic automatic driving decision model based on imitation learning according to different driving style preferences, while reducing a large amount of computing, storage and time overhead. The present invention is based on an attention environment perception framework to enable the automatic driving decision algorithm to focus on more useful key information in changing traffic scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 It is the overall framework of the present invention;

[0049] Figure 2 Input feature example graph for the present invention;

[0050] Figure 3 Output example diagram for the present invention. DETAILED DESCRIPTION

[0051] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments. The illustrative embodiments and descriptions of the present invention are used to explain the present invention but are not intended to limit the present invention.

[0052] The present invention solves the following three technical problems:

[0053] (1) Although the existing autonomous driving decision-making technology based on imitation learning can well complete the autonomous driving task in a single scenario, it cannot adapt to complex and changeable traffic scenarios. For example, in smooth rural road traffic with few cars, autonomous driving cars should consider driving efficiency and comfort more, while in busy and congested urban road traffic, autonomous driving cars should consider driving safety more to prevent accidents. Therefore, the present invention proposes an imitation learning model that supports incremental learning. The basic imitation learning model can be trained with expert data of multiple driving styles and achieve performance that exceeds that of experts in imitating expert driving behavior. Through incremental learning, the model can continuously improve its performance based on the continuously accumulated data without the need to retrain the entire model each time. This will enable autonomous driving vehicles to update model parameters more promptly to adapt to new road scenarios and traffic conditions.

[0054] (2) If the existing technology needs to adjust the driving style, the original model needs to be retrained, which will occupy a lot of computing resources, storage resources and time resources. Even more unfortunately, if the original data set is lost, retraining may even be impossible. In order to solve the above technical problems and speed up training efficiency, the present invention proposes a method for quickly fine-tuning the basic imitation learning-based autonomous driving decision model according to different driving style preferences to solve this technical problem and simultaneously reduce a lot of computing, storage and time overhead.

[0055] (3) Autonomous vehicles need to pay attention to key information based on their state and driving style. For example, during cruising, autonomous vehicles should pay more attention to signs related to cruising in the lane, while at intersections, traffic lights, yield or stop signs, and pedestrians crossing the road become crucial. The present invention aims to solve the above technical problems and proposes an attention-based environment perception framework to enable autonomous driving decision algorithms to pay attention to more useful key information in changing traffic scenarios.

[0056] like Figure 1 As shown, the present invention includes three components: data collection, base model generator and driving style adjustment.

[0057] (1) Data acquisition: The present invention integrates environmental data from different sensors (including cameras, GPS, and IMU). The data acquisition module first encodes the features of lanes, vehicles, and traffic information separately, and then integrates these multimodal features. In addition to environmental data, the present invention also collects the motion distribution demonstrated by reinforcement learning experts, including the distribution of acceleration and steering angle. By integrating environmental data and motion distribution, a synthetic dataset tailored for offline training is formed.

[0058] ① Raw data: In order to obtain the position and direction of the intelligent agent, the present invention uses a particle filter to process the data from the GNSS and IMU sensors. Then, using object detection technology, we extract D u In addition, the present invention also has the characteristics of conventional vehicles ahead in the range. u Traffic signals and road signs can be extracted from image data within a certain range.

[0059] ② Coding of vehicles and pedestrians: After collecting the raw data, the present invention will form Figure 3 The features mentioned in the above. The features of vehicles and pedestrians are represented as F c and F p The present invention is based on vehicle C i For example, Δf t (C i , A) = (vel, θ, dis, θ vel ), among them, vel, θ, dis, θ vel Represents vehicle C i The relative speed, angle, longitudinal distance and speed angle between the vehicle A and the self-driving vehicle A. The present invention selects the front distance D on the three lanes uThe nine nearest conventional vehicles in the range cover three key areas in front, namely, the left front (C1, C2, C3), the front (C4, C5, C6) and the right front (C7, C8, C9). Similarly, there are nine pedestrians distributed in these three areas. When there are not enough vehicles in one area, the present invention uses a feature set (vel, θ, dis, θ vel ) is set to zero. The present invention encodes these features as follows:

[0060] Enc A =Φ A (f A ;w A )

[0061] Enc C =Φ C (Δf t (C i , A); w C )

[0062] Enc P =Φ P (Δf t (P i , A); w P )

[0063] Among them, Φ A , Φ C and Φ P is the embedding layer; w A 、w C and w P is the corresponding parameter; Enc A 、Enc C and Enc P Encode the characteristics of the agent, vehicle, and pedestrian, and use these parameters to calculate the relative displacement at the subsequent time step t.

[0064] ③ Traffic rules encoding: The present invention extracts traffic rules F R The features of traffic lights include traffic lights, stop lines, yield signs, speed limits and zebra crossings. Here, the present invention mainly describes the detailed features of traffic lights, using f tl =(status, dis), where status represents the state of the traffic light tl, and dis represents the relative distance between tl and the autonomous driving vehicle A. R Embedded in the Enc of the network to be introduced later R In this way, the subtle relationships among traffic rules characteristics can be captured.

[0065] ④Waypoint encoding: To guide the driving path of the imitation learning agent, the agent will receive a set of waypoints arranged in order, which represent its forward route. At each time step t, the advanced global routing algorithm will provide 16 waypoints. And embedded as Enc W .

[0066] ⑤ Expert experience: In imitation learning, action distribution plays a key role in demonstrating expert behavior, because action distribution can directly guide the imitation learning agent to select steering angle and acceleration. In order to capture the inherent variability in driving behavior, the present invention uses Beta distribution to model the output behavior of the agent, which contains two distributions: the output distribution of steering angle and the output distribution of acceleration. Fine-tuning these action distributions can produce different driving styles, such as preferring efficiency and caution. Then, the present invention uses filters to eliminate expert experience with low completion rates to ensure high-quality input data for the reinforcement learning agent.

[0067] (2) Basic model generator: The model architecture of the present invention is as follows Figure 2 After collecting data, the present invention uses imitation learning to learn expert behavior using the current environment information in the basic model generator, which is divided into an encoder and a decoder. First, the encoder uses a transformer architecture to assign weights to the environment data, where the schematic diagram of the collected environment data is shown in Figure 3 Second, in the decoder, the imitation learning agent receives weighted environment data, expert-driven waypoints and action distributions as input, and outputs an action distribution policy for future time steps.

[0068] ① Encoder: The perception layer and state layer are encoded by the training data collected at time t. The perception layer identifies key driving elements such as roads and traffic signs and converts them into env_feature through the transformer architecture. The state consists of instructions, speed, and target heading points and is processed into state_feature through a multi-layer perceptron. Each transformer block combines a multi-head self-attention mechanism and a feedforward network (FFN), supplemented by feedforward connections and layer normalization. This structure allows the model to dynamically adjust its focus on features based on different environmental conditions. Enc A 、Enc C 、Enc P 、Enc R and Enc W The features are embedded into a 64-dimensional vector through the embedding layer E. These vector combinations are then defined as a matrix E, which is defined as follows:

[0069] E=[Enc A ,Encc ,Enc P ,Enc R ,Enc W ]

[0070] The matrix E is linearly projected to form the key, query, and value of the Transformer architecture input. The attention weight matrix A and attention output V″ are calculated:

[0071]

[0072] The matrix V" is transformed using a feed-forward network FFN(V") . Then, a multi-layer perceptron (MLP) combines these features and measurements into an end-to-end decision feature:

[0073] action_feature=MLP(env_feature, state_feature)

[0074] To facilitate future decision making, waypoints are passed through a waypoint multilayer perceptron:

[0075] wp_feature = MLP(waypoint)

[0076] Compared with traditional methods, the method of the present invention not only improves the learning efficiency of the imitation learning agent in autonomous tasks and reduces the model complexity, but also uses the attention score to weigh the sensor input according to the current driving situation. This method not only reduces the computational cost but also enhances the generalization of the model.

[0077] ②Decoder: Use the features from the encoder to predict the future actions of the vehicle. Initially, three features from the encoder are obtained, namely waypoint features, action features, and environment features. The policy network processes the action features to obtain the current output action distribution α t and β t Using the action distribution α t and β t And the action feature action_feature input gated recurrent unit (GRU) can determine the action hidden feature h of the next time step t+1 The formula for this process is as follows:

[0078] h t+1 =GRU t (Concat(action_feature t , α t , β t ), h t )

[0079] In order to stabilize the prediction task and reduce the impact of a single behavioral decision on the overall trajectory of the autonomous vehicle, the present invention integrates the action distribution and waypoint features in the GRU training, and generates output to the output layer through a weighted multilayer perceptron. Here, the newly generated h t+1 Connect with wp_feature, then dot-multiply with env_feature to generate the action feature action_feature of the next time step t+1 :

[0080] weight_feature t+1 =Concat(h t+1 ,wp_feature)⊙env__feature

[0081] action_feature t+1 =MLP(Concat(weight_feature t+1 ,h t+1 ))

[0082] action_feature t+1 After passing through the policy network, the action distribution α at time t+1 is generated t+1 and β t+1 .

[0083] The imitation learning agent can roughly predict the action distribution output that should be taken in the next four steps, that is, (α t+1 , β t+1 ) to (α t+4 , β t+4 )’s entire sequence of actions.

[0084] ③ Loss function: When calculating the loss function using the Beta distribution, the standard practice is to use the relative entropy (KL divergence) of the Beta distribution as the negative loss component. The main goal is to ensure that the result distribution of the model is consistent with the required expert distribution by minimizing the relative entropy (KL divergence). This careful alignment helps to achieve consistency between the learning strategy and the expert paradigm.

[0085]

[0086] Among them, Beta (α, β) represents the beta distribution represented by each prediction distribution parameter. Relative entropy (KL divergence) is used to quantify the similarity between the prediction control distribution and the expert control distribution. The beta distribution output by the expert is expressed as express.

[0087] In order to enhance the agent's ability to identify the current state, we extracted features from the policy and value networks of the reinforcement learning expert and integrated them through the L2 loss:

[0088]

[0089] Where V and Represent the predicted value features and the value features obtained from the reinforcement learning expert, respectively. Similarly, F and Represent the predicted policy network features and the features obtained from the expert’s policy network experience, respectively. In the calculation process, we use the weighted average of the policy network features spanning four steps into the future. v and W f is a hyperparameter used to calculate the weighted loss.

[0090] The overall losses are as follows:

[0091] Loss=Loss action +Loss feature

[0092] (3) Driving style adjustment: In order to obtain a variety of driving styles, the present invention adjusts the action distribution of the reinforcement learning expert according to the basic model generator, and adjusts the α and β values ​​of the beta distribution output by it to change the driving style of the reinforcement learning expert for subsequent fine-tuning of the imitation learning model designed by the present invention. Specifically, in order to fine-tune the robust initial performance of the basic model, only a small amount of enhanced training data and additional training cycles are required to adjust the driving style, thereby ensuring its versatility without affecting safety.

[0093] In the present invention, two driving styles are focused on, namely, efficient driving style and cautious driving style.

[0094] (1) The core objective of the present invention is to ensure that the autonomous driving vehicle can adaptively adjust the driving style of the autonomous driving decision algorithm according to the current traffic scenario of the autonomous driving vehicle while maintaining safety. To achieve this goal, the present invention introduces a flexible autonomous driving decision framework based on imitation learning, which has the unique ability to generate decision algorithms with diverse driving styles that adapt to different traffic scenarios. The significance of this innovative framework is that it allows the autonomous driving system to continuously evolve by learning and imitating experts and adapting to various traffic scenarios. It can dynamically adjust driving decisions based on multiple environmental factors such as road conditions to achieve more intelligent, safe and efficient autonomous driving behavior. The framework has a wide range of applications and can be applied to various driving scenarios such as urban roads and highways. It is expected to provide more options and flexibility to meet different user and market needs, from smooth cruising to maneuvering in intense urban traffic. The present invention will help improve the adaptability of autonomous driving vehicles, enabling them to exhibit a richer driving style in various traffic scenarios, thereby better meeting driving needs in different scenarios while ensuring safety. This will hopefully promote the development of autonomous driving technology and enable it to be better integrated into daily traffic and travel.

[0095] (2) The present invention applies fine-tuning technology to autonomous driving scenarios and proposes a method for quickly generating autonomous driving decision algorithm models of different styles based on the fine-tuning idea. This innovative method can efficiently perform fine-tuning on the basis of the original basic model to quickly generate autonomous driving algorithms that adapt to different driving styles, thereby enabling autonomous driving vehicles to operate safely, efficiently and comfortably in complex and changeable traffic scenarios.

[0096] (3) The model designed by the present invention uses a transformer network to weigh environmental inputs. By introducing the transformer network, the model can more effectively process a large amount of complex environmental information, such as road data, traffic signs, vehicle positions, and pedestrian status, and accurately extract and focus on key information. This precise information processing helps autonomous vehicles perceive and understand the surrounding environment more accurately, enabling them to make better decisions to cope with various complex traffic situations, thereby achieving better performance and driving completion rate.

[0097] The technical solution of the present invention is not limited to the above-mentioned specific embodiments. All technical variations made according to the technical solution of the present invention fall within the protection scope of the present invention.

Claims

1. An automatic driving decision-making method with variable driving style based on imitation learning, characterized in that: Includes data collection, base model generator and driving style tuning; The data collection includes environmental data from different sensors and action distributions demonstrated by reinforcement learning experts, the action distributions including the distributions of acceleration and steering angles; the environmental data is encoded separately by the data acquisition module for the features of lanes, vehicles and traffic information, and then the collected features are integrated; the environmental data and action distributions are integrated to form a synthetic data set customized for offline training; The basic model generator is divided into an encoder and a decoder. The encoder uses a transformer architecture to assign weights to environmental data. In the decoder, the imitation learning agent receives weighted environmental data, expert-driven waypoints and action distributions as inputs, and outputs action distribution strategies for future time steps. The driving style adjustment adjusts the action distribution of the reinforcement learning expert according to the basic model generator, and adjusts the α and β values ​​of the beta distribution output by the basic model generator to change the driving style of the reinforcement learning expert for subsequent fine-tuning of the imitation learning model designed by the present invention; The data collection includes the following steps: S11: The data from GNSS and IMU sensors are processed using a particle filter, followed by extraction of D u Characteristics of conventional vehicles ahead in range and D u Extract traffic signals and road signs from image data within the range; S12: Coding of vehicles and pedestrians: The characteristics of vehicles and pedestrians are represented by F c and F p : Δf t (C i ,A)=(υel,θ,dis,θ υel ) Among them, υel, θ, dis, θ υel Represents vehicle C i The relative speed, angle, longitudinal distance and speed angle between the autonomous driving vehicle A and the front distance D on the three lanes u The nine nearest conventional vehicles in the range cover the three key areas in front, namely the left front (C1, C2, C3), the front (C4, C5, C6) and the right front (C7, C8, C9). Similarly, there are nine pedestrians distributed in these three areas. When there are not enough vehicles in an area, a feature set (υel, θ, dis, θ υel ) are set to zero and these features are encoded as follows: Enc A =Φ A (f A ;w A ) Enc C =Φ C (Δf t (C i ,A);w C ) Enc P =Φ P (Δf t (P i ,A);w P ) Among them, Φ A , Φ C and Φ P is the embedding layer; w A 、w C and w P is the corresponding parameter; Enc A 、Enc C and Enc P Encode the characteristics of the agent, vehicle, and pedestrian, and use these parameters to calculate the relative displacement at the subsequent time step t; S13: Traffic rules encoding: Extracting traffic rules F R features, including traffic lights, stop lines, yield signs, speed limits, and zebra crossings, are represented by f tl = (status, dis) describes the detailed characteristics of the traffic light, where status represents the state of the traffic light tl, and dis represents the relative distance between tl and the autonomous driving vehicle A; the entire traffic rule F R Enc embedded in the network R In order to capture the relationship among the traffic rules features; S14: Waypoint Encoding: The agent receives a set of waypoints arranged in order, which represent its forward route. At each time step t, the global routing algorithm provides 16 waypoints and embeds them as Enc W ; S15: Use Beta distribution to model the output behavior of the agent, which includes two distributions: the output distribution of steering angle and the output distribution of acceleration. Fine-tune these action distributions to produce different driving styles. Then, use filters to remove expert experience with low completion rates to ensure high-quality input data for the reinforcement learning agent.

2. The automatic driving decision-making method with variable driving style based on imitation learning according to claim 1, characterized in that: The encoder in the basic model generator has a perception layer and a state layer. The perception layer recognizes roads and traffic signs and converts them into enυ_feature through the converter architecture by encoding the training data collected at time t; the state consists of instructions, speed and target heading points and is processed into state_feature through a multi-layer perceptron; each converter block combines a multi-head self-attention mechanism and a feedforward network, supplemented by feedforward connections and layer normalization; Enc A 、Enc C 、Enc P 、Enc R and Enc W The features are embedded into a 64-dimensional vector through the embedding layer E, and then these vectors are combined to form a matrix E, which is defined as follows: E=[Enc A, Enc c ,Enc P ,Enc R ,Enc W ] The matrix E is linearly projected to form the key, query, and value of the Transformer architecture input, and the attention weight matrix A and attention output V″ are calculated: The matrix V″ is transformed using the feed-forward network FFN(V″), and then the multi-layer perceptron combines these features and measurements into end-to-end decision features: action_feature=MLP(enυ_feature, state_feature) To facilitate future decision making, waypoints are passed through a waypoint multilayer perceptron: wp_feature = MLP(waypoint) The decoder uses the features from the encoder to predict the future actions of the vehicle. Initially, it obtains three features from the encoder, namely waypoint features, action features, and environment features. The policy network processes the action features to obtain the current output action distribution α t and β t ; Using the action distribution α t and β t And the action feature action_feature input gated recurrent unit can determine the action hidden feature h of the next time step t+1 ; The formula for this process is as follows: h t+1 =GRU t (Concat(action_feature t ,a t ,b t ),h t ) The Gated Recurrent Unit (GRU) training integrates the action distribution and waypoint features, and generates output to the output layer through a weighted multi-layer perceptron. The newly generated h t+1 Connect with wp_feature, then multiply with enυ_feature to generate the action feature action_feature of the next time step t+1 : weight_feature t+1 =Concat(h t+1 ,wp_feature)⊙enυ_featureaction_feature t+1 =MLP(Concat(weight_feature t+1 ,h t+1 )) action_feature t+1 After passing through the policy network, the action distribution α at time t+1 is generated t+1 and β t+1 ; The imitation learning agent can roughly predict the action distribution output that should be taken in the next four steps, that is, (α t+1 , β t+1 ) to (α t+4 , β t+4 )’s entire sequence of actions.

3. The automatic driving decision-making method with variable driving style based on imitation learning according to claim 2, characterized in that: When the Beta distribution is used to calculate the loss function in the base model generator, the relative entropy of the Beta distribution is used as the negative loss component. Among them, Beta (α, β) represents the beta distribution represented by each prediction distribution parameter; relative entropy is used to quantify the similarity between the prediction control distribution and the expert control distribution, and the beta distribution output by the expert is express; Features are extracted from the policy and value networks of the reinforcement learning expert and integrated using the L2 loss: Where V and Represent the predicted value features and the value features obtained from the reinforcement learning expert, respectively. Similarly, F and Represent the predicted policy network features and the features obtained from the expert’s policy network experience, respectively. In the calculation process, the weighted average of the policy network features spanning the next four steps is used; W υ and W f is a hyperparameter used to calculate the weighted loss; The overall losses are as follows: Loss=Loss action +Loss feature 。

Citation Information

Patent Citations

  • Driving control method based on self-supervised imitation learning

    CN116353623A