Automatic driving end-to-end method based on multi-modal data

Through the combination of multimodal encoder and reinforcement learning algorithm, the problems of low sample efficiency and poor training stability in urban environments are solved, and efficient autonomous driving is achieved in complex environments.

CN120123744APending Publication Date: 2025-06-10SOUTHEAST UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510277031.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-12-30
Filing Date
2025-03-10
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Existing end-to-end autonomous driving methods based on reinforcement learning face the problems of low sample efficiency and poor training stability in urban environments, and decoupling of perceptual training from driving strategy training may lead to suboptimal strategies.

Method used

The multimodal encoder is used to independently process the navigation map, images, track points and vehicle measurement data, and control commands are output through the SAC algorithm and the PID controller, and the importance of traffic light information is enhanced through auxiliary loss branches.

Benefits of technology

Improve the sample efficiency and training stability of autonomous driving in complex urban environments, ensuring that vehicles can drive autonomously safely and effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123744A_ABST
    Figure CN120123744A_ABST
Patent Text Reader

Abstract

The invention provides an automatic driving end-to-end method based on multi-modal data. The system architecture of the method comprises three main components: a multi-modal encoder, a reinforcement learning algorithm and an auxiliary detection branch. The method comprises the following steps: firstly, inputting a navigation map, an image, track points and vehicle information into a multi-mode encoder by a system, independently processing according to the characteristics of different input data, and mapping original data into high-dimensional feature information; in order to improve comprehensive representation of input on environment information, a traffic light detection auxiliary branch is used, abstract features are guided to be generated, and the perceptual information representation capability of input features is improved. And finally, splicing different input potential representations to form a unified feature vector as the input of a SAC (Soft Actor-Critic) algorithm, then outputting a control command by using the SAC algorithm and a PID (Proportion Integration Differentiation) controller, and outputting a value function by using a Q network. Through the method, the vehicle can be automatically driven in a complex urban road environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of autonomous driving, and specifically relates to an end-to-end method for autonomous driving based on multimodal data. Background Art

[0002] In recent years, significant progress has been made in autonomous driving technology, whose core tasks include environmental perception and driving strategies. Urban driving poses the greatest challenge due to its complexity and unpredictability. To address this challenge, researchers have turned to information fusion techniques and end-to-end methods such as imitation learning (IL) and reinforcement learning (RL). Although IL can learn from expert data, it is limited by the mismatch between expert performance and data distribution. RL learns by interacting with the environment and is not restricted by experts, but has low sample efficiency. Existing end-to-end methods for autonomous driving based on reinforcement learning directly map image data to driving actions, facing problems of sample efficiency and training stability. Such methods are mostly used for tasks with less challenging environmental perception, while urban autonomous driving needs to process a large amount of irrelevant information. Current methods decouple perception training from driving strategy training, which increases stability but may lead to suboptimal strategies. The present invention successfully realizes the simultaneous training of the encoder and the policy network in vision-based urban autonomous driving, aiming to improve the autonomous driving ability in complex urban environments.

[0003] The differences between this application and the prior art are as follows:

[0004] Technical comparison with the patent CN104880182A "An end-to-end autonomous driving method for accelerating reinforcement learning independent of irrelevant information"

[0005] For the patent CN104880182A, the image information collected by the front camera is used, and the original image is used as the input of the autoencoder. While we use navigation maps, images, trajectory points, and vehicle measurement data as the input of the multimodal encoder.

[0006] The patent CN104880182A performs inference and training on an FPGA-based acceleration platform to obtain control commands. While we use the SAC (Soft Actor-Critic) algorithm and combine a PID controller to output control commands.

[0007] Technical comparison with the patent CN116279585A "An end-to-end autonomous driving method, system, and electronic device"

[0008] Patent CN116279585A obtains camera data and lidar data from a camera and a lidar installed on a vehicle respectively, encodes the camera data and lidar data, and obtains encoder features. While we use navigation maps, images, trajectory points, and vehicle measurement data as the input of the multi-modal encoder, and process the data through the multi-modal encoder to obtain encoder features.

[0009] Technical comparison with patent CN117710917A, "End-to-end multi-modal multi-task autonomous driving perception method and device based on long and short time series hybrid coding"

[0010] For the multi-frame input of vehicle-mounted sensors in patent CN117710917A, multiple coding networks are used to extract features of long and short time series; then, the extracted features are subjected to temporal fusion based on the attention mechanism, cross-modal fusion is performed, and a BEV feature map is generated for each task. Then, decoders for different tasks are used to obtain the prediction results for each task. While we use navigation maps, images, trajectory points, and vehicle measurement data as the input of the multi-modal encoder, and independently process the input data through different encoders, and splice their latent representations into a feature vector after obtaining them respectively. Summary of the Invention

[0011] To solve the above technical problems, the present invention proposes an end-to-end method for autonomous driving based on multi-modal data. Through this method, a vehicle can perform autonomous driving in a complex urban road environment.

[0012] To achieve the above object, the technical solution adopted by the present invention is:

[0013] An end-to-end method for autonomous driving based on multi-modal data, characterized by comprising the following steps:

[0014] Step 1: Process the input data through a multi-modal encoder, including independent processing of navigation maps, images, trajectory points, and vehicle measurement data.

[0015] Step 2: Based on the constructed auxiliary loss branch, enhance the importance of traffic light information in the latent representation of the image, complete the classification task of traffic lights, and improve the perceptual information representation ability of the input features.

[0016] Step 3: Process the spliced feature vector based on the SAC algorithm, and combine a PID controller to output a control command.

[0017] As a further improvement of the present invention, the specific steps of the first step include the following steps:

[0018] (1.1) Perform data augmentation on the input image data. The image encoder consists of a convolutional neural network and is used to represent low-dimensional sensor data as high-dimensional features. The final processing flow can be expressed by the formula:

[0019]

[0020] where f i is the image encoder, aug represents the applied data augmentation technique, is the feature representation of the image after encoding, represents the set of image data at past time steps;

[0021] (1.2) The path point encoder based on the WayConv1D method utilizes the input two-dimensional geometric structure by applying one-dimensional convolution with a 2×2 kernel to the two-dimensional coordinates of the next N path points. The output of the one-dimensional convolution is then flattened and processed by a multi-layer perceptron. This process is summarized as:

[0022]

[0023] where f w is the WayConv1D path point encoder, is the feature representation of the path point data after encoding, W t represents the current step;

[0024] (1.3) Based on the vehicle measurement encoder, a multi-layer perceptron MLP is used to process the vehicle measurement data and generate a latent representation. The specific process can be expressed by the formula:

[0025]

[0026] where f v is the vehicle measurement encoder, is the feature representation of the vehicle measurement data after encoding, represents the set of vehicle measurement data at past time steps;

[0027] (1.4) Based on the standard map encoder, a graph convolutional network GCN is used to process the standard map data and generate a latent representation. The specific process can be expressed by the formula:

[0028]

[0029] where f u is the vehicle measurement encoder, is the feature representation of the vehicle measurement data after encoding, represents the set of standard map data at past time steps;

[0030] (1.5) The input data is independently processed by different encoders, and after obtaining their respective latent representations, they are concatenated into a feature vector.

[0031] As a further improvement of the present invention, the specific steps of step two are as follows:

[0032] (3.1) Construct an auxiliary task: traffic light classification, add a traffic light encoder, and define classification labels.

[0033] (3.2) Collect data and calculate the loss. Sample from the experience replay buffer, use the traffic light decoder to classify the sampled images for traffic lights, and then calculate the cross-entropy loss. The calculation formula is as follows:

[0034]

[0035] Where, is the cross-entropy loss, B is the batch size, C is the number of classes, x corresponds to the output value of the traffic light decoder, and y is the true label.

[0036] (3.3) In the backpropagation stage, optimize and update the parameters of the traffic light decoder and the image encoder through the auxiliary loss.

[0037] (3.4) Monitor the performance when dealing with traffic light-related situations during the training process, and adjust the hyperparameters of the auxiliary task as needed.

[0038] As a further improvement of the present invention, the specific steps of step three are as follows:

[0039] (2.1) Receive the feature vector generated in step two, pass through the actor policy and two Q functions, and find the optimal policy through value function update, policy update, and temperature parameter learning. This process can be expressed by the formula:

[0040]

[0041]

[0042]

[0043] Where, and are two value functions, represents the learned optimal policy, i.e., the actor policy, α is the temperature parameter, γ is the discount factor, a t is the action taken at time t, r t represents the reward obtained at time t;

[0044] (2.2) Combine with a PID controller to output control commands;

[0045] (2.3) Encoder gradient blocking, by blocking the gradient propagation from the policy network to the encoder, to prevent the encoder from learning task-irrelevant features;

[0046] (2.4) Hyperparameter tuning and training, monitor the training process, and adjust hyperparameters as needed to accelerate the training speed and improve stability.

[0047] The beneficial effects of the present invention are as follows:

[0048] The present invention discloses an end-to-end method for autonomous driving based on multi-modal data. The system architecture of this method has three main components: a multi-modal encoder, a reinforcement learning algorithm, and an auxiliary detection branch. First, the system inputs the navigation map, images, trajectory points, and vehicle information into the multi-modal encoder, and processes them independently according to the characteristics of different input data, mapping the original data into high-dimensional feature information. To improve the comprehensive representation of the input to environmental information, a traffic light detection auxiliary branch is used to guide the generation of abstract features and improve the perceptual information representation ability of the input features. Finally, the latent representations of different inputs are concatenated to form a unified feature vector, which is used as the input of the SAC (Soft Actor-Critic) algorithm. Then, the SAC algorithm and a PID controller are used together to output control commands, while the Q network is responsible for outputting the value function. Through this method, the vehicle can perform autonomous driving in a complex urban road environment. Brief Description of the Drawings

[0049] Figure 1 It is the overall model architecture diagram of the proposed end-to-end method for autonomous driving;

[0050] Figure 2 It is the flow chart of the proposed end-to-end method for autonomous driving. Detailed Embodiments

[0051] The technical solutions of the application will be further elaborated in detail below with reference to the drawings. The described embodiments are only a part of the embodiments involved in this patent. All non-innovative embodiments made by other researchers in this field based on this embodiment fall within the protection scope of this patent.

[0052] The present invention proposes an end-to-end method for autonomous driving based on multi-modal data, as Figure 1 shown, the implementation process of the method is as Figure 2 shown, and the specific steps are as follows:

[0053] Step 1: Process the input data through a multi-modal encoder, including independently processing the input by an image encoder, a path point encoder, a vehicle measurement data encoder, and a standard map encoder.

[0054] Step 2: Based on the constructed auxiliary loss branch, enhance the importance of traffic light information in the image latent representation and complete the traffic light classification task.

[0055] Step 3: Process the concatenated feature vectors based on the SAC (Soft Actor-Critic) algorithm and output control commands in combination with a PID controller.

[0056] In a preferred embodiment of the present application, Step 1 specifically includes the following steps:

[0057] (1.1) Perform data augmentation on the input image data to improve the generalization of the model. The image encoder consists of a convolutional neural network and is used to represent low-dimensional sensor data as high-dimensional features. The final processing flow can be expressed by the formula:

[0058]

[0059] where f i is the image encoder, aug represents the applied data augmentation technique, is the feature representation of the image after encoding, represents the set of image data at past time steps;

[0060] (1.2) The path point encoder based on the WayConv1D method utilizes the input two-dimensional geometric structure by applying one-dimensional convolution with a 2×2 kernel to the two-dimensional coordinates of the next N path points. The output of the one-dimensional convolution is then flattened and processed by a multi-layer perceptron. This process can be summarized as:

[0061]

[0062] where f w is the WayConv1D path point encoder, is the feature representation of the path point data after encoding, W t represents the current step.

[0063] (1.3) Based on the vehicle measurement encoder, use a multi-layer perceptron (MLP) to process the vehicle measurement data and generate a latent representation. The specific process can be expressed by the formula:

[0064]

[0065] where f v is the vehicle measurement encoder, is the encoded feature representation of vehicle measurement data, representing the set of vehicle measurement data at past time steps.

[0066] (1.4) Based on the standard map encoder, use a graph convolutional network (GCN) to process the standard map data and generate a latent representation. The specific process can be expressed as a formula:

[0067]

[0068] where f u is the vehicle measurement encoder, is the encoded feature representation of vehicle measurement data, representing the set of standard map data at past time steps.

[0069] (1.5) Independently process the input data by different encoders, obtain their latent representations respectively, and then concatenate them into a feature vector

[0070] In a preferred embodiment of the present application, step two specifically includes the following steps:

[0071] (2.1) Construct an auxiliary task: traffic light classification, and add a traffic light encoder., and define classification labels.

[0072] (2.2) Collect data and calculate the loss. Sample from the experience replay buffer, use the traffic light decoder to classify the sampled images for traffic lights, and then calculate the cross-entropy loss. The calculation formula is as follows:

[0073]

[0074] where, is the cross-entropy loss, B is the batch size, C is the number of classes, x corresponds to the output value of the traffic light decoder, and y is the true label.

[0075] (2.3) In the backpropagation stage, optimize and update the parameters of the traffic light decoder and the image encoder through the auxiliary loss.

[0076] (2.4) Monitor the training process and adjust the hyperparameters of the auxiliary task as needed.

[0077] In a preferred embodiment of the present application, step three specifically includes the following steps:

[0078] (3.1) Receive the feature vector generated in claim 2, pass through the actor policy and two Q functions, and find the optimal policy through value function update, policy update, and temperature parameter learning. This process can be expressed as a formula:

[0079]

[0080] Among them, and are two value functions, represents the learned optimal policy, i.e., the actor policy, α is the temperature parameter, γ is the discount factor, a t is the action taken at time t, r t represents the reward obtained at time t.

[0081] (3.2) Combine a PID controller to output a control command.

[0082] (3.3) Encoder gradient blocking, by blocking the gradient propagation from the policy network to the encoder, to prevent the encoder from learning task-irrelevant features.

[0083] (3.4) Hyperparameter tuning and training, monitor the training process, and adjust the hyperparameters as needed to accelerate the training speed and improve stability.

[0084] The above are only the preferred embodiments of the present invention, and do not constitute any other form of limitation to the present invention. Any modification or equivalent change made according to the technical essence of the present invention still falls within the scope claimed by the present invention.

Claims

1. An end-to-end method for autonomous driving based on multimodal data, characterized in that: The steps include: Step 1: Process the input data through a multimodal encoder, including independent processing of navigation maps, images, trajectory points, and vehicle measurement data. Step 2: Based on the constructed auxiliary loss branch, the importance of traffic light information in the potential representation of the image is enhanced, the traffic light classification task is completed, and the perceptual information representation ability of the input features is improved. Step 3: Process the concatenated feature vector based on the SAC algorithm and combine it with a PID controller to output the control command.

2. The end-to-end method for autonomous driving based on multimodal data according to claim 1, characterized in that: The step 1 specifically includes the following steps: (1.1) The input image data is enhanced. The image encoder consists of a convolutional neural network, which is used to represent low-dimensional sensor data as high-dimensional features. The final processing flow can be expressed as the formula: Among them, f i is the image encoder, aug represents the applied data enhancement technology, is the feature representation of the image after encoding. A collection of image data representing past time steps; (1.2) The waypoint encoder based on the WayConv1D method exploits the 2D geometric structure of the input by applying a 1D convolution with a 2×2 kernel to the 2D coordinates of the next N waypoints. The output of the 1D convolution is then flattened and processed by a multilayer perceptron. This process can be summarized as: Among them, f w is the WayConv1D waypoint encoder, is the feature representation of the path point data after encoding, W t Indicates the current step; (1.3) Based on the vehicle measurement encoder, a multi-layer perceptron MLP is used to process the vehicle measurement data and generate a potential representation. The specific process can be expressed as the formula: Among them, f v It is a vehicle measurement encoder. is the feature representation of the vehicle measurement data after encoding, represents a collection of vehicle measurement data at past time steps; (1.4) Based on the standard map encoder, the graph convolutional network GCN is used to process the standard map data and generate a potential representation. The specific process can be expressed as the formula: Among them, f u It is a vehicle measurement encoder. is the feature representation of the vehicle measurement data after encoding, A collection of standard map data representing past time steps; (1.5) Different encoders process the input data independently, obtain their potential representations and concatenate them into a feature vector 3. The end-to-end method for autonomous driving based on multimodal data according to claim 1, characterized in that: The step 2 specifically includes the following steps: (3.1) Construct auxiliary tasks: traffic light classification, add traffic light encoder, and define classification labels; (3.2) Collect data and calculate loss. Sample from the experience replay buffer, use the traffic light decoder to classify the sampled image into traffic lights, and then calculate the cross entropy loss. The calculation formula is as follows: in, is the cross entropy loss, B is the batch size, C is the number of categories, x corresponds to the output value of the traffic light decoder, and y is the true label; (3.3) Back-propagation stage, the parameters of the traffic light decoder and image encoder are optimized and updated through auxiliary loss; (3.4) Monitor the performance of the training process when handling traffic light related situations, and adjust the hyperparameters of the auxiliary tasks as needed.

4. The end-to-end method for autonomous driving based on multimodal data according to claim 1, characterized in that: The step three specifically includes the following steps: (2.1) Receive the feature vector generated in step 2, and find the optimal strategy through the actor strategy and two Q functions, by updating the value function, updating the strategy, and learning the temperature parameters. This process can be expressed as the formula: in, and are two value functions, represents the optimal strategy learned, That is, the actor strategy, α is the temperature parameter, γ is the discount factor, a t is the action taken at time t, r t represents the reward obtained at time t; (2.2) Combine with a PID controller to output control commands; (2.3) Encoder gradient blocking, by blocking the gradient from the policy network from propagating to the encoder, preventing the encoder from learning features irrelevant to the task; (2.4) Hyperparameter adjustment and training, monitor the training process and adjust hyperparameters as needed to speed up training and improve stability.

Citation Information

Patent Citations

  • Method for separating error coefficients of biaxial rotation based laser gyro assembly

    CN104880182A