Progressive end-to-end trajectory planning method and system based on BEV characteristics

By adopting a progressive trajectory planning method based on BEV features in end-to-end autonomous driving systems, the problems of unclear dependencies between modules and insufficient feature expression are solved, and more efficient system performance and security are achieved.

CN120143673APending Publication Date: 2025-06-13NANJING UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510267548.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-13

Smart Images

  • Figure CN120143673A_ABST
    Figure CN120143673A_ABST
Patent Text Reader

Abstract

The invention discloses a progressive end-to-end trajectory planning method and system based on BEV features, and the method comprises the steps: extracting rich scene features of an online mapping module, a three-dimensional occupation module, an action prediction module and the like through multi-module parallel pre-training, and generating BEV features fused with comprehensive information; then, a lightweight trajectory planner is constructed, the planner performs trajectory planning only by depending on the BEV features, a probability planning method is introduced to deal with uncertainty in planning, probability distribution of actions is output, and one action is sampled from the probability distribution to control the vehicle; the system comprises a storage medium and a processor, and is used for operating according to the instruction to execute the method. According to the method, by simplifying the model structure, interference among multiple tasks is avoided, and the operation efficiency and the trajectory planning precision are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of autonomous driving trajectory planning, and particularly to a progressive end-to-end trajectory planning method and system based on BEV features. Background Art

[0002] In recent years, autonomous driving technology has become a research hotspot in academia and industry, and end-to-end autonomous driving is an important task in the field of autonomous driving. End-to-end autonomous driving directly learns the vehicle's trajectory planning from sensor data (such as camera images, lidar, radar, etc.) through a single neural network model, avoiding the complex modular processing in traditional autonomous driving systems. Through a unified model framework, end-to-end autonomous driving effectively reduces the error accumulation in information transmission and calculation processes, improves the response speed and stability of the system, and gradually becomes the core direction of the development of autonomous driving technology.

[0003] Although significant progress has been made in end-to-end autonomous driving technology, it still faces a series of challenges in practical applications. First of all, end-to-end autonomous driving usually integrates multiple sub-task modules (such as object detection, object tracking, scene mapping, trajectory prediction, etc.), and these tasks are processed through a unified network. The model structure is complex, and it is difficult to coordinate and integrate between tasks. Moreover, the problems of information redundancy and target conflict are relatively prominent. It is difficult to balance the performance of each module during the training process, which affects the performance of the overall system. And the existing algorithms follow a deterministic paradigm and directly regress actions, making it difficult to handle the diversity of human driving behaviors and non-convex solution spaces, and easily outputting compromise or dominant trajectories (such as parking or going straight), resulting in potential safety hazards and a decline in planning performance. Therefore, a more efficient algorithm for end-to-end trajectory planning is needed. Summary of the Invention

[0004] The purpose of the present invention is to address the problems in the background art and propose a progressive end-to-end trajectory planning method and system based on BEV features.

[0005] The technical solution of the present invention, the first aspect of the present invention provides a progressive end-to-end trajectory planning method based on BEV features, including the following specific steps:

[0006] S1. Obtain the omnidirectional view pictures in the autonomous driving scenario;

[0007] S2. Obtain the two-dimensional features of each omnidirectional view picture in the autonomous driving scenario;

[0008] S3. Construct a BEV encoder module, and generate BEV features according to the two-dimensional features of each omnidirectional view picture in step S2;

[0009] S4. Build an online mapping module, an action prediction module, and a 3D occupancy prediction module; the three modules share the BEV features in step S3, perform multi-module parallel pre-training, and generate updated BEV features that integrate comprehensive scene information.

[0010] S5. Build a trajectory planning module, use the BEV features generated in step S4 as a basis, and use them for further training of the trajectory planning module to obtain the final trajectory prediction result.

[0011] Further, in step S1, the surround-view images are captured by on-vehicle cameras, which are respectively: front-view image, left-front-view image, right-front-view image, rear-view image, left-rear-view image, and right-rear-view image.

[0012] Further, the obtaining of the two-dimensional features of each surround-view image in step S2 specifically includes:

[0013] Use the deep residual network ResNet to extract features from each surround-view image. The extracted two-dimensional features are multi-channel feature tensors output by the deep residual network, including local geometric features such as edges and textures, and semantic features such as object categories and shapes. Combine all the features to form a feature map, that is, obtain the feature map F of the surround-view image.

[0014] Further, the generation of the BEV features in step S3 specifically includes:

[0015] Use the Encoder module proposed in the BEVFormer algorithm as the BEV encoder module to encode the feature map F of the surround-view image into BEV features B; the specific structure of the BEV encoder module includes a temporal self-attention module and a spatial cross-attention module.

[0016] Further, the generation of the updated BEV features that integrate comprehensive information in step S4 specifically includes:

[0017] Step S41. Build an online mapping module and generate an online vector high-precision map based on the BEV features.

[0018] Step S42. Build an action prediction module and predict the future movement trajectories of all agents based on the BEV features.

[0019] Step S43. Build a 3D occupancy prediction module and predict the 3D occupancy of the space around the vehicle based on the BEV features to generate a 3D occupancy grid map.

[0020] Step S44. Generate the final BEV features through multi-module parallel pre-training.

[0021] Among them, the online mapping module, the motion prediction module, and the 3D occupancy prediction module are pre-trained in parallel to generate an online high-precision map, predict future motion trajectories, and generate a 3D occupancy grid map respectively; each module optimizes the initial BEV features into supervised BEV features through supervised learning with the loss function corresponding to the task.

[0022] Further, it is characterized in that generating the online vector high-precision map described in step S41 specifically includes:

[0023] Step S411, defining a set of instance-level queries and a set of point-level queries where i is the index of the map element, j is the index of the point corresponding to the map element, N is the number of map elements, and N v is the number of points included in each map element;

[0024] These queries are shared by all instances and are used to encode the features of the map elements. Each map element can correspond to different features, including lane dividers, road boundaries, or pedestrian crosswalks;

[0025] Step S412, defining a set of hierarchical queries for each map element i, where the hierarchical queries are obtained by adding the instance-level queries and the point-level queries, and the calculation formula is:

[0026]

[0027] where is the instance-level query, is the point-level query, is the hierarchical query of the j-th point of the i-th map element;

[0028] Step S413, inputting all the hierarchical queries into the cascaded decoder layers for processing. Each decoder layer extracts the map features by iteratively updating the hierarchical queries; in each decoder layer, the multi-head self-attention mechanism is adopted to enable the hierarchical queries to exchange information with each other and perform information interaction between and within instances;

[0029] Step S414, enabling the hierarchical queries to interact with the BEV features through the deformable attention mechanism;

[0030] Step S415, outputting a classification branch and a point regression branch through the prediction head; the classification branch predicts the instance category scores of each map element, and the point regression branch predicts the positions of the point sets, where the point sets contain the normalized BEV coordinates of each point of the map element.

[0031] Further, generating the future motion trajectories of all agents described in step S42 specifically includes:

[0032] This action prediction task uses the action prediction module in the UniAD algorithm. The algorithm defines agent queries, promotes the interaction between query features through the self-attention mechanism, and interacts the query features with BEV features through the cross-attention mechanism. Finally, trajectory decoding is performed to output multiple possible future motion trajectories.

[0033] Further, the generation of the three-dimensional occupancy grid map described in step S43 specifically includes:

[0034] This three-dimensional occupancy prediction task uses the SurroundOcc algorithm. The algorithm converts BEV features into three-dimensional features through a 2D-3D spatial attention module and decodes the three-dimensional spatial occupancy semantics to predict the three-dimensional occupancy grid map.

[0035] Further, the obtaining of the final trajectory prediction result described in step S5 specifically includes:

[0036] Step S51, extract the planning vocabulary V from the nuScenes dataset, and use the farthest trajectory sampling strategy to select N representative trajectories, N = 4096, as the reference actions in the model planning and prediction process;

[0037] Step S52, encode each trajectory in the planning vocabulary V into a high-dimensional embedding representation E(a), where a represents the trajectory, to capture the features and information of the trajectory;

[0038] Step S53, encode the driving command into an embedding representation E cmd and add it to the high-dimensional embedding representation E(a) of the planning vocabulary obtained in step S52 to obtain a composite feature representation;

[0039] This embedding representation uses the trigonometric function-based position information encoding method proposed in the VADv2 algorithm. This method converts continuous position information into a high-dimensional embedding representation through trigonometric functions to effectively capture the spatial features of the trajectory.

[0040] Step S54, input the composite feature representation into a multi-layer perceptron MLP for encoding, and then perform a max-pooling operation on the modal dimension to obtain an aggregated planning query;

[0041] Step S55, use a cascaded Transformer decoder to interact the aggregated planning query with the BEV features described in step S4 through the cross-attention mechanism to obtain an optimized planning query;

[0042] Step S56, fuse the optimized planning query with the ego-vehicle state information, input it into a multi-layer perceptron MLP for encoding, and output the probability distribution of the trajectories: a set of N trajectories and their corresponding probability values {p(a 1 ), p(a2 ), …, p(a N )};

[0043] Step S57, select the trajectory with the highest probability in the probability distribution as the finally predicted trajectory.

[0044] The second aspect of the present invention provides a progressive end-to-end trajectory planning system based on BEV features, which uses the above method for trajectory planning, including an image acquisition module, a feature extraction module, a BEV encoder module, an online mapping module, an action prediction module, a three-dimensional occupancy prediction module, and a trajectory planning module;

[0045] The image acquisition module is used to acquire panoramic view pictures in the autonomous driving scenario;

[0046] The feature extraction module extracts two-dimensional features according to the acquired panoramic view pictures;

[0047] The BEV encoder module generates BEV features according to the two-dimensional features;

[0048] The online mapping module generates an online vector high-precision map according to the obtained BEV features;

[0049] The action prediction module predicts the future movement trajectories of all agents according to the obtained BEV features;

[0050] The three-dimensional occupancy prediction module predicts the three-dimensional occupancy of the space around the vehicle according to the obtained BEV features and generates a three-dimensional occupancy grid map;

[0051] The trajectory planning module uses the BEV features to interact through the cross-attention mechanism to obtain an optimized planning query; fuses the optimized planning query with the ego-vehicle state information, inputs it into a multi-layer perceptron for encoding, outputs the probability distribution of the trajectory, and selects the trajectory with the highest probability in the probability distribution as the finally predicted trajectory.

[0052] Compared with the prior art, the present invention has the following beneficial technical effects:

[0053] The method of the present invention can effectively solve the problems of unclear dependencies between modules and insufficient feature expression in end-to-end autonomous driving. This method uses progressive pre-training to gradually increase the complexity and difficulty of the training task, enabling the model to accumulate knowledge at each stage and finally efficiently complete the target task.

[0054] First, in the first stage, the method pre-trains intermediate modules such as perception, mapping, 3D occupancy, and motion prediction, and adopts a multi-module parallel approach to extract richer scene features, thereby generating BEV features that fuse comprehensive information. Compared with traditional serial or hybrid designs, this parallel pre-training method not only improves the feature extraction efficiency but also reduces the mutual interference between modules, significantly optimizing the collaborative performance between tasks. Second, in the second stage, the method designs a lightweight planner that relies only on the BEV features generated in the pre-training stage for trajectory planning, and introduces a probabilistic planning method into the planner to handle uncertainties in planning, outputting the probability distribution of actions, and sampling an action from it to control the vehicle. By focusing on a single task, the method of the present invention effectively avoids interference between multiple tasks, significantly simplifies the model structure and training process, and improves the overall safety of operation and deployment. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 is a flowchart of the method of the present invention.

[0056] Figure 2 is a schematic diagram of a panoramic view picture of an autonomous driving scenario.

[0057] Figure 3 is a schematic diagram of the distribution of on-vehicle cameras.

[0058] Figure 4 is a diagram of the trajectory planning result. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0059] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the protection scope of the present invention.

[0060] The application principle of the present invention will be described in detail below with reference to the drawings.

[0061] Embodiment 1

[0062] A progressive end-to-end trajectory planning method based on BEV features proposed by the present invention, see Figure 1 , specifically includes:

[0063] Step 1, obtaining panoramic view pictures in the autonomous driving scenario; the panoramic view pictures are taken using on-vehicle cameras, and are respectively: a front view picture, a left front view picture, a right front view picture, a rear view picture, a left rear view picture, and a right rear view picture;

[0064] Step 2, obtaining two-dimensional features of each panoramic view picture in the autonomous driving scenario;

[0065] Step 3: Construct a BEV encoder module, and generate BEV features based on the two-dimensional features of each panoramic view image in Step 2.

[0066] Step 4: Construct an online mapping module, an action prediction module, and a 3D occupancy prediction module. The three modules share the BEV features in Step 3 for multi-module parallel pre-training to generate updated BEV features that integrate comprehensive scene information.

[0067] Step 5: Construct a trajectory planning module, and use the BEV features generated in Step 4 as a basis for further training of the trajectory planning module to obtain the final trajectory prediction result.

[0068] Preferably, Step 2 includes:

[0069] Use the Deep Residual Network ResNet (Reference: He K, Zhang X, Ren S, et al. Deep residual learning for image recognition[C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2016: 770-778.) to extract features from each panoramic view image, and form all the features into a feature map, that is, obtain the feature map F of the panoramic view.

[0070] Preferably, Step 3 includes:

[0071] Use the Encoder module proposed in the BEVFormer algorithm (Reference: He K, Zhang X, Ren S, et al. Deep residual learning for image recognition[C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2016: 770-778.) as the BEV encoder module, and encode the feature map F of the panoramic view into BEV features B. The specific structure of the BEV encoder module includes a temporal self-attention module and a spatial cross-attention module.

[0072] Preferably, Step 4 includes:

[0073] Step 4-1: Construct an online mapping module, and generate an online vector high-precision map based on the BEV features.

[0074] Step 4-2: Construct an action prediction module, and predict the future movement trajectories of all agents based on the BEV features.

[0075] Step 4-3: Construct a 3D occupancy prediction module to predict the 3D occupancy of the space around the vehicle based on the BEV features and generate a 3D occupancy grid map;

[0076] Step 4-4: Generate the final BEV features through multi-module parallel pre-training.

[0077] Preferably, Step 4-1 includes:

[0078] Step 4-1-1: Define a set of instance-level queries and a set of point-level queries for each map element, where i is the index of the map element, j is the index of the point corresponding to the map element, N is the number of map elements, and N v is the number of points included in each map element.

[0079] These queries are shared by all instances and are used to encode the features of the map elements. Each map element can correspond to different features, such as lane dividers, road boundaries, or pedestrian crosswalks, etc.;

[0080] Step 4-1-2: Define a set of hierarchical queries for each map element i, where the hierarchical queries are obtained by adding the instance-level queries and the point-level queries, and the calculation formula is:

[0081]

[0082] where, is the instance-level query, is the point-level query, is the hierarchical query of the j-th point of the i-th map element;

[0083] Step 4-1-3: Input all the hierarchical queries into the cascaded decoder layers for processing. Each decoder layer extracts the map features by iteratively updating the hierarchical queries. In each decoder layer, the multi-head self-attention mechanism (MHSA) is adopted to enable the hierarchical queries to exchange information with each other and perform information interaction between and within instances;

[0084] Step 4-1-4: Enable the hierarchical queries to interact with the BEV features through the deformable attention mechanism. This mechanism can handle map elements with irregular shapes and effectively capture long-range context dependencies to ensure accurate modeling of complex road environments;

[0085] Step 4-1-5: Output a classification branch and a point regression branch through the prediction head. The classification branch predicts the instance category scores of each map element, and the point regression branch predicts the positions of the point sets, where the point sets contain the normalized BEV coordinates of the respective points of the map elements;

[0086] Preferably, step 4-2 includes:

[0087] The action prediction task uses the action prediction module in the UniAD algorithm (Reference: Hu Y, Yang J, Chen L, et al. Planning-oriented autonomous driving[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2023: 17853-17862.). The algorithm defines agent queries, promotes the interaction between query features through the self-attention mechanism, and interacts the query features with BEV features through the cross-attention mechanism. Finally, trajectory decoding is performed to output multiple possible future motion trajectories.

[0088] Preferably, step 4-3 includes:

[0089] The 3D occupancy prediction task uses the SurroundOcc algorithm (Reference: Wei Y, Zhao L, Zheng W, et al. Surroundocc: Multi-camera 3d occupancy prediction for autonomous driving[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2023: 21729-21740.). The algorithm transforms BEV features into 3D features through a 2D-3D spatial attention module and decodes the 3D space occupancy semantics to predict a 3D occupancy grid map.

[0090] Preferably, step 4-4 includes:

[0091] The online mapping module, the action prediction module, and the 3D occupancy prediction module generate an online high-precision map, predict future motion trajectories, and generate a 3D occupancy grid map respectively through parallel pre-training. Each module performs supervised learning with the loss function corresponding to the task, so as to optimize the initial BEV features into supervised BEV features.

[0092] Preferably, step 5 includes:

[0093] Step 5-1, extract the planning vocabulary V from the nuScenes dataset, and use the farthest trajectory sampling strategy to select N representative trajectories (N = 4096) as the reference actions in the model planning and prediction process;

[0094] Step 5-2: Encode each trajectory in the planning vocabulary V into a high-dimensional embedding representation E(a), where a represents the trajectory, to capture the features and information of the trajectory;

[0095] This embedding representation uses the trigonometric-based position information encoding method proposed in the VADv2 algorithm (Reference: Chen S, Jiang B, Gao H, et al. Vadv2: End-to-end vectorized autonomous driving via probabilistic planning[J]. arXiv preprint arXiv:2402.13243, 2024.). This method converts continuous position information into a high-dimensional embedding representation through trigonometric functions to effectively capture the spatial features of the trajectory.

[0096] Step 5-3: Encode the driving command into an embedding representation E cmd and add it to the high-dimensional embedding representation E(a) of the planning vocabulary obtained in Step 5-2 to obtain a composite feature representation;

[0097] Step 5-4: Input the composite feature representation into a multi-layer perceptron (MLP) for encoding, and then perform a max-pooling operation on the modality dimension to obtain an aggregated planning query;

[0098] Step 5-5: Use a cascaded Transformer decoder (Reference: Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need[J]. Advances in neural information processing systems, 2017, 30.) to interact the aggregated planning query with the BEV features described in Step 4 through a cross-attention mechanism to obtain an optimized planning query;

[0099] Step 5-6: Fuse the optimized planning query with the ego-vehicle state information (speed, acceleration, heading angle), input it into a multi-layer perceptron (MLP) for encoding, and output the probability distribution of the trajectory: a set of N trajectories and their corresponding probability values {p(a 1 ), p(a 2 ), …, p(a N )}.

[0100] Step 5-7: Select the trajectory with the highest probability in the probability distribution as the finally predicted trajectory. Figure 4This is an example of the trajectory planning result for this embodiment. This figure is an aerial view. Point A in the figure represents the current position of the autonomous vehicle. There are a total of 6 points, P1 - P6, representing the trajectory points predicted by the autonomous driving system within the next 3 seconds. One point is predicted every 0.5 seconds, and all the points are connected to obtain a trajectory route. The rectangular frames in the figure represent other vehicles or pedestrians in the current scenario, and the blue gradient curve represents the predicted future trajectories of these moving objects. It can be seen that this embodiment can more accurately predict the future trajectory of the ego vehicle in complex driving scenarios including sharp curves and multi - obstacle interactions, significantly improving the system's dynamic path tracking ability on unstructured roads, and enhancing the accuracy of the trajectory planning task and the safety of autonomous driving.

[0101] Embodiment 2

[0102] This embodiment provides a progressive end - to - end trajectory planning system based on BEV features, which uses the method in Embodiment 1 for trajectory planning. Specifically, it includes an image acquisition module, a feature extraction module, a BEV encoder module, an online mapping module, an action prediction module, a 3D occupancy prediction module, and a trajectory planning module;

[0103] The image acquisition module is used to collect panoramic view pictures in the autonomous driving scenario;

[0104] The feature extraction module extracts two - dimensional features based on the collected panoramic view pictures;

[0105] The BEV encoder module generates BEV features based on the two - dimensional features;

[0106] The online mapping module generates an online vector high - precision map based on the obtained BEV features;

[0107] The action prediction module predicts the future movement trajectories of all agents based on the obtained BEV features;

[0108] The 3D occupancy prediction module predicts the 3D occupancy of the space around the vehicle based on the obtained BEV features and generates a 3D occupancy grid map;

[0109] The trajectory planning module uses the BEV features to interact through a cross - attention mechanism to obtain an optimized planning query; fuses the optimized planning query with the ego - vehicle state information, inputs it into a multi - layer perceptron for encoding, and outputs the probability distribution of the trajectory. Selects the trajectory with the highest probability in the probability distribution as the finally predicted trajectory.

[0110] Embodiment 3

[0111] The present invention also provides a progressive end - to - end trajectory planning system based on BEV features, including a storage medium and a processor;

[0112] The storage medium is used to store instructions;

[0113] The processor is configured to operate according to the instructions to execute the steps of the method according to any one of the first aspect.

[0114] The present invention also provides a computer-readable storage medium, on which a program or instructions are stored, and when the program or instructions are executed by a processor, the three-dimensional space occupancy recognition method as described above is implemented.

[0115] In a specific implementation, the present application provides a computer storage medium and a corresponding data processing unit. Among them, the computer storage medium can store a computer program, and when the computer program is executed by the data processing unit, it can run the content of the invention of a progressive end-to-end trajectory planning method and system based on basic BEV features and some or all of the steps in each embodiment. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), or the like.

[0116] Those skilled in the art can clearly understand that the technical solutions in the embodiments of the present invention can be implemented by means of a computer program and its corresponding general hardware platform. Based on such an understanding, the essence of the technical solutions in the embodiments of the present invention, or the part that contributes to the prior art, can be embodied in the form of a computer program, that is, a software product. The computer program software product can be stored in a storage medium, including several instructions for causing a device (which can be a personal computer, a server, a single-chip microcomputer, an MCU, or a network device, etc.) including a data processing unit to execute the methods described in each embodiment or some parts of the embodiments of the present invention.

[0117] The present invention provides an idea and method for a progressive end-to-end trajectory planning method and system based on BEV features. There are many methods and ways to specifically implement this technical solution. The above description is only the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by the prior art.

Claims

1. A progressive end-to-end trajectory planning method based on BEV features, characterized in that: The specific steps include: S1. Obtain a surround view image in the autonomous driving scene; S2, obtaining two-dimensional features of each surround view image in the autonomous driving scene; S3, constructing a BEV encoder module, generating BEV features according to the two-dimensional features of each surround view image in step S2; S4, constructing an online mapping module, an action prediction module and a 3D occupancy prediction module; The three modules share the BEV features in step S3, perform multi-module parallel pre-training, and generate updated BEV features that integrate comprehensive scene information; S5. Construct a trajectory planning module, using the BEV features generated in step S4 as a basis for further training the trajectory planning module to obtain the final trajectory prediction result.

2. The progressive end-to-end trajectory planning method based on BEV features according to claim 1 is characterized in that: In step S1, the surround view pictures are taken using a vehicle-mounted camera, which are: a front view picture, a left front view picture, a right front view picture, a rear view picture, a left rear view picture, and a right rear view picture.

3. The progressive end-to-end trajectory planning method based on BEV features according to claim 1 is characterized in that: The step S2 of obtaining the two-dimensional features of each surround view image in the autonomous driving scene specifically includes: A deep residual network ResNet is used to extract features from each surround view image. The extracted two-dimensional features are multi-channel feature tensors output by the deep residual network, including local geometric features such as edges and textures, as well as semantic features such as object categories and shapes. All features are combined into a feature map to obtain the feature map F of the surround view.

4. The progressive end-to-end trajectory planning method based on BEV features according to claim 1 is characterized in that: The generating of BEV features in step S3 specifically includes: The Encoder module proposed in the BEVFormer algorithm is used as the BEV encoder module to encode the feature map F of the surround view into the BEV feature B; the specific structure of the BEV encoder module includes a temporal self-attention module and a spatial cross-attention module.

5. The progressive end-to-end trajectory planning method based on BEV features according to claim 1, characterized in that: The step S4 of generating the updated BEV features integrating comprehensive information specifically includes: Step S41, constructing an online mapping module to generate an online vector high-precision map based on BEV features; Step S42, constructing an action prediction module to predict the future motion trajectories of all agents based on the BEV features; Step S43, constructing a three-dimensional occupancy prediction module, predicting the three-dimensional occupancy of the space around the vehicle according to the BEV characteristics, and generating a three-dimensional occupancy grid map; Step S44, generating the final BEV features through multi-module parallel pre-training; Among them, the online mapping module, motion prediction module and three-dimensional occupancy prediction module are pre-trained in parallel to generate online high-precision maps, predict future motion trajectories and generate three-dimensional occupancy grid maps respectively; each module performs supervised learning with the loss function of the corresponding task to optimize the initial BEV features into supervised BEV features.

6. The progressive end-to-end trajectory planning method based on BEV features according to claim 5, characterized in that: It is characterized in that The generation of an online vector high-precision map in step S41 specifically includes: Step S411, define a set of instance-level queries for each map element and a set of point-level queries Where i is the index of the map element, j is the index of the corresponding point of the map element, N is the number of map elements, and N v The number of points contained in each map element; These queries are shared by all instances and are used to encode features of map elements, each of which can correspond to different features, including lane dividers, road boundaries, or pedestrian crosswalks; Step S412: define a set of hierarchical queries for each map element i Among them, the hierarchical query is obtained by adding the instance-level query and the point-level query, and the calculation formula is: in, For instance-level queries, For point-level queries, Hierarchical query for the jth point of the i-th map element; Step S413, all hierarchical queries are input into the cascaded decoder layers for processing, and each decoder layer extracts map features by iteratively updating the hierarchical queries; in each decoder layer, a multi-head self-attention mechanism is used to enable the hierarchical queries to exchange information with each other and perform information interaction between instances and within instances; Step S414, enabling the hierarchical query to interact with the BEV features through a deformable attention mechanism; Step S415, outputting a classification branch and a point regression branch through the prediction head; the classification branch predicts the category score of each map element instance, and the point regression branch predicts the position of a point set, and the point set contains the normalized BEV coordinates of each point of the map element.

7. The progressive end-to-end trajectory planning method based on BEV features according to claim 5, characterized in that: The generation of future motion trajectories of all agents in step S42 specifically includes: This action prediction task uses the action prediction module in the UniAD algorithm. The algorithm defines the agent query, promotes the interaction between query features through the self-attention mechanism, and interacts the query features with the BEV features through the cross-attention mechanism. Finally, trajectory decoding is performed to output multiple possible future motion trajectories.

8. The progressive end-to-end trajectory planning method based on BEV features according to claim 5, characterized in that: The step S43 of generating a three-dimensional occupancy grid map specifically includes: The three-dimensional occupancy prediction task uses the SurroundOcc algorithm. The algorithm converts BEV features into three-dimensional features through a 2D-3D spatial attention module and decodes the three-dimensional space occupancy semantics to predict the three-dimensional occupancy grid map.

9. The progressive end-to-end trajectory planning method based on basic BEV characteristics according to claim 1, characterized in that: The final trajectory prediction result obtained in step S5 specifically includes: Step S51, extracting the planning vocabulary V from the nuScenes dataset, using the farthest trajectory sampling strategy, selecting N representative trajectories N = 4096 as reference actions in the model planning and prediction process; In step S52, each trajectory in the planning vocabulary V is encoded into a high-dimensional embedding representation E(a), where a represents the trajectory, capturing the characteristics and information of the trajectory. Step S53, encoding the driving command into an embedded representation E cmd , and add it to the high-dimensional embedding representation E(a) of the planned vocabulary obtained in step S52 to obtain a composite feature representation; Step S54, inputting the composite feature representation into the multi-layer perceptron MLP for encoding, and then performing a maximum pooling operation on the modal dimension to obtain an aggregated planning query; Step S55, using a cascaded Transformer decoder, the aggregated planning query interacts with the BEV features described in step S4 through a cross-attention mechanism to obtain an optimized planning query; Step S56, the optimized planning query is integrated with the vehicle state information, input into the multi-layer perceptron MLP for encoding, and the probability distribution of the trajectory is output: a set of N trajectories and their corresponding probability values ​​{pa1, pa2, ..., pa N }; Step S57, selecting a trajectory with the highest probability in the probability distribution as the final predicted trajectory.

10. A progressive end-to-end trajectory planning system based on BEV features, using the method described in any one of claims 1-9 for trajectory planning, characterized in that: It includes image acquisition module, feature extraction module, BEV encoder module, online mapping module, action prediction module, 3D occupancy prediction module and trajectory planning module; The image acquisition module is used to collect surround view images in the autonomous driving scene; The feature extraction module extracts two-dimensional features based on the collected surround view images; The BEV encoder module generates BEV features based on the two-dimensional features; The online mapping module generates an online vector high-precision map based on the obtained BEV features; The action prediction module predicts the future motion trajectories of all agents based on the obtained BEV features; The three-dimensional occupancy prediction module predicts the three-dimensional occupancy of the space around the vehicle based on the obtained BEV features and generates a three-dimensional occupancy grid map; The trajectory planning module uses BEV features to interact through a cross-attention mechanism to obtain an optimized planning query; the optimized planning query is fused with the vehicle state information, input into a multi-layer perceptron for encoding, and the probability distribution of the output trajectory is selected as the final predicted trajectory.

Citation Information

Cited By

  • Lightweight online high-precision map construction method, device and system and storage medium

    CN120685107A

  • Automatic driving decision-making method based on three-dimensional scene reconstruction and multi-modal large model

    CN120894490A

  • Autonomous driving decision-making method based on three-dimensional scene reconstruction and multi-modal large model

    CN120894490B