An autonomous driving planning method and device combining prior information of a city-level neural radiance field
Through the autonomous driving planning method combining the prior information of urban-level neural radiation field, the accuracy and robustness of the autonomous driving system in complex environments is solved, safe and accurate path planning is achieved, and the system's adaptability and response speed in variable driving scenarios is improved.
Patent Information
- Application Number
- CN202510512113.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-23
AI Technical Summary
The autonomous driving system faces the challenges of accuracy, robustness and real-time in complex and changing road environments, especially in severe weather and low visibility conditions, the stability and reliability of the perception system are insufficient, the construction and update of high-precision maps are difficult to ensure safe and reasonable decision-making and path planning, and other key technologies such as multi-objective tracking and prediction, decision-making and planning.
Combining the autonomous driving planning method with a priori information of the city-level neural radiation field, by obtaining real-time image features and converting them into bird's-eye viewing information, a priori information is extracted using the pre-trained urban-level neural radiation field model, and real-time and priori information are imported into the autonomous driving planning model to generate a safe and accurate planning path.
It realizes accurate perception of static and dynamic environments in complex environments, ensures control safety and planning accuracy, and improves the adaptability and response speed of the autonomous driving system in variable driving scenarios.
Smart Images

Figure CN120024358B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving, and in particular, to an autonomous driving planning method and device combining prior information of a city-level neural radiance field. Background Art
[0002] Autonomous driving technology, as the forefront of the integration of artificial intelligence and intelligent transportation systems, has witnessed a research and application boom globally in recent years. Thanks to the rapid progress of computing power, sensor technology, and deep learning algorithms, autonomous driving technology has gradually advanced from theoretical exploration in the laboratory to the stage of practical application, demonstrating great commercial potential and value in multiple fields such as passenger vehicles, commercial vehicles, logistics transportation, shared mobility, and unmanned delivery.
[0003] Automobile manufacturers, technology giants, and startups worldwide have actively engaged in the research and development of autonomous driving systems, forming a highly competitive landscape. Tesla leads the popularization of semi-autonomous driving technology with its Autopilot system, while Waymo has made breakthroughs in fully driverless taxi services. Meanwhile, domestic companies such as Baidu, NIO, XPeng, and Li Auto have also continuously innovated in the field of autonomous driving, promoting the overall development of the industry. Baidu's Apollo platform has further accelerated the application of autonomous driving technology in passenger and commercial vehicles through large-scale tests in multiple cities across the country.
[0004] However, behind the booming development of autonomous driving technology, there are still many technical pain points and challenges. Firstly, the perception problem in complex environments is particularly prominent. Autonomous driving systems need to accurately identify and process diverse road environments, including traffic signs, pedestrians, obstacles, etc. Especially in adverse weather and low visibility conditions, the stability and reliability of the perception system are severely tested. Secondly, the construction and update of high-precision maps are also a major challenge. Autonomous driving relies on accurate map data for navigation and decision-making, but the real-time update, coverage, and rapid changes in dynamic scenarios of maps pose higher requirements for the accuracy and timeliness of maps. In addition, key technologies such as multi-object tracking and prediction, decision-making, and planning also face challenges. Autonomous driving systems need to handle various uncertainties in dynamic traffic environments to ensure safe and reasonable decision-making and path planning.
[0005] In summary, there is an urgent need for a method that enables autonomous driving systems to balance accuracy, robustness, and real-time performance in complex and changing road environments. Summary of the Invention
[0006] The objective of the embodiments of the present invention is to provide an autonomous driving planning method and device combining prior information of a city-level neural radiance field to ensure the safety and accuracy of autonomous driving planning.
[0007] In a first aspect, the present invention provides an autonomous driving planning method that combines prior information of a city-level neural radiance field. The method includes:
[0008] Obtain real-time images collected by a camera device on a target vehicle, perform feature extraction on the real-time images, and convert the extracted features into real-time feature information in a bird's-eye view.
[0009] Extract prior information of the city-level neural radiance field based on a pre-trained city-level neural radiance field model.
[0010] Convert the prior information of the city-level neural radiance field into prior feature information in a bird's-eye view.
[0011] Import the real-time feature information and the prior feature information into a pre-trained autonomous driving planning model, and output autonomous driving planning information for the target vehicle.
[0012] In an optional implementation, the step of performing feature extraction on the real-time images and converting the extracted features into real-time feature information in a bird's-eye view includes:
[0013] Perform a convolution operation on the real-time images to extract various different types of features of the real-time images.
[0014] Perform a non-linear activation function processing on the various features to introduce non-linear features, and encode the processed features into feature vectors.
[0015] Use a geometric transformation matrix to convert the feature vectors in the image coordinate system into real-time feature vectors in the bird's-eye view coordinate system.
[0016] In an optional implementation, the method further includes a step of pre-training the city-level neural radiance field model, and this step includes:
[0017] Obtain a plurality of historical images collected by camera devices on multiple vehicles in a target city during a historical period.
[0018] Perform point sampling based on the light directions of the camera devices corresponding to the respective historical images to obtain a plurality of sampling points.
[0019] For each of the sampling points, obtain the 3D position information, view direction, and video identifier of the sampling point.
[0020] Import the 3D position information, view direction, and video identifier corresponding to each of the sampling points into a constructed initial model, perform processing based on the initial model, and output corresponding density information, color information, and semantic feature information.
[0021] Guided by the constructed loss function, based on the density information, color information, and semantic feature information output by the initial model, as well as the density information, color information, and semantic feature information corresponding to the region information to which the sampling points belong, the initial model is iteratively trained multiple times until a trained city-level neural radiance field model is obtained when the preset requirements are met.
[0022] In an alternative embodiment, the step of performing point sampling on the light directions of the camera devices corresponding to the respective historical images to obtain a plurality of sampling points includes:
[0023] Mapping the pose information of the camera devices corresponding to the respective historical images to a data point;
[0024] Performing clustering processing on the plurality of data points to divide the plurality of data points into a plurality of sub-regions, and then performing clustering processing on the data points within each of the sub-regions to divide them into a plurality of sub-fields;
[0025] For each of the historical images, performing point sampling on the light direction of the camera device corresponding to the historical image to obtain a plurality of sampling points, and determining the region information to which each of the sampling points belongs, where the region information is sub-field information.
[0026] In an alternative embodiment, the step of performing clustering processing on the plurality of data points to divide the plurality of data points into a plurality of sub-regions includes:
[0027] Setting the number of clusters for clustering based on the scale information of the target city, and randomly generating the centroids of each cluster;
[0028] Based on the distances between the respective data points and the centroids of each cluster, dividing each of the data points into the corresponding cluster, and updating the centroids of the clusters according to the position information of the data points included in each cluster. After multiple iterations, until the centroids of each cluster reach stability, a plurality of sub-regions are obtained based on the obtained plurality of clusters.
[0029] In an alternative embodiment, the step of pre-training to obtain the city-level neural radiance field model further includes:
[0030] Identifying the pixel points of potential moving objects in each of the historical images through a semantic segmentation mask;
[0031] Setting the pixel values of the identified pixel points to a specific mask value.
[0032] In an alternative embodiment, the step of extracting city-level neural radiance field prior information based on the pre-trained city-level neural radiance field model includes:
[0033] Obtaining a plurality of prior images collected by camera devices on multiple vehicles within the target city;
[0034] Import the multiple prior images into a pre-trained city-level neural radiance field model. For each of the prior images, collect a plurality of sampling points for the projection light direction of the camera device corresponding to the prior image;
[0035] Filter out surface points from the multiple sampling points based on the cumulative transmittance and opacity metrics;
[0036] Construct city-level neural radiance field prior information based on the multiple surface points corresponding to the multiple prior images. The city-level neural radiance field prior information includes road structures, traffic signs, and obstacle positions.
[0037] In an alternative embodiment, the autonomous driving planning model includes multiple modules, and the method further includes the step of pre-training the autonomous driving planning model, which includes:
[0038] Collect training samples, where the training samples are the fusion results of images and prior information;
[0039] Import the training samples into the autonomous driving planning model, and after injecting a noise set into each module of the autonomous driving planning model, perform training under the guidance of the constructed loss function;
[0040] In each round of iteration, calculate the weight values of the respective modules in the autonomous driving planning model for the next round of iteration, and calculate the total loss function value according to the weight values of the respective modules in the next round of iteration until the preset requirements are met, obtaining a trained autonomous driving planning model.
[0041] In an alternative embodiment, the step of importing the real-time feature information and the prior feature information into a pre-trained autonomous driving planning model and outputting autonomous driving planning information for the target vehicle includes:
[0042] Import the real-time feature information and the prior feature information into a pre-trained autonomous driving planning model, and perform fusion processing on the real-time feature information and the prior feature information to obtain fusion features;
[0043] Obtain perception information, tracking information, map construction information, and motion prediction information based on the fusion features;
[0044] Combine the perception information, tracking information, map construction information, and motion prediction information to generate autonomous driving planning path information.
[0045] In a second aspect, the present invention provides an autonomous driving planning device combined with city-level neural radiance field prior information, and the device includes:
[0046] The first extraction module is used to obtain the real-time image collected by the camera device on the target vehicle, extract features from the real-time image, and convert the extracted features into real-time feature information in a bird's-eye view.
[0047] The second extraction module is used to extract the prior information of the city-level neural radiance field based on the pre-trained city-level neural radiance field model.
[0048] The conversion module is used to convert the prior information of the city-level neural radiance field into prior feature information in a bird's-eye view.
[0049] The processing module is used to import the real-time feature information and the prior feature information into the pre-trained autonomous driving planning model, and output the autonomous driving planning information for the target vehicle.
[0050] The present invention provides an autonomous driving planning method and device combining prior information of a city-level neural radiance field. By extracting features from the real-time image collected by the camera device on the obtained target vehicle, converting the extracted features into real-time feature information in a bird's-eye view, and extracting the prior information of the city-level neural radiance field based on the pre-trained city-level neural radiance field model. The prior information of the city-level radiance field is converted into prior feature information in a bird's-eye view. Finally, the real-time feature information and the prior feature information are imported into the pre-trained autonomous driving planning model, and the autonomous driving planning information for the target vehicle is output. In this solution, by combining the real-time image and the prior information of the city-level neural radiance field for autonomous driving planning, the static environment and dynamic environment around the vehicle can be accurately perceived, the safety of control can be ensured, and moreover, by converting the real-time features and prior features into features in a bird's-eye view, planning can be carried out from a global perspective, ensuring the accuracy of planning. Description of the Drawings
[0051] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments of the present invention. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0052] Figure 1 It is a flowchart of the autonomous driving planning method provided by the embodiment of the present invention;
[0053] Figure 2 For Figure 1 It is a flowchart of the sub-steps included in S11 in
[0054] Figure 3In the autonomous driving planning method provided by an embodiment of the present invention, it is a flowchart of a method for training a city-level neural radiance field model;
[0055] Figure 4 For Figure 3 It is a flowchart of the sub-steps included in S22 in
[0056] Figure 5 For Figure 1 It is a flowchart of the sub-steps included in S12 in
[0057] Figure 6 In the autonomous driving planning method provided by an embodiment of the present invention, it is a flowchart of a method for training an autonomous driving planning model;
[0058] Figure 7 For Figure 1 It is a flowchart of the sub-steps included in S14 in
[0059] Figure 8 It is a functional module block diagram of an autonomous driving planning device provided by an embodiment of the present invention;
[0060] Figure 9 It is a structural block diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0061] Next, the technical solutions in the embodiments of the present invention will be described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0062] Please refer to Figure 1 , which is a flowchart of an autonomous driving planning method that combines city-level neural radiance field prior information provided by an embodiment of the present invention. This autonomous driving planning method can be executed by an autonomous driving planning device, which can be implemented by software and / or hardware and can be configured in an electronic device. The electronic device can be a computer device, a server, etc., such as a server in a backend control platform, etc. The detailed steps of this autonomous driving planning method are introduced as follows.
[0063] S11, Obtain the real-time image collected by the camera device on the target vehicle, extract features from the real-time image, and convert the extracted features into real-time feature information in a bird's-eye view.
[0064] S12, Extract the city-level neural radiance field prior information based on the pre-trained city-level neural radiance field model.
[0065] S13, Convert the city-level neural radiance field prior information into prior feature information in a bird's-eye view.
[0066] S14. Import the real-time feature information and the prior feature information into a pre-trained autonomous driving planning model, and output the autonomous driving planning information for the target vehicle.
[0067] In this embodiment, the so-called target vehicle refers to any vehicle that needs to perform autonomous driving control. Camera devices are provided at positions such as the body and the bottom of the target vehicle. During the driving process of the target vehicle, the camera devices can continuously collect images within the shooting field of view of the devices, which are herein named real-time images.
[0068] In order to perform autonomous driving planning from a global perspective, convert the features of the real-time image into real-time feature information from a bird's-eye view.
[0069] In addition, in this embodiment, a city-level neural radiance field model is pre-trained. Based on this model, historical images can be processed to extract city-level neural radiance field prior information. Historical images refer to images collected by camera devices on multiple vehicles within the city during a historical period. The prior information includes, for example, road structures, traffic signs, obstacle positions, etc. Such prior information will affect the autonomous driving planning of the vehicle.
[0070] In order to enable the effective fusion of the prior information and the real-time image information, therefore, convert the prior information into prior feature information from a bird's-eye view.
[0071] In addition, in this embodiment, an autonomous driving planning model is also pre-trained. The autonomous driving planning model can, based on the input real-time feature information and prior feature information, output the autonomous driving planning information for the target vehicle under the real-time feature information, such as including autonomous driving path information, etc.
[0072] In this embodiment, combining the real-time image and the city-level neural radiance field prior information for autonomous driving planning can accurately perceive the static environment and dynamic environment around the vehicle, ensure the safety of control, and moreover, converting the real-time features and prior features into features from a bird's-eye view can perform planning from a global perspective and ensure the accuracy of the planning.
[0073] The following first introduces the processing process of the collected real-time image:
[0074] Please refer to Figure 2 , the above step S11 may include the following sub-steps:
[0075] S111. Perform a convolution operation on the real-time image to extract various different types of features of the real-time image.
[0076] S112. Perform a non-linear activation function process on the various features to introduce non-linear features and encode the processed features into feature vectors.
[0077] S113. Convert the feature vector in the image coordinate system into a real-time feature vector in the bird's-eye view coordinate system using a geometric transformation matrix.
[0078] In this embodiment, a residual network ResNet can be used to process the real-time image. During the feature extraction process, ResNet performs a convolution operation on the input real-time image through its convolutional layers. Each convolutional kernel is responsible for extracting a specific type of feature from the real-time image, such as edges, textures, shapes, etc. These convolutional kernels slide over the image and, through weighted summation and bias term processing, extract different features of the image. Subsequently, these features are processed through a non-linear activation function to introduce non-linearity. Finally, these features are encoded into feature vectors, providing a basis for subsequent image analysis and understanding.
[0079] During the bird's-eye view feature transformation process, for the extracted image feature vector Perform further processing to generate a real-time feature vector in the bird's-eye view . To this end, first define a geometric transformation matrix T, which is responsible for converting the image coordinate system into the bird's-eye view coordinate system. The geometric transformation matrix T contains transformation parameters such as rotation, scaling, and translation, and can accurately map each point in the image to the corresponding position in the bird's-eye view.
[0080] For each pixel point (u, v) in the real-time image and its corresponding feature vector (u, v), calculate its position in the bird's-eye view coordinate system through the transformation matrix T ( ). This transformation can be expressed as:
[0081]
[0082] where T is a 3×3 homogeneous transformation matrix, which includes a rotation matrix R, a scaling factor S, and a translation vector t.
[0083] After completing the coordinate transformation, map the real-time feature vector (u, v) to the corresponding position in the bird's-eye view coordinate system ( ) to form a bird's-eye view feature map . This process can be expressed as:
[0084] ( ) (u, v)
[0085] After this conversion process, the vehicle can obtain a more comprehensive and unified environmental representation (i.e., ), which not only covers the spatial layout and object distribution of the surrounding environment but also effectively eliminates the interference caused by perspective changes. Such environmental representation enables the autonomous driving system to better grasp the surrounding environment and thus make more informed decisions in aspects such as path planning, obstacle recognition, and avoidance.
[0086] In this embodiment, the real-time feature vector of the real-time image is combined with the prior information of the city-level neural radiance field extracted based on a pre-trained city-level neural radiance field model and is used together for autonomous driving planning.
[0087] The implementation method of the pre-trained city-level neural radiance field model is introduced as follows:
[0088] Please refer to Figure 3 , in this embodiment, the city neural radiance field model is pre-trained through the following method:
[0089] S21, Obtain a plurality of historical images collected by camera devices on multiple vehicles in the target city during a historical period.
[0090] S22, Based on the light directions of the camera devices corresponding to the respective historical images, perform point sampling to obtain a plurality of sampling points.
[0091] S23, For each of the sampling points, obtain the 3D position information, view direction, and video identifier of the sampling point.
[0092] S24, Import the 3D position information, view direction, and video identifier corresponding to each of the sampling points into the constructed initial model, and perform processing based on the initial model to output corresponding density information, color information, and semantic feature information.
[0093] S25, Guided by the constructed loss function, based on the density information, color information, and semantic feature information output by the initial model, and the density information, color information, and semantic feature information corresponding to the area information to which the sampling points belong, perform multiple iterative trainings on the initial model until the trained city-level neural radiance field model is obtained when the preset requirements are met.
[0094] In this embodiment, the target city refers to the urban area where the target vehicle is located, and the multiple vehicles refer to the vehicles communicating with the backend control platform. The server in the backend control platform can obtain the historical images collected by the camera devices of the multiple vehicles during a historical period.
[0095] During implementation, the collected historical images should cover as many different conditions as possible, such as different lighting conditions (day or night), weather (sunny, rainy, snowy, etc.), and other scenarios. At the same time, the collected historical images should cover most of the road scenes in the target city, meeting a certain degree of spatial coverage and being able to reflect the dynamic changes of the target city in both the time domain and the spatial domain. In addition, to enhance the scene representation ability of the image data, a pre-trained convolutional neural network can be used to extract features from each historical image to generate pixel-level image descriptors. Through the convolutional and pooling operations of the deep convolutional neural network layer by layer, these descriptors can efficiently extract edges, textures, color distributions, and more advanced semantic information in the image, such as specific objects like lane lines, traffic signs, pedestrians, vehicles, and the relationships between them. In addition, they can also capture dynamic features such as environmental color differences brought about by seasonal changes, different road surface conditions caused by weather changes, and the density and behavior patterns of the flow of people and vehicles at different time periods. These detailed and expressive features lay a solid foundation for subsequent advanced autonomous driving tasks such as image classification, object detection, scene understanding, and path planning.
[0096] To improve the data processing efficiency and the accuracy of model construction, during implementation, based on the pose of the camera device corresponding to each historical image, that is, the pose of the camera device itself when collecting the historical image, sub-regions and sub-fields can be divided. Specifically, please refer to Figure 4 The above step S22 can be implemented in the following way:
[0097] S221, map the pose information of the camera device corresponding to each of the historical images to a data point.
[0098] S222, perform clustering processing on multiple data points to divide the multiple data points into multiple sub-regions, and then perform clustering processing on the data points within each sub-region to divide them into multiple sub-fields.
[0099] S223, for each of the historical images, perform point sampling on the light direction of the camera device corresponding to the historical image to obtain multiple sampling points, and determine the region information to which each of the sampling points belongs, where the region information is sub-field information.
[0100] First, map the pose information of each camera device when collecting each historical image to a data point, and each data point contains the spatial position and viewing direction information of the camera device.
[0101] When dividing sub - regions, the number of clusters for clustering is set based on the scale information of the target city, and the centroids of each cluster are randomly generated. Based on the distances between each data point and the centroids of each cluster, each data point is assigned to the corresponding cluster, and the centroids of the clusters are updated according to the position information of the data points included in each cluster. After multiple iterations, until the centroids of each cluster reach stability, multiple sub - regions are obtained based on the multiple clusters obtained.
[0102] In this embodiment, the K - means clustering algorithm is applied iteratively. In each iteration, the algorithm updates the centroid of each cluster (i.e., sub - region) and re - divides the clusters according to the distance between the data points and the centroids. This process is carried out multiple times until the division of the clusters no longer changes or reaches the preset number of iterations. Finally, according to the clustering results, the boundaries of each sub - region are determined. The pose information of the camera devices within the same sub - region is close to each other, indicating that their positions in the urban space are relatively concentrated. The pose information of different sub - regions is relatively dispersed, indicating that they represent different regions of the city respectively.
[0103] The centroid calculation formula for each sub - region is as follows:
[0104]
[0105] Among them, is the centroid of the j - th sub - region, is the number of data points in this sub - region, is the i - th data point belonging to this sub - region.
[0106] By the above method of dividing data points into sub - regions, the purpose of reducing the computational complexity can be achieved. Specifically, the city - level neural radiance field model needs to process a large amount of sensor data (such as images, camera poses, etc.). Directly modeling the entire city will lead to an excessive amount of calculation. By dividing the city into multiple sub - regions, the problem can be decomposed into multiple smaller sub - problems, and each sub - region is processed independently, thereby reducing the computational complexity.
[0107] In addition, the model accuracy can be improved. The urban environments in different regions may have different characteristics (such as building density, road structure, traffic flow, etc.). By dividing the city into sub - regions, more refined modeling can be carried out according to the characteristics of each sub - region, thereby improving the accuracy and adaptability of the model.
[0108] Moreover, local updates can be supported. The urban environment is dynamically changing (such as road construction, new buildings, etc.). By dividing sub - regions, only the sub - regions that have changed need to be updated, rather than reconstructing the entire city model, thus saving computational resources.
[0109] Furthermore, considering that each sub-region (urban area divided by K-Means clustering) may still contain a large number of camera pose data points, in order to further improve the efficiency of data management and query, the sub-region is further divided into multiple sub-fields. Each sub-field represents a smaller spatial unit within the sub-region and has its own centroid (i.e., the central position of the sub-field).
[0110] Similarly, the K-Means clustering algorithm is executed on the data points within each sub-region, clustering of sub-fields is performed based on the pose data points of the camera devices, and the centroid of each sub-field is calculated. In this way, during the rendering process, the sub-field intersecting with the query ray can be quickly located.
[0111] First, input the pose data points within the sub-region, then set the number of clusters K' to the number of sub-fields, then apply the K-means clustering algorithm, iterate and optimize until convergence, and finally output the clustering result, that is, the centroid of the sub-field. Each clustering center represents the center of a sub-field.
[0112] The formula for calculating the centroid of the sub-field is as follows:
[0113]
[0114] where, is the centroid of the j-th sub-field, is the number of data points in the sub-field, is the i-th data point belonging to the sub-field.
[0115] By clustering and dividing sub-fields in the above way, the efficiency of data management and rendering can be further improved. During the rendering process, it is necessary to quickly locate the region intersecting with the query ray. If the sub-region is too large, the query efficiency will decrease. The division of sub-fields enables the data within each sub-region to be managed and stored more precisely. Each sub-field can be stored and queried independently, reducing the computational amount of data processing.
[0116] During the rendering process of the Neural Radiance Field (NeRF), it is necessary to sample multiple points along the ray and query the features of these points. By dividing sub-fields, the sub-field intersecting with the query ray can be quickly located, thereby reducing unnecessary calculations and improving the rendering efficiency.
[0117] Based on the above, a multi-resolution hash grid H can be defined for each sub-field to store and query spatial features. The resolution of the grid is adjusted according to the computational efficiency and representation accuracy.
[0118] Point sampling of the ray directions of each camera device can be understood as sequentially sampling along the shooting perspective direction of the camera device on the historical images.
[0119] In addition, a loss function is constructed in this embodiment, and the training of the city-level neural radiance field model is a training process guided by the loss function, aiming at minimizing the difference between the output information of the model for each sampling point and the actual information of each sampling point in its respective sub-field.
[0120] Among them, the information input into the model includes the 3D position information, view direction, and video identifier of each sampling point. The 3D position information (x) refers to the three-dimensional coordinates (x, y, z) of a certain sampling point in the scene, representing the specific position of the sampling point in space. In the Neural Radiance Field (NeRF), the 3D position information is used to describe the spatial position of a certain sampling point in the scene.
[0121] The view direction (d) refers to the direction vector from the camera device (or viewpoint) to this 3D position, usually represented by a three-dimensional vector (such as pitch angle, yaw angle, etc.). The view direction is used to describe the direction of the light from the camera device to this sampling point, helping the model understand the perspective information of the light.
[0122] The video identifier (vid) refers to a unique identifier used to distinguish different historical images or scenes. Since the city-level NeRF model may be constructed based on multiple historical images, the video identifier is used to distinguish the lighting conditions, scene features, etc. in different historical images, ensuring that the model can process specific information in different historical images.
[0123] There is a multi-layer perceptron in the city-level neural radiance field model, and the training of the city-level neural radiance field model includes the training of the multi-layer perceptron. Among them, the input information includes the above-mentioned 3D position information, view direction, and video identifier, and the output information includes density information, color information, and semantic feature information.
[0124] Density information The prediction formula is as follows:
[0125]
[0126] Color information The prediction formula is as follows:
[0127]
[0128] Among them, represents the spherical harmonic encoding of the direction, represents the cross-video embedding for considering video-specific lighting conditions, ) It is used to predict color information. Among them, cross-video embedding captures specific conditions under each historical image by assigning a unique embedding vector to each historical image. This embedding vector is input into a multi-layer perceptron to help the model distinguish the lighting and scene differences in different images. That is to say, cross-video embedding represents specific conditions (such as lighting, weather, etc.) in different images through the embedding vector, helping the model distinguish the lighting differences in different images, so as to more accurately predict colors.
[0129] Semantic feature information The prediction formula is as follows:
[0130]
[0131] Among them, ) It is used to predict semantic feature information.
[0132] This process models semantic features with an invariant direction. Generally speaking, the final output of the multi-layer perceptron can be expressed as follows:
[0133]
[0134] In addition, each sampling point is assigned to the corresponding sub-field. The assignment method is based on the distance between the sampling point and the centroid of each sub-field. The sub-field with the smallest distance is determined as the sub-field to which the sampling point belongs. The assignment process is defined as follows:
[0135]
[0136] Among them, represents the centroid of the j-th sub-field.
[0137] Since the hash grid H of each sub-field stores spatial features, and this spatial feature includes density information, color information, and semantic feature information.
[0138] Therefore, for each sampling point, determine the sub-field to which it belongs, and then query the corresponding density information, color information, and semantic feature information from the hash grid of the sub-field to which it belongs, which can be characterized as follows:
[0139]
[0140] In addition, for the part related to the sky in each historical image, the prediction of color information and semantic feature information is mainly considered. Therefore, for the sky feature, its output color information and semantic feature information can be characterized as follows:
[0141]
[0142] The output of the model for each sampling point is finally rendered, which is an integration along the light direction of the camera device, incorporating the integrated color information and semantic feature information , which can be characterized as follows:
[0143]
[0144]
[0145] and respectively represent the cumulative transmittance and per-slice opacity, and the calculation formulas are as follows:
[0146]
[0147]
[0148] In addition, the loss function constructed in this embodiment is as follows:
[0149]
[0150] For color and semantic feature reconstruction the L2 loss is adopted, and the binary cross-entropy loss is adopted for sky features . In addition, an inter-layer loss and a distortion loss are introduced to further improve the performance of the model. Among them, , , , respectively represent weight coefficients.
[0151] Guided by the constructed loss function, compare the density information, color information, and semantic feature information of the sampling points output by the model, as well as the density information, color information, and language information saved in the hash grid of the sub-field to which the sampling points belong, and perform multiple iterative trainings. When the final loss function converges, or the number of iterations reaches the preset maximum number, or the iteration duration reaches the preset maximum duration, etc., it can be determined that the preset requirements are met, and at this time, a trained city-level neural radiance field model can be obtained.
[0152] In this embodiment, in order to avoid potential moving objects, such as vehicles, pedestrians, bicycles, etc., which may introduce noise during the rendering process because their positions and appearances change over time. Therefore, in this embodiment, the following processing methods can also be added during the process of training the city-level neural radiance field model:
[0153] Identify the pixel points of potential moving objects in each historical image through a semantic segmentation mask; set the pixel values of the identified pixel points to a specific mask value.
[0154] Among them, the set mask value can be 0 or the background value, so that during the process of the model processing each sampling point in the historical image, the influence of these pixel points can be ignored.
[0155] On the basis of training the city-level neural radiance field model, extract the city-level neural radiance field prior information based on the city-level neural radiance field model. Please refer to Figure 5 , the above step S12 can be implemented through the following steps:
[0156] S121, Obtain multiple prior images collected by camera devices on multiple vehicles within the target city.
[0157] S122, Import the multiple prior images into the pre-trained city-level neural radiance field model, and for each of the prior images, collect multiple sampling points in the projection light direction of the camera device corresponding to the prior image.
[0158] S123, Screen out surface points from multiple sampling points based on the cumulative transmittance and opacity metrics.
[0159] S124, Construct the city-level neural radiance field prior information based on multiple surface points corresponding to the multiple prior images. The city-level neural radiance field prior information includes road structure, traffic signs, and obstacle positions.
[0160] Among them, the prior image is also an image collected by camera devices on multiple vehicles during a historical period.
[0161] For each prior image, adopt a processing method similar to the above historical image, and sample multiple sampling points according to the light direction of the camera device. Specifically, select N sampling points along the light direction of the camera device . By determining the sampling points where the cumulative transmittance and opacity first exceed the set threshold, they are determined as surface points, and the representation method is as follows:
[0162]
[0163] After successfully identifying each surface point and obtaining their corresponding semantic features, summarize all the surface points from the prior images. In order to streamline the point cloud data, a voxel-based downsampling method is adopted to calculate the mean value of the features within each voxel. In this way, a series of feature-rich voxels can be obtained, which will be used as the city-level neural radiance field prior information to enhance the robustness of the online perception model, including road structure, traffic signs, obstacle positions, etc.
[0164] It should be noted that this extraction process is only executed once after the construction of the city-level neural radiance field model, and the extracted prior information will be properly stored. Therefore, neither the slow rendering process nor the large NeRF model will have any impact on the online perception model used to obtain real-time feature information. During actual deployment, the online perception model only needs to store and use this pre-extracted prior information.
[0165] Finally, in order to effectively fuse this prior information with the real-time features extracted by the online perception model, the same processing method as the features of the real-time image mentioned above is adopted to convert the prior information into prior information features from a bird's-eye view.
[0166] After obtaining the real-time feature information of the real-time image and the prior feature information of the prior information in the above manner, the real-time feature information and the prior feature information are imported into the pre-trained autonomous driving planning model to obtain autonomous driving planning information.
[0167] Among them, the autonomous driving planning model includes multiple modules, such as an efficient feature fusion module, a trajectory formation module, a map formation module, a motion prediction module, and a planning module, etc. Please refer to Figure 6 , and the implementation method of the pre-trained autonomous driving planning model is introduced as follows:
[0168] S31, collect training samples, where the training samples are the fusion results of images and prior information.
[0169] S32, import the training samples into the autonomous driving planning model, and inject a set of noises into each module of the autonomous driving planning model, and then train under the guidance of the constructed loss function.
[0170] S33, in each round of iteration, calculate the weight values of each module in the autonomous driving planning model in the next round of iteration, and calculate the total loss function value according to the weight values of each module in the next round of iteration until the preset requirements are met, and the trained autonomous driving planning model is obtained.
[0171] Since the information input into the autonomous driving planning model subsequently is the combination of real-time image information and prior information, therefore, during the training stage of the autonomous driving planning model, the fusion results of images and prior information are used as training samples.
[0172] To enhance the robustness of the model, a noise set can be injected into each module. This noise set is an artificially added noise used to simulate sensor (such as camera device) errors, environmental interference, or malicious attacks. By adding noise, the model can learn how to cope with noise and interference during the training process, thereby improving its robustness in practical applications. Various extreme situations (such as sensor failures, sudden environmental changes, etc.) can also be simulated to ensure that the model can still maintain stable performance under these circumstances.
[0173] During the training process, although each module has different weights in the loss function, the overall goal is ultimately guiding, rather than the loss of each module. This approach ensures that noise is generated from the overall view of the model, that is, using the overall loss for backpropagation, rather than focusing on the individual module losses that may be contradictory and have a negative impact on the robustness of the overall decision-making.
[0174] First, define the model output, which represents on the basis of the input data, after injecting the noise set ={ , , } into each module, the final output result is specifically achieved through the following function combination:
[0175]
[0176] where , , successively represent the specific adversarial perturbations (i.e., the noise set) injected into the m-th perception module ( ), the k-th prediction module ( ), and the planning module ( ).
[0177] Secondly, in order to find the optimal amount of noise injection, the following optimization problem needs to be solved, that is:
[0178] =
[0179] where C is the constraint set of the noise, is the total loss function, is the true label, that is, the autonomous driving planning information executed in the actual situation.
[0180] To manage the different contributions of each module during the training process, dynamic weight accumulation adaptation is introduced, which adaptively adjusts the loss weight of each module to the overall objective according to its contribution during the noise injection process. This method introduces a normalization weight function to eliminate the dimension difference, accelerate the model convergence speed, reduce the impact of outliers on model training, and improve its stability and model generalization ability.
[0181] First, to extend the concept of multi-tasking to multiple modules, the loss of each module at the current time step t The ratio relative to the previous value is calculated as:
[0182]
[0183] where is the loss of module j at time step , is the loss of module j at time step , is the loss ratio of module j at time step t.
[0184] Then, based on these loss ratios, the normalization weight formula is used to update the weights:
[0185]
[0186] where is the normalized weight of module at time step , is the average of the loss ratios of all modules at time step , ( ) is the normalization function, is the total number of modules.
[0187] Finally, the total loss at time step is calculated according to the updated weights:
[0188] =
[0189] This method can ensure that the weights dynamically adapt to the performance of each module over time, thereby improving stability and overall performance.
[0190] On the basis of training the autonomous driving planning model in the above manner, please refer to Figure 7 . In the application stage, by the following method, the autonomous driving planning model is used to obtain the autonomous driving planning information for the target vehicle based on the real-time feature information and the prior feature information:
[0191] S141. Import the real-time feature information and the prior feature information into a pre-trained autonomous driving planning model, and perform fusion processing on the real-time feature information and the prior feature information to obtain fused features.
[0192] S142. Obtain perception information, tracking information, map construction information, and motion prediction information based on the fused features.
[0193] S143. Combine the perception information, tracking information, map construction information, and motion prediction information to generate autonomous driving planning path information.
[0194] As can be seen from the above, the autonomous driving planning module includes an efficient feature fusion module, a trajectory formation module, a map formation module, a motion prediction module, a planning module, etc.
[0195] The efficient feature fusion module is responsible for fusing the prior feature information provided by the city-level neural radiance field model with the real-time feature information captured by the vehicle's sensors in real time. This fusion process not only has a light computational load and hardly adds extra computational overhead, but also cleverly avoids any modification to the original model architecture. This fusion strategy injects rich environmental context information into the autonomous driving system in an extremely efficient manner, thereby greatly improving its adaptability and response speed to dynamic and changing driving scenarios while maintaining the simplicity of the system.
[0196] On the basis of fusing the prior feature information and the real-time feature information, the trajectory formation module is further upgraded to use these enhanced features to more accurately identify newly emerging obstacles and continuously and stably track the detected targets. Through a set of carefully designed tracking query vectors, the integration of detection and multi-object tracking can be achieved, ensuring the real-time performance and accuracy of the system.
[0197] The map formation module uses map query vectors to more finely segment various map elements such as lane lines, sidewalks, intersections, etc. These map elements fused with prior information provide detailed and accurate environmental background information for subsequent path planning and decision-making.
[0198] On this basis, the motion prediction module deeply analyzes the complex interactions between objects and the environment, accurately predicts the future trajectories of each dynamic entity, generates multi-modal future action predictions, and provides forward-looking environmental change information for the planning module.
[0199] Finally, after receiving the comprehensive information provided by the foregoing modules, the planning module comprehensively considers the future trajectory prediction of the object and the state of the host vehicle, and generates optimal planned path information to ensure that the autonomous driving vehicle can drive safely and efficiently in a complex and changeable traffic environment. The close cooperation of this series of modules jointly constitutes the core architecture of the end-to-end autonomous driving system, realizing the full-automatic process from environmental perception to decision-making and planning.
[0200] Based on the same inventive concept, please refer to Figure 8 , this embodiment of the present invention also provides a schematic diagram of functional modules of an autonomous driving planning device combining prior information of a city-level neural radiance field. This embodiment can divide the functional modules of the autonomous driving planning device according to the above method embodiment. For example, each functional module can be corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in this embodiment of the present invention is illustrative, only a logical function division, and there may be other division methods in actual implementation.
[0201] For example, in the case of dividing each functional module corresponding to each function, Figure 8 The shown autonomous driving planning device is only a schematic diagram of a device. The autonomous driving planning device may include a first extraction module, a second extraction module, a conversion module, and a processing module. The functions of each functional module of the autonomous driving planning device will be elaborated in detail below.
[0202] The first extraction module is used to obtain real-time images collected by a camera device on a target vehicle, extract features from the real-time images, and convert the extracted features into real-time feature information in a bird's-eye view.
[0203] The second extraction module is used to extract prior information of a city-level neural radiance field based on a pre-trained city-level neural radiance field model.
[0204] The conversion module is used to convert the prior information of the city-level neural radiance field into prior feature information in a bird's-eye view.
[0205] The processing module is used to import the real-time feature information and the prior feature information into a pre-trained autonomous driving planning model, and perform fusion processing on the real-time feature information and the prior feature information to obtain fusion features.
[0206] The autonomous driving planning device provided in this embodiment can be used to execute the autonomous driving planning method in any implementation manner in the above embodiment. For details not elaborated in this embodiment, reference can be made to the corresponding description in the above embodiment, and this embodiment will not be repeated here.
[0207] Please refer to Figure 9 FIG. Figure 9 , which is a structural block diagram of the electronic device provided by the embodiment of the present invention. The electronic device may be a computer device, a server, etc. in an autonomous driving control platform. The electronic device includes a memory, a processor, and a communication module. Each element of the memory, the processor, and the communication module is electrically connected directly or indirectly to each other to realize data transmission or interaction. For example, these elements may be electrically connected to each other through one or more communication buses or signal lines.
[0208] Among them, the memory is used to store computer programs or data. The memory may be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read only memory (PROM), an erasable programmable read only memory (EPROM), an electrically erasable programmable read only memory (EEPROM), etc.
[0209] The processor is used to read / write the data or programs stored in the memory and execute the autonomous driving planning method combining the prior information of the urban-level neural radiance field provided by any embodiment of the present invention.
[0210] The communication module is used to establish a communication connection between the electronic device and other communication terminals through a network and is used to receive and transmit data through the network.
[0211] It should be understood that Figure 9 the structure shown is only a schematic structural diagram of the electronic device, and the electronic device may further include more or fewer components than those shown in Figure 9 or have a different configuration from that shown in Figure 9 FIG. Figure 9 .
[0212] Furthermore, the embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are executed, the autonomous driving planning method combining the prior information of the urban-level neural radiance field provided by the above embodiment is realized.
[0213] Specifically, the computer-readable storage medium can be a general storage medium, such as a removable disk, a hard disk, etc. When the computer program on the computer-readable storage medium is run, it can execute the above-mentioned autonomous driving planning method combined with the prior information of the urban-level neural radiance field. Regarding the process involved when the computer and its executable instructions in the computer-readable storage medium are run, reference can be made to the relevant descriptions in the above method embodiments, which will not be elaborated here.
[0214] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces, and the indirect coupling or communication connection of the devices or units can be in electrical, mechanical or other forms.
[0215] In addition, the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0216] Furthermore, in each embodiment of the present invention, the functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.
[0217] It should be noted that if the function is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or this part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other various media that can store program codes.
[0218] In this document, relational terms such as first and second are used solely to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.
[0219] The above description is only for the embodiments of the present invention and is not intended to limit the protection scope of the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An autonomous driving planning method that combines prior information of a city-level neural radiance field, characterized in that The method includes: Obtain a real-time image collected by a camera device on a target vehicle, extract features from the real-time image, and convert the extracted features into real-time feature information in a bird's-eye view; Extract urban-level neural radiance field prior information based on a pre-trained urban-level neural radiance field model; Convert the urban-level neural radiance field prior information into prior feature information in a bird's-eye view; Import the real-time feature information and the prior feature information into a pre-trained autonomous driving planning model, and output autonomous driving planning information for the target vehicle; The method further includes the step of pre-training the urban-level neural radiance field model, and this step includes: Obtain a plurality of historical images collected by camera devices on multiple vehicles within a target city during a historical period; perform point sampling based on the light directions of the camera devices corresponding to the respective historical images to obtain a plurality of sampling points; for each of the sampling points, obtain the 3D position information, view direction, and video identifier of the sampling point; import the 3D position information, view direction, and video identifier corresponding to each sampling point into a constructed initial model, perform processing based on the initial model, and output corresponding density information, color information, and semantic feature information; guided by a constructed loss function, based on the density information, color information, and semantic feature information output by the initial model, and the density information, color information, and semantic feature information corresponding to the area information to which the sampling point belongs, perform multiple iterative trainings on the initial model until a trained urban-level neural radiance field model is obtained when a preset requirement is met; The step of extracting urban-level neural radiance field prior information based on a pre-trained urban-level neural radiance field model includes: Obtain a plurality of prior images collected by camera devices on multiple vehicles within a target city; import the plurality of prior images into a pre-trained urban-level neural radiance field model, and for each of the prior images, collect a plurality of sampling points for the projection light direction of the camera device corresponding to the prior image; screen out surface points from the plurality of sampling points based on the cumulative transmittance and opacity metrics; construct urban-level neural radiance field prior information based on the plurality of surface points corresponding to the plurality of prior images, and the urban-level neural radiance field prior information includes road structure, traffic signs, and obstacle positions.
2. The autonomous driving planning method combining prior information of urban-level neural radiation fields according to claim 1, wherein The step of extracting features from the real-time image and converting the extracted features into real-time feature information in a bird's-eye view includes: Perform a convolution operation on the real-time image to extract various different types of features of the real-time image; Perform a non-linear activation function process on the various features to introduce non-linear features and encode the processed features into feature vectors; Use a geometric transformation matrix to convert the feature vector in the image coordinate system into a real-time feature vector in the bird's-eye view coordinate system.
3. The autonomous driving planning method combining prior information of urban-level neural radiation field according to claim 1, characterized in that, The step of performing point sampling based on the light directions of the camera devices corresponding to the respective historical images to obtain a plurality of sampling points includes: Map the pose information of the camera device corresponding to each historical image to a data point; Cluster multiple data points to divide them into multiple sub-regions, and then cluster the data points in each of the sub-regions to divide them into multiple sub-fields; For each of the historical images, perform point sampling on the light direction of the camera device corresponding to the historical image to obtain multiple sampling points, and determine the region information to which each of the sampling points belongs, where the region information is sub-field information.
4. The autonomous driving planning method combining prior information of urban-level neural radiation fields according to claim 3, wherein The step of clustering multiple data points to divide them into multiple sub-regions includes: Set the number of clusters for clustering based on the scale information of the target city, and randomly generate the centroids of each cluster; Based on the distances between each of the data points and the centroids of each of the clusters, divide each of the data points into the corresponding cluster, and update the centroids of the clusters according to the position information of the data points included in each cluster. After multiple iterations, until the centroids of each cluster reach stability, obtain multiple sub-regions based on the multiple clusters obtained.
5. The autonomous driving planning method incorporating prior information of the urban-level neural radiation field according to claim 3, characterized in that The step of pre-training to obtain the city-level neural radiance field model further includes: Identify the pixel points of potential moving objects in each of the historical images through a semantic segmentation mask; Set the pixel values of the identified pixel points to specific mask values.
6. The autonomous driving planning method combining prior information of urban-level neural radiation field according to claim 1, wherein The autonomous driving planning model includes multiple modules, and the method further includes the step of pre-training to obtain the autonomous driving planning model, and this step includes: Collect training samples, where the training samples are the fusion results of images and prior information; Import the training samples into the autonomous driving planning model, and after injecting a noise set into each module of the autonomous driving planning model, perform training under the guidance of the constructed loss function; In each round of iteration, calculate the weight values of each module in the autonomous driving planning model in the next round of iteration, and calculate the total loss function value according to the weight values of each module in the next round of iteration until the preset requirements are met, and obtain the trained autonomous driving planning model.
7. The autonomous driving planning method combining prior information of urban-level neural radiation field according to claim 1, wherein The step of importing the real-time feature information and the prior feature information into the pre-trained autonomous driving planning model and outputting the autonomous driving planning information for the target vehicle includes: Import the real-time feature information and the prior feature information into the pre-trained autonomous driving planning model, and perform fusion processing on the real-time feature information and the prior feature information to obtain fusion features; Obtain perception information, tracking information, map construction information, and motion prediction information based on the fusion features; Combine the perception information, tracking information, map construction information, and motion prediction information to generate autonomous driving planning path information.
8. An autonomous driving planning device that combines prior information of a city-level neural radiance field, characterized in that, For implementing the autonomous driving planning method combining city-level neural radiance field prior information according to any one of claims 1-7, the device includes: A first extraction module, configured to obtain a real-time image collected by a camera device on a target vehicle, perform feature extraction on the real-time image, and convert the extracted features into real-time feature information in a bird's-eye view; A second extraction module, configured to extract city-level neural radiance field prior information based on a pre-trained city-level neural radiance field model; A conversion module for converting the prior information of the city-level neural radiance field into prior feature information from a bird's-eye view; A processing module for importing the real-time feature information and the prior feature information into a pre-trained autonomous driving planning model, and performing fusion processing on the real-time feature information and the prior feature information to obtain fused features.
Citation Information
Patent Citations
Monocular depth estimation method based on surface normal vector and neural radiation field
CN118570273A
Methods and internet of things systems for managing traffic road cleaning in smart city
US20230386327A1