Automatic driving planning method and device in combination with prior information of city-level neural radiation field
Patent Information
- Application Number
- CN202510512113.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-23
Smart Images

Figure CN120024358A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving technology, and in particular to an autonomous driving planning method and device combining city-level neural radiation field prior information. Background Art
[0002] Autonomous driving technology, as the frontier of the integration of artificial intelligence and intelligent transportation systems, has set off a wave of research and development and application around the world in recent years. Thanks to the rapid progress in computing power, sensor technology and deep learning algorithms, autonomous driving technology has gradually moved from theoretical exploration in the laboratory to the stage of practical application, and has shown great commercial potential and value in many fields such as passenger cars, commercial vehicles, logistics and transportation, shared travel and unmanned delivery.
[0003] Automakers, technology giants and startups around the world have devoted themselves to the research and development and testing of autonomous driving systems, forming a competitive landscape where a hundred flowers bloom. Tesla has led the popularization of semi-autonomous driving technology with its Autopilot system, while Waymo has made breakthrough progress in fully driverless taxi services. At the same time, domestic companies such as Baidu, Weilai, Xiaopeng, and Ideal have also continued to innovate in the field of autonomous driving, promoting the overall development of the industry. Baidu's Apollo platform has further accelerated the pace of application of autonomous driving technology in passenger cars and commercial vehicles through large-scale testing in many cities in China.
[0004] However, behind the booming development of autonomous driving technology, there are still many technical pain points and challenges. First, the perception problem of complex environments is particularly prominent. The autonomous driving system needs to accurately identify and handle changing road environments, including traffic signs, pedestrians, obstacles, etc., especially in bad weather and low visibility conditions. The stability and reliability of the perception system face severe tests. Secondly, the construction and updating of high-precision maps is also a major problem. Autonomous driving relies on accurate map data to assist navigation and decision-making, but the real-time update of maps, coverage, and rapid changes in dynamic scenes all place higher demands on the accuracy and timeliness of maps. In addition, key technologies such as multi-target tracking and prediction, decision-making and planning also face challenges. The autonomous driving system needs to deal with multiple uncertainties in a dynamic traffic environment to ensure safe and reasonable decision-making and path planning.
[0005] In summary, there is an urgent need for a method that enables the autonomous driving system to balance accuracy, robustness and real-time performance in complex and changeable road environments. Summary of the invention
[0006] The purpose of the embodiments of the present invention is to provide an autonomous driving planning method and device that combines city-level neural radiation field prior information to ensure the safety and accuracy of autonomous driving planning.
[0007] In a first aspect, the present invention provides an autonomous driving planning method combining city-level neural radiation field prior information, the method comprising: Acquire a real-time image captured by a camera device on a target vehicle, perform feature extraction on the real-time image, and convert the extracted features into real-time feature information from a bird's-eye view; Based on the pre-trained city-level neural radiation field model, the city-level neural radiation field prior information is extracted; Converting the city-level neural radiation field prior information into prior feature information from a bird's-eye view; The real-time feature information and the prior feature information are imported into a pre-trained autonomous driving planning model, and the autonomous driving planning information for the target vehicle is output.
[0008] In an optional implementation, the step of extracting features from the real-time image and converting the extracted features into real-time feature information from a bird's-eye view includes: Performing a convolution operation on the real-time image to extract multiple different types of features of the real-time image; Performing nonlinear activation function processing on the multiple features to introduce nonlinear features, and encoding the processed features into feature vectors; The geometric transformation matrix is used to transform the feature vector in the image coordinate system into the real-time feature vector in the bird's-eye view coordinate system.
[0009] In an optional embodiment, the method further includes a step of pre-training to obtain the city-level neural radiation field model, the step comprising: Acquire multiple historical images collected by camera devices on multiple vehicles in a target city during a historical period; Performing point sampling based on the light direction of the camera device corresponding to each of the historical images to obtain a plurality of sampling points; For each of the sampling points, obtaining 3D position information, a view direction, and a video identifier of the sampling point; Importing the 3D position information, view direction and video identifier corresponding to each of the sampling points into the constructed initial model, processing based on the initial model, and outputting corresponding density information, color information and semantic feature information; Guided by the constructed loss function, based on the density information, color information and semantic feature information output by the initial model, and the density information, color information and semantic feature information corresponding to the area information to which the sampling point belongs, the initial model is iteratively trained multiple times until a trained city-level neural radiation field model is obtained when the preset requirements are met.
[0010] In an optional implementation, the step of performing point sampling based on the light direction of the camera device corresponding to each of the historical images to obtain a plurality of sampling points includes: Mapping the position information of the camera device corresponding to each of the historical images into a data point; Performing clustering processing on the multiple data points to divide the multiple data points into multiple sub-areas, and then performing clustering processing on the data points in each of the sub-areas to divide them into multiple sub-fields; For each of the historical images, point sampling is performed on the light direction of the camera device corresponding to the historical image to obtain a plurality of sampling points, and the region information to which each of the sampling points belongs is determined, where the region information is subfield information.
[0011] In an optional implementation, the step of clustering the multiple data points to divide the multiple data points into multiple sub-areas includes: The number of clusters is set based on the size information of the target city, and the centroid of each cluster is randomly generated; Based on the distance between each data point and the centroid of each cluster, each data point is divided into a corresponding cluster, and the centroid of the cluster is updated according to the position information of the data points contained in each cluster. After multiple iterations, until the centroid of each cluster is stable, multiple sub-areas are obtained based on the obtained multiple clusters.
[0012] In an optional implementation, the step of pre-training to obtain the city-level neural radiation field model further includes: Identifying pixels of potential moving objects in each of the historical images through a semantic segmentation mask; The pixel value of the identified pixel is set to a specific mask value.
[0013] In an optional implementation, the step of extracting city-level neural radiation field prior information based on a pre-trained city-level neural radiation field model includes: Acquire multiple prior images captured by camera devices on multiple vehicles in the target city; Importing the multiple prior images into a pre-trained city-level neural radiation field model, and for each of the prior images, collecting multiple sampling points in the direction of the projection light of the camera device corresponding to the prior image; Filtering surface points from multiple sampling points based on cumulative transmittance and opacity indicators; Based on multiple surface points corresponding to multiple prior images, city-level neural radiation field prior information is constructed, and the city-level neural radiation field prior information includes road structure, traffic signs, and obstacle locations.
[0014] In an optional embodiment, the autonomous driving planning model includes a plurality of modules, and the method further includes a step of pre-training to obtain the autonomous driving planning model, the step comprising: Collecting training samples, where the training samples are the fusion results of images and prior information; Importing the training samples into the autonomous driving planning model, and injecting a noise set into each module of the autonomous driving planning model, and then training with the constructed loss function as a guide; In each round of iteration, the weight values of each module in the autonomous driving planning model in the next round of iteration are calculated, and the total loss function value is calculated according to the weight values of each module in the next round of iteration, until the preset requirements are met, and the trained autonomous driving planning model is obtained.
[0015] In an optional embodiment, the step of importing the real-time feature information and the prior feature information into a pre-trained autonomous driving planning model and outputting the autonomous driving planning information for the target vehicle includes: Importing the real-time feature information and the prior feature information into a pre-trained autonomous driving planning model, and fusing the real-time feature information and the prior feature information to obtain a fused feature; Obtaining perception information, tracking information, map construction information and motion prediction information based on the fused features; The perception information, tracking information, map construction information and motion prediction information are combined to generate autonomous driving planning path information.
[0016] In a second aspect, the present invention provides an autonomous driving planning device in combination with city-level neural radiation field prior information, the device comprising: A first extraction module is used to obtain a real-time image captured by a camera device on a target vehicle, perform feature extraction on the real-time image, and convert the extracted features into real-time feature information from a bird's-eye view; The second extraction module is used to extract the city-level neural radiation field prior information based on the pre-trained city-level neural radiation field model; A conversion module, used to convert the city-level neural radiation field prior information into prior feature information under a bird's-eye view; A processing module is used to import the real-time feature information and the prior feature information into a pre-trained autonomous driving planning model, and output autonomous driving planning information for the target vehicle.
[0017] The present invention provides an automatic driving planning method and device combined with city-level neural radiation field prior information, by extracting features from the real-time image acquired by the camera device on the target vehicle, converting the extracted features into real-time feature information from a bird's-eye view, and extracting the city-level neural radiation field prior information based on the pre-trained city-level neural radiation field model. The city-level radiation field prior information is converted into prior feature information from a bird's-eye view. Finally, the real-time feature information and the prior feature information are imported into the pre-trained automatic driving planning model, and the automatic driving planning information for the target vehicle is output. In this solution, automatic driving planning is performed in combination with real-time images and city-level neural radiation field prior information, which can accurately perceive the static and dynamic environments around the vehicle and ensure the safety of control. Moreover, the real-time features and the prior features are converted into features from a bird's-eye view, so that planning can be performed from a global perspective to ensure the accuracy of planning. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments of the present invention are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0019] Figure 1 A flowchart of an autonomous driving planning method provided by an embodiment of the present invention; Figure 2 for Figure 1 A flowchart of the sub-steps included in S11; Figure 3 A flowchart of a city-level neural radiation field model training method in an autonomous driving planning method provided in an embodiment of the present invention; Figure 4 for Figure 3 A flowchart of the sub-steps included in S22; Figure 5 for Figure 1 A flowchart of the sub-steps included in S12; Figure 6 A flowchart of an autonomous driving planning model training method in an autonomous driving planning method provided in an embodiment of the present invention; Figure 7 for Figure 1 A flowchart of the sub-steps included in S14; Figure 8 A functional module block diagram of an autonomous driving planning device provided by an embodiment of the present invention; Fig. 9 This is a structural block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0020] The technical solutions in the embodiments of the present invention will be described below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0021] See also Figure 1 , is a flow chart of an autonomous driving planning method combined with city-level neural radiation field prior information provided by an embodiment of the present invention, the autonomous driving planning method can be executed by an autonomous driving planning device, the autonomous driving planning device can be implemented by software and / or hardware, and can be configured in an electronic device, the electronic device can be a computer device, a server, etc., for example, it can be a server in a back-end control platform, etc. The detailed steps of the autonomous driving planning method are described as follows.
[0022] S11, obtaining a real-time image captured by a camera device on a target vehicle, performing feature extraction on the real-time image, and converting the extracted features into real-time feature information from a bird's-eye view.
[0023] S12, extracting city-level neural radiation field prior information based on the pre-trained city-level neural radiation field model.
[0024] S13, converting the city-level neural radiation field prior information into prior feature information from a bird's-eye view.
[0025] S14, importing the real-time feature information and the prior feature information into a pre-trained autonomous driving planning model, and outputting autonomous driving planning information for the target vehicle.
[0026] In this embodiment, the so-called target vehicle refers to any vehicle that needs to be controlled by automatic driving. Camera equipment is installed on the body and bottom of the target vehicle. During the driving process of the target vehicle, the camera equipment can continuously capture images within the field of view of the device, which are named real-time images here.
[0027] In order to perform autonomous driving planning from a global perspective, the features of the real-time image are converted into real-time feature information from a bird's-eye view.
[0028] In addition, in this embodiment, a city-level neural radiation field model is pre-trained, based on which historical images can be processed to extract city-level neural radiation field prior information. Historical images refer to images captured by cameras on multiple vehicles in the city during a historical period. Prior information includes road structure, traffic signs, obstacle locations, etc. This type of prior information will affect the vehicle's autonomous driving planning.
[0029] In order to effectively integrate the prior information with the real-time image information, the prior information is converted into prior feature information from a bird's-eye view.
[0030] In addition, in this embodiment, an autonomous driving planning model is pre-trained. The autonomous driving planning model can output autonomous driving planning information for the target vehicle under the real-time feature information based on the input real-time feature information and prior feature information, such as autonomous driving path information.
[0031] In this embodiment, the automatic driving planning is performed by combining real-time images and city-level neural radiation field prior information, which can accurately perceive the static and dynamic environments around the vehicle and ensure the safety of control. In addition, the real-time features and prior features are converted into features from a bird's-eye view, so that planning can be performed from a global perspective to ensure the accuracy of the planning.
[0032] The following first introduces the processing process of the acquired real-time image: See also Figure 2 , the above step S11 may include the following sub-steps: S111, performing a convolution operation on the real-time image to extract a plurality of different types of features of the real-time image.
[0033] S112, performing nonlinear activation function processing on the multiple features to introduce nonlinear features, and encoding the processed features into feature vectors.
[0034] S113, using a geometric transformation matrix to convert the feature vector in the image coordinate system into a real-time feature vector in the bird's-eye view coordinate system.
[0035] In this embodiment, the residual network ResNet can be used to process the real-time image. In the feature extraction process, ResNet performs a convolution operation on the input real-time image through its convolution layer, and each convolution kernel is responsible for extracting a specific type of feature from the real-time image, such as edge, texture, shape, etc. These convolution kernels slide on the image, and extract different features of the image through weighted summation and bias processing. Subsequently, these features are processed by a nonlinear activation function to introduce nonlinear characteristics. Finally, these features are encoded as feature vectors, which provide a basis for subsequent image analysis and understanding.
[0036] In the process of bird’s-eye view feature conversion, the extracted image feature vector Further processing to generate real-time feature vectors from a bird's eye view To this end, we first define a geometric transformation matrix T, which is responsible for converting the image coordinate system into the bird's-eye view coordinate system. The geometric transformation matrix T contains transformation parameters such as rotation, scaling, and translation, which can accurately map each point in the image to the corresponding position in the bird's-eye view.
[0037] For each pixel (u, v) in the real-time image and its corresponding feature vector (u,v), and calculate its position in the bird's-eye view coordinate system through the transformation matrix T ( ). This transformation can be expressed as:
[0038] Among them, T is a 3×3 homogeneous transformation matrix, which includes the rotation matrix R, the scaling factor S and the translation vector t.
[0039] After the coordinate transformation is completed, the real-time feature vector (u,v) is mapped to the corresponding position in the bird's-eye view coordinate system ( ), forming a bird's-eye view feature map . This process can be expressed as: ( ) (u,v) After this conversion process, the vehicle can obtain a more comprehensive and unified representation of the environment (i.e. ), which not only covers the spatial layout and object distribution of the surrounding environment, but also effectively eliminates the interference caused by changes in perspective. Such environmental representation enables the autonomous driving system to grasp the surrounding environment more accurately, and then make more intelligent decisions in terms of path planning, obstacle recognition and avoidance.
[0040] In this embodiment, the real-time feature vector of the real-time image is combined with the city-level neural radiation field prior information extracted based on the pre-trained city-level neural radiation field model, and used together for autonomous driving planning.
[0041] The following first introduces the implementation method of the pre-trained city-level neural radiation field model: See also Figure 3 In this embodiment, the urban neural radiation field model is pre-trained in the following manner: S21, obtaining a plurality of historical images captured by camera devices on a plurality of vehicles in a target city within a historical period.
[0042] S22, performing point sampling based on the light direction of the camera device corresponding to each of the historical images to obtain a plurality of sampling points.
[0043] S23: For each of the sampling points, obtain the 3D position information, the viewing direction and the video identifier of the sampling point.
[0044] S24, importing the 3D position information, view direction and video identifier corresponding to each sampling point into the constructed initial model, performing processing based on the initial model, and outputting corresponding density information, color information and semantic feature information.
[0045] S25, guided by the constructed loss function, based on the density information, color information and semantic feature information output by the initial model, and the density information, color information and semantic feature information corresponding to the area information to which the sampling point belongs, the initial model is iteratively trained for multiple times until a trained city-level neural radiation field model is obtained when the preset requirements are met.
[0046] In this embodiment, the target city refers to the city range where the target vehicle is located, and the multiple vehicles refer to the vehicles communicating with the back-end control platform. The server in the back-end control platform can obtain the historical images collected by the camera equipment of the multiple vehicles in the historical period.
[0047] During implementation, the collected historical images should contain as many scenes as possible under different lighting conditions (day or night), weather (sunny or rainy and snowy), etc. At the same time, the collected historical images should contain most of the road scenes in the target city, meet a certain degree of spatial coverage, and be able to reflect the dynamic changes of the target city in the time domain and space domain. In addition, in order to enhance the scene representation ability of image data, the pre-trained convolutional neural network can be used to extract features from each historical image to generate pixel-level image descriptors. These descriptors can efficiently extract edges, textures, color distributions, and higher-level semantic information in the image through layer-by-layer convolution and pooling operations of deep convolutional neural networks, such as lane lines, traffic signs, pedestrians, vehicles and other specific objects and their relationships. In addition, they can also capture dynamic features such as environmental color differences caused by seasonal changes, different road conditions caused by weather changes, and the density and behavior patterns of pedestrian and vehicle flows in different time periods. These detailed and expressive features lay a solid foundation for subsequent advanced autonomous driving tasks such as image classification, target detection, scene understanding, and path planning.
[0048] In order to improve the data processing efficiency and the accuracy of model building, in implementation, the sub-areas and sub-fields can be divided based on the posture of the camera device corresponding to each historical image, that is, the posture of the camera device itself when collecting historical images. For details, please refer to Figure 4 , the above step S22 can be implemented in the following way: S221, mapping the posture information of the camera device corresponding to each of the historical images into a data point.
[0049] S222, clustering the multiple data points to divide the multiple data points into multiple sub-areas, and then clustering the data points in each of the sub-areas to divide them into multiple sub-fields.
[0050] S223, for each of the historical images, perform point sampling on the light direction of the camera device corresponding to the historical image to obtain multiple sampling points, and determine the area information to which each of the sampling points belongs, where the area information is subfield information.
[0051] First, the posture information of each camera device when collecting each historical image is mapped into a data point, and each data point contains the spatial position and viewing direction information of the camera device.
[0052] When dividing the sub-areas, the number of clusters is set based on the scale information of the target city, and the centroid of each cluster is randomly generated; based on the distance between each data point and the centroid of each cluster, each data point is divided into the corresponding cluster, and the centroid of the cluster is updated according to the position information of the data points contained in each cluster. After multiple iterations, when the centroid of each cluster is stable, multiple sub-areas are obtained based on the multiple clusters obtained.
[0053] In this embodiment, the K-means clustering algorithm is applied iteratively. In each iteration, the algorithm updates the centroid of each cluster (i.e., sub-area) and redivides the clusters according to the distance between the data point and the centroid. This process is repeated many times until the division of the clusters no longer changes or the preset number of iterations is reached. Finally, the boundary of each sub-area is determined based on the clustering results. The position information of the camera devices in the same sub-area is close to each other, indicating that their positions in the urban space are relatively concentrated. The position information of different sub-areas is relatively scattered, indicating that they represent different areas of the city.
[0054] The centroid calculation formula for each sub-area is as follows:
[0055] in, is the centroid of the jth subregion, is the number of data points in the subregion, is the i-th data point belonging to this subregion.
[0056] By dividing the data points into sub-areas, the purpose of reducing computational complexity can be achieved. Specifically, the city-level neural radiation field model needs to process a large amount of sensor data (such as images, camera postures, etc.), and directly modeling the entire city will result in excessive computation. By dividing the city into multiple sub-areas, the problem can be decomposed into multiple smaller sub-problems, and each sub-area is processed independently, thereby reducing computational complexity.
[0057] In addition, the accuracy of the model can be improved. The urban environment in different areas may have different characteristics (such as building density, road structure, traffic flow, etc.). By dividing the city into sub-areas, more detailed modeling can be performed based on the characteristics of each sub-area, thereby improving the accuracy and adaptability of the model.
[0058] In addition, it can also support local updates. The urban environment is changing dynamically (such as road construction, new buildings, etc.). By dividing the sub-areas, only the changed sub-areas can be updated without rebuilding the model of the entire city, thus saving computing resources.
[0059] Furthermore, considering that each sub-area (urban area divided by K-Means clustering) may still contain a large number of camera device pose data points, in order to further improve the efficiency of data management and query, the sub-area is further divided into multiple sub-fields. Each sub-field represents a smaller spatial unit within the sub-area and has its own centroid (i.e., the center position of the sub-field).
[0060] Similarly, the K-Means clustering algorithm is performed on the data points in each sub-area, the sub-fields are clustered according to the pose data points of the camera device, and the centroid of each sub-field is calculated. In this way, during the rendering process, the sub-fields that intersect with the query ray can be quickly located.
[0061] First, input the pose data points in the sub-area, then set the number of clusters K' to the number of sub-fields, then apply the K-means clustering algorithm, iterate and optimize until convergence, and finally output the clustering result, that is, the centroid of the sub-field. Each cluster center represents the center of a sub-field.
[0062] The centroid of a subfield is calculated as follows:
[0063] in, is the centroid of the jth subfield, is the number of data points in this subfield, is the i-th data point belonging to this subfield.
[0064] The above clustering and sub-field division method can further improve the efficiency of data management and rendering. During the rendering process, it is necessary to quickly locate the area that intersects with the query ray. If the sub-area is too large, the query efficiency will be reduced. The division of sub-fields allows the data in each sub-area to be managed and stored more finely. Each sub-field can be stored and queried independently, reducing the amount of calculation for data processing.
[0065] In the rendering process of NeRF, it is necessary to sample multiple points along the ray and query the features of these points. By dividing the subfields, the subfields that intersect with the query ray can be quickly located, thereby reducing unnecessary calculations and improving rendering efficiency.
[0066] Based on the above, a multi-resolution hash grid H can be defined for each subfield to store and query spatial features. The resolution of the grid is adjusted according to computational efficiency and representation accuracy.
[0067] Point sampling of the light direction of each camera device can be understood as sequential sampling on the historical image according to the shooting angle of view direction of the camera device.
[0068] In addition, a loss function is constructed in this embodiment, and the training of the city-level neural radiation field model is guided by the loss function, with the goal of minimizing the difference between the output information of the model for each sampling point and the actual information of each sampling point itself in the subfield to which it belongs.
[0069] The information input into the model includes the 3D position information, view direction, and video identifier of each sampling point. The 3D position information (x) refers to the three-dimensional coordinates (x, y, z) of a sampling point in the scene, indicating the specific position of the sampling point in space. In Neural Radiance Field (NeRF), 3D position information is used to describe the spatial position of a sampling point in the scene.
[0070] The view direction (d) refers to the direction vector from the camera (or viewpoint) to the 3D position, usually expressed as a three-dimensional vector (such as pitch angle, yaw angle, etc.). The view direction is used to describe the direction of light from the camera to the sampling point, helping the model understand the perspective information of light.
[0071] Video identifier (vid) refers to a unique identifier used to distinguish different historical images or scenes. Since the city-level NeRF model may be built based on multiple historical images, the video identifier is used to distinguish the lighting conditions, scene features, etc. in different historical images to ensure that the model can process specific information in different historical images.
[0072] The city-level neural radiation field model has a multi-layer perceptron, and the training of the city-level neural radiation field model includes the training of the multi-layer perceptron. The input information includes the above-mentioned 3D position information, view direction and video identifier, and the output information includes density information, color information and semantic feature information.
[0073] Density information The prediction formula is as follows:
[0074] Color Information The prediction formula is as follows:
[0075] in, Spherical harmonic code representing direction, represents a cross-video embedding that takes into account video-specific lighting conditions, ) is used to predict color information. Among them, cross-video embedding captures the specific conditions under the historical image by assigning a unique embedding vector to each historical image. This embedding vector will be input into the multi-layer perceptron to help the model distinguish the lighting and scene differences in different images. In other words, cross-video embedding represents the specific conditions (such as lighting, weather, etc.) in different images through embedding vectors, helping the model distinguish the lighting differences between different images, thereby predicting colors more accurately.
[0076] Semantic feature information The prediction formula is as follows:
[0077] in, ) is used to predict semantic feature information.
[0078] This process models direction-invariant semantic features. In general, the final output of the multilayer perceptron can be expressed as follows:
[0079] In addition, each sampling point is assigned to a corresponding subfield according to the distance between the sampling point and the centroid of each subfield, and the subfield with the smallest distance is determined as the subfield to which the sampling point belongs. The allocation process is defined as follows:
[0080] in, represents the centroid of the j-th subfield.
[0081] Since the hash grid H of each subfield stores spatial features, the spatial features include density information, color information, and semantic feature information.
[0082] Therefore, for each sampling point, determine the subfield to which it belongs, and then query the hash grid of the subfield to obtain the corresponding density information, color information, and semantic feature information, which can be represented as follows:
[0083] In addition, for the parts of each historical image involving the sky, the prediction of color information and semantic feature information is mainly considered. Therefore, for the sky feature, the output color information and semantic feature information It can be characterized as follows:
[0084] The output of the model for each sampling point is finally rendered. The rendering is the integration along the light direction of the camera device, integrating the color information of the integration. and semantic feature information , which can be characterized as follows:
[0085]
[0086] and They represent the cumulative transmittance and sheet-by-sheet opacity respectively, and the calculation formula is as follows:
[0087]
[0088] In addition, the loss function constructed in this embodiment is as follows:
[0089] Color and semantic feature reconstruction L2 loss is used, and binary cross entropy loss is used for sky features. In addition, the inter-layer loss is introduced and distortion loss , to further improve the performance of the model. , , , They represent weight coefficients respectively.
[0090] Based on the constructed loss function as a guide, the density information, color information and semantic feature information of the sampling points output by the model are compared with the density information, color information and language information stored in the hash grid of the subfield to which the sampling points belong, and multiple iterations of training are performed. When the final loss function reaches convergence, or the number of iterations reaches the preset maximum number of times, or the iteration duration reaches the preset maximum duration, it can be determined that the preset requirements are met, and the trained city-level neural radiation field model can be obtained.
[0091] In this embodiment, in order to avoid the potential moving objects, such as vehicles, pedestrians, bicycles, etc., noise may be introduced in the rendering process because their positions and appearances change over time. Therefore, in this embodiment, the following processing methods can also be added during the training of the city-level neural radiation field model: The pixel points of potential moving objects in each historical image are identified through the semantic segmentation mask; and the pixel values of the identified pixel points are set as specific mask values.
[0092] The set mask value may be 0 or a background value, so that the influence of these pixels can be ignored when the model processes each sampling point in the historical image.
[0093] Based on the training of the city-level neural radiation field model, the city-level neural radiation field prior information is extracted based on the city-level neural radiation field model. Figure 5 , the above step S12 can be implemented by the following steps: S121, obtaining a plurality of prior images captured by camera devices on a plurality of vehicles in a target city.
[0094] S122, importing the multiple prior images into a pre-trained city-level neural radiation field model, and for each of the prior images, collecting multiple sampling points in the direction of the projection light of the camera device corresponding to the prior image.
[0095] S123, selecting surface points from the plurality of sampling points based on the cumulative transmittance and the opacity index.
[0096] S124, constructing city-level neural radiation field prior information based on multiple surface points corresponding to multiple prior images, wherein the city-level neural radiation field prior information includes road structure, traffic signs, and obstacle locations.
[0097] The prior images are also images collected by camera devices on multiple vehicles during a historical period.
[0098] For each prior image, a similar processing method as the above historical image is used to sample multiple sampling points according to the light direction of the camera device. Specifically, N sampling points are selected along the light direction of the camera device. The sampling point where the cumulative transmittance and opacity first exceed the set threshold is determined as the surface point, and the characterization method is as follows:
[0099] After successfully identifying each surface point After obtaining their corresponding semantic features, all surface points from the prior image are summarized. In order to streamline the point cloud data, a voxel-based downsampling method is used to calculate the mean of the features within each voxel. In this way, a series of feature-rich voxels can be obtained, which will serve as the city-level neural radiation field prior information to improve the robustness of the online perception model, including road structure, traffic signs, obstacle locations, etc.
[0100] It is important to note that this extraction process is only performed once after the city-level neural radiation field model is built, and the extracted prior information is properly stored. Therefore, neither the slow rendering process nor the huge NeRF model will have any impact on the online perception model used to obtain real-time feature information. In actual deployment, the online perception model only needs to store and use this pre-extracted prior information.
[0101] Finally, in order to effectively integrate these prior information with the real-time features extracted by the online perception model, the same processing method as the above-mentioned features for real-time images is adopted to convert the prior information into prior information features from a bird's-eye view.
[0102] After obtaining the real-time feature information of the real-time image and the prior feature information of the prior information in the above manner, the real-time feature information and the prior feature information are imported into the pre-trained autonomous driving planning model to obtain the autonomous driving planning information.
[0103] The autonomous driving planning model includes multiple modules, such as efficient feature fusion module, trajectory formation module, map formation module, motion prediction module and planning module. Figure 6 , the following introduces the implementation method of the pre-trained autonomous driving planning model: S31, collecting training samples, where the training samples are the fusion results of images and prior information.
[0104] S32, importing the training samples into the autonomous driving planning model, injecting a noise set into each module of the autonomous driving planning model, and then performing training guided by the constructed loss function.
[0105] S33, in each round of iteration, calculate the weight value of each module in the autonomous driving planning model in the next round of iteration, and calculate the total loss function value according to the weight value of each module in the next round of iteration, until the preset requirements are met, and the trained autonomous driving planning model is obtained.
[0106] Since the information subsequently input into the autonomous driving planning model is a combination of real-time image information and prior information, in the stage of training the autonomous driving planning model, the fusion result of the image and prior information is used as a training sample.
[0107] In order to enhance the robustness of the model, a noise set can be injected into each module. This noise set is a kind of artificially added noise used to simulate sensor (such as camera equipment) errors, environmental interference or malicious attacks. By adding noise, the model can learn how to deal with noise and interference during training, thereby improving its robustness in practical applications. It can also simulate various extreme situations (such as sensor failures, sudden environmental changes, etc.) to ensure that the model can still maintain stable performance in these situations.
[0108] During training, although each module occupies a different weight in the loss function, the overall objective is ultimately guided rather than the loss of each module. This approach ensures that noise is generated with a holistic view of the model, i.e., backpropagation is performed using the overall loss rather than focusing on individual module losses that may contradict each other and negatively affect the overall decision robustness.
[0109] First, define the model output, which represents the input data Based on this, a noise set is injected into each module ={ , , }, the final output result is achieved through the following function combination:
[0110] in , , represents the injection of the mth perception module ( ), the kth prediction module ( ) and planning modules ( ) is a specific adversarial perturbation (i.e., a set of noise).
[0111] Secondly, in order to find the optimal amount of noise injection, the following optimization problem needs to be solved, namely: =
[0112] Where C is the set of noise constraints, is the total loss function, is the true label, that is, the autonomous driving planning information executed in actual situations.
[0113] In order to manage the different contributions of each module during the training process, dynamic weight accumulation adaptation is introduced to adaptively adjust the loss weight of each module to the overall goal according to its contribution during the noise injection process. This method introduces a normalized weight function to eliminate dimensional differences and speed up model convergence, reduce the impact of outliers on model training, and improve its stability and model generalization ability.
[0114] First, in order to extend the concept of multi-task to multiple modules, the loss of each module at the current time step t Ratio relative to the previous value Calculated as:
[0115] in, is the module j at time step of loss, is the module j at time step of loss, is the loss ratio of module j at time step t.
[0116] The weights are then updated based on these loss ratios using the normalized weight formula:
[0117] in, It is a module At time step The normalized weight of is the number of modules at time step The average loss ratio of ( ) is the normalization function, is the total number of modules.
[0118] Finally, the updated weights are used to calculate the Total loss: =
[0119] This approach ensures that the weights dynamically adapt to the performance of each module over time, thereby increasing stability and improving overall performance.
[0120] Based on the autonomous driving planning model trained by the above method, please refer to Figure 7 In the application stage, the autonomous driving planning model is used to obtain the autonomous driving planning information for the target vehicle based on the real-time feature information and the prior feature information in the following ways: S141, importing the real-time feature information and the prior feature information into a pre-trained autonomous driving planning model, and fusing the real-time feature information and the prior feature information to obtain a fused feature.
[0121] S142, obtaining perception information, tracking information, map construction information and motion prediction information based on the fusion features.
[0122] S143, combining the perception information, tracking information, map construction information and motion prediction information to generate autonomous driving planning path information.
[0123] From the above, it can be seen that the autonomous driving planning module includes an efficient feature fusion module, a trajectory formation module, a map formation module, a motion prediction module and a planning module.
[0124] The efficient feature fusion module is responsible for fusing the prior feature information provided by the city-level neural radiation field model with the real-time feature information captured by the vehicle's sensors. This fusion process is not only computationally light and adds almost no additional computational overhead, but also cleverly avoids any changes to the original model architecture. This fusion strategy injects rich environmental context information into the autonomous driving system in an extremely efficient way, thereby greatly improving its adaptability and response speed to dynamic and changing driving scenarios while maintaining the simplicity of the system.
[0125] Based on the integration of prior feature information and real-time feature information, the trajectory formation module is further upgraded to use these enhanced features to more accurately identify new obstacles and continuously and stably track the detected targets. Through a set of carefully designed tracking query vectors, the integration of detection and multi-target tracking can be achieved, ensuring the real-time and accuracy of the system.
[0126] The map formation module uses map query vectors to more finely segment various map elements such as lane lines, sidewalks, intersections, etc. These map elements that integrate prior information provide detailed and accurate environmental background information for subsequent path planning and decision-making.
[0127] On this basis, the motion prediction module deeply analyzes the complex interactions between objects and the environment, accurately predicts the future trajectory of each dynamic entity, generates multimodal future action predictions, and provides forward-looking environmental change information for the planning module.
[0128] Finally, after receiving the comprehensive information provided by the above modules, the planning module comprehensively considers the future trajectory prediction of the object and the state of the vehicle, generates the optimal planning path information, and ensures that the autonomous vehicle can drive safely and efficiently in a complex and changing traffic environment. The close collaboration of this series of modules together constitutes the core architecture of the end-to-end autonomous driving system, realizing a fully automated process from environmental perception to decision-making and planning.
[0129] Based on the same inventive concept, please refer to Figure 8, an embodiment of the present invention also provides a functional module diagram of an autonomous driving planning device that combines city-level neural radiation field prior information. This embodiment can divide the functional modules of the autonomous driving planning device according to the above method embodiment. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present invention is schematic, which is only a logical function division. There may be other division methods in actual implementation.
[0130] For example, when each functional module is divided into corresponding functional modules, Figure 8 The autonomous driving planning device shown is only a schematic diagram of the device. The autonomous driving planning device may include a first extraction module, a second extraction module, a conversion module and a processing module. The functions of each functional module of the autonomous driving planning device are described in detail below.
[0131] A first extraction module is used to obtain a real-time image captured by a camera device on a target vehicle, perform feature extraction on the real-time image, and convert the extracted features into real-time feature information from a bird's-eye view; The second extraction module is used to extract the city-level neural radiation field prior information based on the pre-trained city-level neural radiation field model; A conversion module, used to convert the city-level neural radiation field prior information into prior feature information under a bird's-eye view; A processing module is used to import the real-time feature information and the prior feature information into a pre-trained autonomous driving planning model, and fuse the real-time feature information and the prior feature information to obtain a fused feature.
[0132] The autonomous driving planning device provided in this embodiment can be used to execute the autonomous driving planning method under any implementation method in the above embodiments. For matters not detailed in this embodiment, please refer to the corresponding description of the above embodiments, and this embodiment will not be repeated here.
[0133] See also Fig. 9 , is a block diagram of an electronic device provided in an embodiment of the present invention, and the electronic device may be a computer device, a server, etc. in an autonomous driving control platform. The electronic device includes a memory, a processor, and a communication module. The memory, the processor, and the communication module are electrically connected to each other directly or indirectly to achieve data transmission or interaction. For example, these components may be electrically connected to each other via one or more communication buses or signal lines.
[0134] The memory is used to store computer programs or data. The memory can be, but is not limited to, random access memory (RAM), read only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.
[0135] The processor is used to read / write data or programs stored in the memory, and execute the autonomous driving planning method combined with city-level neural radiation field prior information provided by any embodiment of the present invention.
[0136] The communication module is used to establish a communication connection between the electronic device and other communication terminals through the network, and is used to send and receive data through the network.
[0137] It should be understood that Fig. 9 The structure shown is only a schematic diagram of the structure of the electronic device. The electronic device may also include Fig. 9 More or fewer components as shown, or with Fig. 9 Different configurations shown.
[0138] Furthermore, an embodiment of the present invention also provides a computer-readable storage medium, which stores machine-executable instructions. When the machine-executable instructions are executed, the autonomous driving planning method combined with city-level neural radiation field prior information provided in the above embodiment is implemented.
[0139] Specifically, the computer-readable storage medium can be a general storage medium, such as a mobile disk, a hard disk, etc. When the computer program on the computer-readable storage medium is executed, the above-mentioned automatic driving planning method combined with the city-level neural radiation field prior information can be executed. Regarding the process involved when the computer-readable storage medium and its executable instructions are executed, the relevant description in the above-mentioned method embodiment can be referred to, and no further details are given here.
[0140] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.
[0141] In addition, the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0142] Furthermore, the functional modules in the various embodiments of the present invention may be integrated together to form an independent part, or each module may exist independently, or two or more modules may be integrated to form an independent part.
[0143] It should be noted that if the function is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.
[0144] In this document, relational terms such as first and second, etc. are used merely to distinguish one entity or operation from another entity or operation, but do not necessarily require or imply any such actual relationship or order between these entities or operations.
[0145] The above description is only an embodiment of the present invention and is not intended to limit the protection scope of the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An autonomous driving planning method combining city-level neural radiation field prior information, characterized in that: The method comprises: Acquire a real-time image captured by a camera device on a target vehicle, perform feature extraction on the real-time image, and convert the extracted features into real-time feature information from a bird's-eye view; Based on the pre-trained city-level neural radiation field model, the city-level neural radiation field prior information is extracted; Converting the city-level neural radiation field prior information into prior feature information from a bird's-eye view; The real-time feature information and the prior feature information are imported into a pre-trained autonomous driving planning model, and the autonomous driving planning information for the target vehicle is output.
2. The autonomous driving planning method according to claim 1, characterized in that: The step of extracting features from the real-time image and converting the extracted features into real-time feature information from a bird's-eye view includes: Performing a convolution operation on the real-time image to extract multiple different types of features of the real-time image; Perform nonlinear activation function processing on multiple features to introduce nonlinear features, and encode the processed features into feature vectors; The geometric transformation matrix is used to transform the feature vector in the image coordinate system into the real-time feature vector in the bird's-eye view coordinate system.
3. The autonomous driving planning method according to claim 1, characterized in that: The method further comprises the step of pre-training to obtain the city-level neural radiation field model, the step comprising: Acquire multiple historical images collected by camera devices on multiple vehicles in a target city during a historical period; Performing point sampling based on the light direction of the camera device corresponding to each of the historical images to obtain a plurality of sampling points; For each of the sampling points, obtaining 3D position information, a view direction, and a video identifier of the sampling point; Importing the 3D position information, view direction and video identifier corresponding to each of the sampling points into the constructed initial model, processing based on the initial model, and outputting corresponding density information, color information and semantic feature information; Guided by the constructed loss function, based on the density information, color information and semantic feature information output by the initial model, and the density information, color information and semantic feature information corresponding to the area information to which the sampling point belongs, the initial model is iteratively trained multiple times until a trained city-level neural radiation field model is obtained when the preset requirements are met.
4. The autonomous driving planning method according to claim 3, characterized in that: The step of performing point sampling based on the light direction of the camera device corresponding to each of the historical images to obtain a plurality of sampling points comprises: Mapping the position information of the camera device corresponding to each of the historical images into a data point; Performing clustering processing on the multiple data points to divide the multiple data points into multiple sub-areas, and then performing clustering processing on the data points in each of the sub-areas to divide them into multiple sub-fields; For each of the historical images, point sampling is performed on the light direction of the camera device corresponding to the historical image to obtain a plurality of sampling points, and the region information to which each of the sampling points belongs is determined, where the region information is subfield information.
5. The method for autonomous driving planning in combination with city-level neural radiation field prior information according to claim 4, characterized in that: The step of clustering the multiple data points to divide the multiple data points into multiple sub-areas includes: The number of clusters is set based on the size information of the target city, and the centroid of each cluster is randomly generated; Based on the distance between each data point and the centroid of each cluster, each data point is divided into a corresponding cluster, and the centroid of the cluster is updated according to the position information of the data points contained in each cluster. After multiple iterations, until the centroid of each cluster is stable, multiple sub-areas are obtained based on the obtained multiple clusters.
6. The autonomous driving planning method in combination with city-level neural radiation field prior information according to claim 4, characterized in that: The step of pre-training to obtain the city-level neural radiation field model also includes: Identifying pixels of potential moving objects in each of the historical images through a semantic segmentation mask; The pixel value of the identified pixel is set to a specific mask value.
7. The autonomous driving planning method according to claim 1, characterized in that: The step of extracting prior information of the city-level neural radiation field based on the pre-trained city-level neural radiation field model includes: Acquire multiple prior images captured by camera devices on multiple vehicles in the target city; Importing the multiple prior images into a pre-trained city-level neural radiation field model, and for each of the prior images, collecting multiple sampling points in the direction of the projection light of the camera device corresponding to the prior image; Filtering surface points from multiple sampling points based on cumulative transmittance and opacity indicators; Based on multiple surface points corresponding to multiple prior images, city-level neural radiation field prior information is constructed, and the city-level neural radiation field prior information includes road structure, traffic signs, and obstacle locations.
8. The method for autonomous driving planning in combination with city-level neural radiation field prior information according to claim 1, characterized in that: The autonomous driving planning model includes a plurality of modules. The method further includes a step of pre-training to obtain the autonomous driving planning model, the step including: Collecting training samples, where the training samples are the fusion results of images and prior information; Importing the training samples into the autonomous driving planning model, and injecting a noise set into each module of the autonomous driving planning model, and then training with the constructed loss function as a guide; In each round of iteration, the weight values of each module in the autonomous driving planning model in the next round of iteration are calculated, and the total loss function value is calculated according to the weight values of each module in the next round of iteration, until the preset requirements are met, and the trained autonomous driving planning model is obtained.
9. The autonomous driving planning method incorporating city-level neural radiation field prior information according to claim 1, characterized in that: The step of importing the real-time feature information and the prior feature information into a pre-trained autonomous driving planning model and outputting the autonomous driving planning information for the target vehicle includes: Importing the real-time feature information and the prior feature information into a pre-trained autonomous driving planning model, and fusing the real-time feature information and the prior feature information to obtain a fused feature; Obtaining perception information, tracking information, map construction information and motion prediction information based on the fused features; The perception information, tracking information, map construction information and motion prediction information are combined to generate autonomous driving planning path information.
10. An autonomous driving planning device combining city-level neural radiation field prior information, characterized in that: The device comprises: A first extraction module is used to obtain a real-time image captured by a camera device on a target vehicle, perform feature extraction on the real-time image, and convert the extracted features into real-time feature information from a bird's-eye view; The second extraction module is used to extract the city-level neural radiation field prior information based on the pre-trained city-level neural radiation field model; A conversion module, used to convert the city-level neural radiation field prior information into prior feature information under a bird's-eye view; A processing module is used to import the real-time feature information and the prior feature information into a pre-trained autonomous driving planning model, and fuse the real-time feature information and the prior feature information to obtain a fused feature.
Citation Information
Patent Citations
Offline autonomous three-dimensional reconstruction method and system, terminal and storage medium
CN118429522A
Urban airspace map construction method and device based on image acquisition equipment
CN118533161A
Monocular depth estimation method based on surface normal vector and neural radiation field
CN118570273A
Automatic driving 3D operation scene generation method based on aerial view perception
CN119625677A
Programmable logic device and semiconductor device
KR102154633B1
Cited By
4D space-time field construction method and device based on neural field reconstruction
CN121120981A