Method and System for Generating Autonomous Driving Scene Data Based on Implicit Neural Rendering
Through implicit neural rendering technology, high-precision maps are built, object detection and tracking, scene detection and editing, and finally multi-sensor simulation rendering is performed, solving the problem that existing simulators cannot generate highly realistic autonomous driving data, and achieving efficient data generation for diverse scenarios.
Patent Information
- Application Number
- CN202310994135.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-08
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2043-08-08
AI Technical Summary
The existing autonomous driving simulators cannot effectively generate highly realistic multi-sensor data, and the scene library is limited, which cannot meet the data needs of diversified autonomous driving scenarios.
Using an implicit neural rendering method, the scene is reconstructed by constructing offline high-precision maps, performing object detection and tracking, using multi-view implicit surface reconstruction technology, and editing vehicle trajectories in combination with road network information, and finally real-time simulation rendering of multiple types of sensors is performed.
It effectively solves the problem of unrealistic simulation multi-sensor data, is suitable for large-scale dynamic autonomous driving scenarios, and generates highly realistic sensor data through editing and rendering to meet the needs of diversified autonomous driving scenarios.
Smart Images

Figure CN117036607B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of autonomous driving technology, and in particular, to a method and system for generating autonomous driving scenario data based on implicit neural rendering. Background Art
[0002] Currently, autonomous driving is one of the hot topics of concern in the industrial and academic fields. In order to adapt to the complex and changeable driving environments in reality, a large number of complex algorithms are included in autonomous driving systems, and the development and verification of algorithms require a large amount of test data. However, it is difficult to meet the requirements of safety tests only by collecting data from real vehicles. Therefore, simulation has become the only solution to solve the problem of insufficient test mileage data.
[0003] Currently, commonly used simulators such as CARLA can provide relatively complete traffic flow management and configuration schemes for different sensors. However, due to the obvious differences between the simulated world and the real world, they cannot generate highly realistic autonomous driving data well, and the simulated scene library is limited, unable to meet the data requirements of diverse scenarios for autonomous driving. Summary of the Invention
[0004] The embodiments of the present application provide a method and system for generating autonomous driving scenario data based on implicit neural rendering. Through the method of implicit surface reconstruction, the problem of unrealistic multi-sensor data in simulation can be effectively solved, and it is applicable to large-scale dynamic autonomous driving scenarios. At the same time, through the trajectory editing and generation algorithm based on road network information, various long-tail scenarios can be generated by editing a single scenario.
[0005] To solve the above technical problems, in a first aspect, the embodiments of the present application provide a method and system for generating autonomous driving scenario data based on implicit neural rendering, including the following steps: First, based on sensor data, an offline high-precision map is constructed; the sensor data includes lidar point cloud data and pose data; then, based on the offline high-precision map, offline target detection and tracking processing are performed on the sensor data to obtain the target detection and tracking results; next, the method of multi-view implicit surface reconstruction is used to reconstruct the scenario of autonomous driving to obtain the scenario reconstruction result; then, based on the offline high-precision map, the target detection and tracking results, and the road network structure, the trajectory of the vehicle is edited to generate the vehicle trajectory; finally, based on the offline high-precision map, the target detection and tracking results, the scenario reconstruction result, and the vehicle trajectory, real-time simulation rendering is performed using a variety of different types of sensors to obtain a variety of types of highly realistic sensor data.
[0006] In some exemplary embodiments, an offline high-precision map is constructed based on sensor data, including: processing each frame of point cloud in the lidar point cloud data, filtering non-ground points to obtain processed lidar point cloud; based on pose data, stitching all the processed lidar point clouds to obtain full-scene lidar point cloud data; converting the full-scene lidar point cloud data into a top-view image to obtain a road image; wherein, the pixels of the road image correspond one-to-one with the point cloud range; the gray value of a pixel is obtained by calculating the average value of the point cloud intensities within the point cloud range corresponding to the pixel; based on the road image, lanes and road network structures are drawn.
[0007] In some exemplary embodiments, based on the offline high-precision map, the sensor data is subjected to offline object detection and tracking processing to obtain object detection and tracking results, including: based on an object detector, using every five frames of point cloud as the input for one detection to perform object detection to obtain detection results; using raw point features and voxel features to refine the detection results to obtain refined detection results; using a two-stage data association method to obtain multiple sets of tracking trajectories, and associating the multiple sets of tracking trajectories together through a position-aware similarity matching score to obtain a tracking result; predicting the geometric shape, position, and confidence attributes of an object, and further optimizing the refined detection results and the tracking result to obtain object detection and tracking results.
[0008] In some exemplary embodiments, the input of the multi-view implicit surface reconstruction method is camera images and camera internal and external parameter data, lidar data and lidar external parameter data, and 3D detection data of dynamic objects obtained in the offline object detection and tracking part; the multi-view implicit surface reconstruction method is used to reconstruct the autonomous driving scene, and the reconstruction includes background reconstruction and foreground reconstruction; wherein, the foreground represents moving vehicles and pedestrians, and the background represents other parts excluding the foreground; during the background reconstruction process, the space is divided into three parts: near view, far view, and sky, and the near view, far view, and sky are processed separately to reconstruct the background; during the foreground reconstruction process, the foreground is dynamically reconstructed according to the 3D detection box data of moving objects at different times and the foreground asset library.
[0009] In some exemplary embodiments, a multi-vehicle decision-making and planning method is used to edit the trajectory of a vehicle to generate a vehicle trajectory; the input of the multi-vehicle decision-making and planning method includes road network topology structure, route information, vehicle state, and trajectory prediction for uncontrolled vehicles; the output of the multi-vehicle decision-making and planning method is the original ego-vehicle and other-vehicle trajectories and the edited other-vehicle trajectories; the multi-vehicle decision-making and planning method includes: simultaneously generating rough and interpretable discrete decision sequences for all autonomous vehicles to solve the interaction behavior between vehicles in the scene; based on the discrete decision sequences, generating continuous kinematically feasible trajectories for each controlled autonomous vehicle.
[0010] In a second aspect, an embodiment of the present application further provides an autonomous driving scenario data generation system based on implicit neural rendering, including: a high-precision map construction module, an object detection and tracking module, an implicit three-dimensional reconstruction module, a trajectory editing and generation module, and a real-time rendering module that are connected in sequence; the high-precision map construction module is used to construct an offline high-precision map according to sensor data; the sensor data includes lidar point cloud data and pose data; the object detection and tracking module is used to perform offline object detection and tracking processing on the sensor data according to the offline high-precision map to obtain an object detection and tracking result; the implicit three-dimensional reconstruction module is used to reconstruct the autonomous driving scenario by using the method of multi-view implicit surface reconstruction to obtain a scenario reconstruction result; the trajectory editing and generation module is used to edit the vehicle trajectory according to the offline high-precision map, the object detection and tracking result, and the road network structure to generate a vehicle trajectory; the real-time rendering module is used to perform real-time simulation rendering by using a variety of different types of sensors according to the offline high-precision map, the object detection and tracking result, the scenario reconstruction result, and the vehicle trajectory to obtain a variety of highly realistic sensor data.
[0011] In some exemplary embodiments, the high-precision map construction module includes a data preprocessing module, an image conversion module, and a lane and road network structure drawing module; the data preprocessing module is used to process each frame of point cloud in the lidar point cloud data, filter non-ground points to obtain processed lidar point cloud; and based on the pose data, splice all the processed lidar point cloud to obtain full-scene lidar point cloud data; the image conversion module is used to convert the full-scene lidar point cloud data into a top view image to obtain a road image; wherein, the pixels of the road image correspond one-to-one with the point cloud range; the gray value of the pixel is obtained by calculating the average value of the point cloud intensities within the point cloud range corresponding to the pixel; the lane and road network structure drawing module is used to draw the lane and road network structure according to the road image.
[0012] In some exemplary embodiments, the object detection and tracking module includes an object detector, a point intensity perception module, a multi-object tracking module, and a prediction module; the object detector is used to use every five frames of point cloud as the input of one detection for object detection to obtain a detection result; the point intensity perception module is used to refine the detection result by using the original point feature and voxel feature to obtain a refined detection result; the multi-object tracking module is used to obtain multiple sets of tracking trajectories according to the two-stage data association method, and associate the multiple sets of tracking trajectories together through the position-aware similarity matching score to obtain a tracking result; the prediction module is used to predict the geometric shape, position, and confidence attributes of the object, and further optimize the refined detection result and the tracking result to obtain an object detection and tracking result.
[0013] In some exemplary embodiments, the implicit three-dimensional reconstruction module includes a background reconstruction module and a foreground reconstruction module; the background reconstruction module is configured to divide the space into a near view, a far view, and the sky, and perform background reconstruction processing on the near view, the far view, and the sky respectively; the foreground reconstruction module is configured to perform dynamic reconstruction on the foreground according to the 3D detection box data of the moving target at different times and the foreground asset library.
[0014] In some exemplary embodiments, the trajectory editing and generation module includes a decision-making module and a trajectory planning module; wherein, the decision-making module is configured to generate a rough and interpretable discrete decision sequence for all autonomous driving vehicles simultaneously to solve the interaction behavior between the scenario vehicles; the trajectory planning module is configured to receive the discrete decision sequence and generate a continuous kinematically feasible trajectory for each controlled autonomous driving vehicle.
[0015] The technical solution provided by the embodiments of the present application has at least the following advantages:
[0016] The embodiments of the present application provide a method and system for generating autonomous driving scenario data based on implicit neural rendering. The method includes the following steps: First, based on sensor data, an offline high-precision map is constructed; the sensor data includes lidar point cloud data and pose data; then, based on the offline high-precision map, offline target detection and tracking processing are performed on the sensor data to obtain a target detection and tracking result; next, a multi-view implicit surface reconstruction method is used to reconstruct the autonomous driving scenario to obtain a scenario reconstruction result; then, based on the offline high-precision map, the target detection and tracking result, and the road network structure, the trajectory of the vehicle is edited to generate a vehicle trajectory; finally, based on the offline high-precision map, the target detection and tracking result, the scenario reconstruction result, and the vehicle trajectory, real-time simulation rendering is performed using a variety of different types of sensors to obtain a variety of highly realistic sensor data.
[0017] This application provides a method and system for generating autonomous driving scenario data based on implicit neural rendering. By presenting a complete system for generating autonomous driving data, which includes five modules: offline object detection and tracking, high-precision map generation, implicit three-dimensional reconstruction, trajectory generation, and new sensor simulation rendering. Trajectories are generated in combination with traffic flow and multi-scene rendering is performed in combination with the reconstruction results. To generate highly realistic sensor data, on the one hand, combined with road network information, the trajectories of vehicles are edited to generate more vehicle-dynamics-compliant and more realistic vehicle trajectories; on the other hand, through implicit surface reconstruction technology and combined with a foreground asset library, a highly realistic scene reconstruction of the autonomous driving scenario is carried out; finally, the generated trajectories are used to edit and render the scene to generate various types of highly realistic sensor data. In addition, when rendering sensor data, this application supports the simulation of 16 different models of lidar, cameras, and millimeter-wave radars, and also supports 3D Radar and 4D Radar, meeting the requirements of various sensor configurations in autonomous driving scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] One or more embodiments are illustrated by way of example in the accompanying drawings, which do not constitute a limitation on the embodiments unless otherwise stated. The figures in the drawings do not constitute a scale limitation.
[0019] Figure 1 It is a schematic flow chart of a method for generating autonomous driving scenario data based on implicit neural rendering provided by an embodiment of this application;
[0020] Figure 2 It is a schematic structural diagram of a data generation system for implicit surface reconstruction based on multi-views provided by an embodiment of this application;
[0021] Figure 3 It is a schematic simulation flow chart of a data generation system for implicit surface reconstruction based on multi-views provided by an embodiment of this application;
[0022] Figure 4 It is a high-precision map with road network information provided by an embodiment of this application;
[0023] Figure 5 It is a schematic diagram of the implicit surface reconstruction result provided by an embodiment of this application;
[0024] Figure 6 It is an HMI display diagram of the simulation rendering of the VLP16 lidar provided by an embodiment of this application;
[0025] Figure 7 It is an HMI display diagram of the simulation rendering of the Pandar_qt lidar provided by an embodiment of this application;
[0026] Figure 8Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0027] As can be seen from the background art, although existing simulators can provide relatively complete traffic flow management and configuration solutions for different sensors, due to the obvious differences between the simulation world and the real world, they cannot generate highly realistic autonomous driving data well, and the simulation scenario library is limited and cannot meet the data requirements of diverse scenarios for autonomous driving.
[0028] To ensure the safety of autonomous driving, it is necessary to test a large amount of autonomous driving data. However, the real world is too complex, especially with the long-tail effect. These boundary scenarios are crucial for safe driving, but they are diverse and difficult to encounter, so testing these scenarios in the real world is very expensive and dangerous. Therefore, generating various autonomous driving data through simulation is crucial for solving the long-tail problem and improving the safety of autonomous driving.
[0029] The Neural Radiance Field (NeRF) reconstructs the scene and generates data from a new perspective in an implicit representation manner. Compared with simulators based on physical models, it can synthesize high-fidelity sensor data, so it has received more attention in the field of autonomous driving sensor simulation in recent years. UniSim is a representative algorithm for current autonomous driving sensor simulation using three-dimensional implicit reconstruction. UniSim reconstructs the static background and dynamic vehicles in the scene by constructing a neural network feature grid and synthesizes them together to simulate LiDAR and camera data from a new perspective. UniSim also supports editing of autonomous driving scenarios, and completes the insertion, removal, and manipulation of vehicles by replacing the available space with dynamic vehicles. To better process the synthesized views, UniSim combines the dynamic vehicles with learnable priors and uses a convolutional network to complete the reconstruction of invisible regions.
[0030] However, current novel view synthesis methods implicitly represent the scene as a Neural Radiance Field (NeRF) and use neural networks to perform volume rendering. Although these methods can represent complex geometries and appearances and have achieved photorealistic rendering, they are only applicable to small static scenes and cannot handle large-scale autonomous driving scenes with a large number of moving vehicles. Existing autonomous driving simulation platforms such as CARLA and AirSim can generate road network structures, but due to the high labor cost, the created simulation scenes are limited and it is difficult to cover all the scenes required for autonomous driving testing. Moreover, the generated sensor data is not realistic enough. In addition, data-driven autonomous driving sensor simulation combines real data and computer vision techniques to construct new simulated sensor data. However, these methods have problems such as being unable to render novel view sensor data or the rendered data being unrealistic.
[0031] To solve the above technical problems, the present application provides an autonomous driving scene data generation method based on implicit neural rendering. The method includes the following steps: First, based on sensor data, an offline high-precision map is constructed; the sensor data includes lidar point cloud data and pose data; then, based on the offline high-precision map, offline object detection and tracking processing are performed on the sensor data to obtain object detection and tracking results; next, a multi-view implicit surface reconstruction method is used to reconstruct the autonomous driving scene to obtain a scene reconstruction result; then, based on the offline high-precision map, object detection and tracking results, and road network structure, the vehicle trajectory is edited to generate a vehicle trajectory; finally, based on the offline high-precision map, object detection and tracking results, scene reconstruction result, and vehicle trajectory, real-time simulation rendering is performed using multiple different types of sensors to obtain highly realistic sensor data of multiple types. The method proposed in the present application can effectively solve the problem of unrealistic multi-sensor data in simulation through the implicit surface reconstruction method, and is applicable to large-scale dynamic autonomous driving scenes. At the same time, through the trajectory editing and generation algorithm based on road network information, various long-tail scenes can be edited and generated for a single scene.
[0032] The following will elaborate on the embodiments of the present application in conjunction with the accompanying drawings. However, those of ordinary skill in the art can understand that in the embodiments of the present application, many technical details are proposed to help readers better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented.
[0033] See Figure 1 , the embodiments of the present application provide an autonomous driving scene data generation method and system based on implicit neural rendering, including the following steps:
[0034] Step S1: Construct an offline high-precision map based on sensor data, where the sensor data includes lidar point cloud data and pose data.
[0035] Step S2: Based on the offline high-precision map, perform offline object detection and tracking on the sensor data to obtain the object detection and tracking results.
[0036] Step S3: Reconstruct the autonomous driving scenario using the method of multi-view implicit surface reconstruction to obtain the scene reconstruction results.
[0037] Step S4: Edit the vehicle trajectory based on the offline high-precision map, object detection and tracking results, and road network structure to generate the vehicle trajectory.
[0038] Step S5: Based on the offline high-precision map, object detection and tracking results, scene reconstruction results, and vehicle trajectory, perform real-time simulation rendering using multiple different types of sensors to obtain highly realistic sensor data of multiple types.
[0039] In some embodiments, constructing the offline high-precision map based on the sensor data in Step S1 includes the following steps:
[0040] Step S101: Process each frame of the lidar point cloud data, filter out non-ground points, and obtain the processed lidar point cloud.
[0041] Step S102: Based on the pose data, splice all the processed lidar point clouds to obtain the full-scene lidar point cloud data.
[0042] Step S103: Convert the full-scene lidar point cloud data into a top-view image to obtain the road image, where the pixels of the road image correspond one-to-one with the point cloud range, and the gray value of the pixel is obtained by calculating the average value of the point cloud intensities within the point cloud range corresponding to the pixel.
[0043] Step S104: Draw lanes and road network structure based on the road image.
[0044] In some embodiments, performing offline object detection and tracking on the sensor data based on the offline high-precision map in Step S2 to obtain the object detection and tracking results includes the following steps:
[0045] Step S201: Based on the object detector, use every five frames of point cloud as the input for each detection to perform object detection and obtain the detection results.
[0046] Step S202: Refine the detection results using the original point features and voxel features to obtain the refined detection results.
[0047] Step S203: Obtain multiple groups of tracking trajectories using a two-stage data association method, and associate the multiple groups of tracking trajectories together through the position perception similarity matching score to obtain the tracking result.
[0048] Step S204: Predict the geometric shape, position, and confidence attributes of the object, and further optimize the refined detection result and the tracking result to obtain the target detection and tracking result.
[0049] In some embodiments, the input of the multi-view implicit surface reconstruction method in step S3 is camera images and camera internal and external parameter data, lidar data and lidar external parameter data, and 3D detection data of dynamic objects obtained in the offline object detection and tracking part; the multi-view implicit surface reconstruction method is used to reconstruct the autonomous driving scene, and the reconstruction includes background reconstruction and foreground reconstruction; wherein, the foreground represents moving vehicles and pedestrians, and the background represents other parts excluding the foreground; during the background reconstruction process, the space is divided into three parts: near view, far view, and sky, and the near view, far view, and sky are processed separately to reconstruct the background; during the foreground reconstruction process, the foreground is dynamically reconstructed according to the 3D detection box data of moving objects at different times and the foreground asset library.
[0050] In some embodiments, in step S4, a multi-vehicle decision-making and planning method is used to edit the trajectory of the vehicle to generate a vehicle trajectory; the input of the multi-vehicle decision-making and planning method includes road network topology structure, route information, vehicle state, and trajectory prediction for uncontrolled vehicles; the output of the multi-vehicle decision-making and planning method is the original ego-vehicle and other-vehicle trajectories and the edited other-vehicle trajectories. Among them, the multi-vehicle decision-making and planning method includes the following steps: simultaneously generate rough and interpretable discrete decision sequences for all autonomous vehicles to solve the interaction behavior between vehicles in the scene; based on the discrete decision sequences, generate continuous kinematically feasible trajectories for each controlled autonomous vehicle.
[0051] See Figure 2, the embodiment of the present application also provides an autonomous driving scenario data generation system based on implicit neural rendering, including: a high-precision map construction module 101, a target detection and tracking module 102, an implicit three-dimensional reconstruction module 103, a trajectory editing and generation module 104, and a real-time rendering module 105 connected in sequence; the high-precision map construction module 101 is used to construct an offline high-precision map according to sensor data; the sensor data includes lidar point cloud data and pose data; the target detection and tracking module 102 is used to perform offline target detection and tracking processing on the sensor data according to the offline high-precision map to obtain target detection and tracking results; the implicit three-dimensional reconstruction module 103 is used to reconstruct the autonomous driving scenario by using the method of multi-view implicit surface reconstruction to obtain a scene reconstruction result; the trajectory editing and generation module 104 is used to edit the vehicle trajectory according to the offline high-precision map, the target detection and tracking results, and the road network structure to generate a vehicle trajectory; the real-time rendering module 105 is used to perform real-time simulation rendering by using a variety of different types of sensors according to the offline high-precision map, the target detection and tracking results, the scene reconstruction result, and the vehicle trajectory to obtain a variety of highly realistic sensor data.
[0052] The present application proposes a brand-new editable multi-sensor simulation system for autonomous driving scenarios based on implicit surface reconstruction. The system flow schematic diagram is as Figure 3 shown. The system covers complete modules such as high-precision map generation, target detection and tracking, implicit three-dimensional reconstruction, trajectory editing and generation, and multi-sensor real-time rendering simulation. In order to generate highly accurate sensor data, on the one hand, combined with road network information, the vehicle trajectory is edited to generate a more vehicle-dynamics-compliant and more realistic vehicle trajectory; on the other hand, through implicit surface reconstruction technology and combined with a foreground asset library, a highly accurate scene reconstruction of the autonomous driving scenario is performed; finally, the generated trajectory is combined to edit and render the scene to generate various highly accurate sensor data.
[0053] In some embodiments, the high-precision map construction module 101 includes a data preprocessing module, an image conversion module, and a lane and road network structure drawing module; the data preprocessing module is used to process each frame of point cloud in the lidar point cloud data, filter non-ground points to obtain processed lidar point cloud; and based on the pose data, splice all the processed lidar point cloud to obtain full-scene radar point cloud data; the image conversion module is used to convert the full-scene radar point cloud data into a top-view image to obtain a road image; wherein, the pixels of the road image correspond one-to-one with the point cloud range; the gray value of the pixel is obtained by calculating the average value of the point cloud intensities within the point cloud range corresponding to the pixel; the lane and road network structure drawing module is used to draw the lane and road network structure according to the road image.
[0054] Specifically, in the process of constructing a high-precision map, based on the given lidar point cloud data and pose data, first, each frame of point cloud is processed to filter out non-ground points, and then the pose data is used to stitch all the processed lidar point clouds to obtain a complete full-scene lidar point cloud file. Next, the lidar point cloud data is converted into a top-down view image, where one pixel of the image corresponds to the point cloud within a fixed range, and the gray value of the pixel is calculated from the average value of the point cloud intensities within the corresponding range. After obtaining the road image, the lane and road network structure are drawn so that the trajectory generation module can edit the original trajectory and generate a new trajectory based on the road network information. Figure 4 The generated high-precision map with road network information is shown, and each lane will have a corresponding number, and the information of adjacent lanes will also be recorded.
[0055] In some embodiments, the object detection and tracking module 102 includes an object detector, a point intensity perception module, a multi-object tracking module, and a prediction module; the object detector is used to take every five frames of point cloud as the input of one detection, perform object detection, and obtain a detection result; the point intensity perception module is used to refine the detection result by using the original point features and voxel features to obtain a refined detection result; the multi-object tracking module is used to obtain multiple sets of tracking trajectories according to the two-stage data association method, and associate the multiple sets of tracking trajectories together through the position-aware similarity matching score to obtain a tracking result; the prediction module is used to predict the geometric shape, position, and confidence attributes of the object, and further optimize the refined detection result and the tracking result to obtain an object detection and tracking result.
[0056] Specifically, for the object detection part, the CenterPoint model is used as the basic object detector. First, every five frames of point cloud are taken as the input of one detection, and a point intensity perception module is designed to use the original point features and voxel features to achieve accurate result refinement.
[0057] Next is multi-object tracking. This module uses a two-stage data association strategy to reduce the possibility of incorrect matching. The boxes detected according to the confidence level are divided into two different groups. The pre-existing object trajectories are initially only associated with the high group for data association. Subsequently, the successfully associated boxes are used to update the existing trajectories. The unupdated trajectories are further associated with the low group, and the unassociated boxes are discarded. In addition, the life cycle of the object is allowed to persist until the sequence ends, and then the redundant boxes that have not been updated will be deleted. Reverse the time order and perform the above tracking process again to generate another set of tracking results. Then these trajectories are associated together through the position-aware similarity matching score.
[0058] Finally, the geometric shape, position, and confidence attributes of the object are predicted by three different modules respectively to further optimize the accuracy of object tracking and detection. The detection box and tracking number of the object in each frame obtained are combined with the high-precision map to obtain the lane information where the object is located in each frame, which facilitates the vehicle trajectory control and generalization generation in the trajectory generation part.
[0059] In some embodiments, the implicit 3D reconstruction module 103 includes a background reconstruction module and a foreground reconstruction module; the background reconstruction module is used to divide the space into the near view, far view, and sky, and perform background reconstruction processing on the near view, far view, and sky respectively; the foreground reconstruction module is used to perform dynamic reconstruction on the foreground according to the 3D detection box data of moving objects at different times and the foreground asset library.
[0060] Specifically, the implicit 3D reconstruction part uses the method of multi-view implicit surface reconstruction to reconstruct the scene of autonomous driving. The input of this algorithm is the camera image and its internal and external parameter data, lidar data and its external parameters, and the 3D detection data of dynamic objects obtained in the offline object detection and tracking part. The reconstruction is divided into background reconstruction and foreground reconstruction. The foreground represents moving vehicles and pedestrians, and the other parts are the background. During the background reconstruction process, to address the challenges brought by the unbounded space, the space is divided into three parts: the near view, far view, and sky, and the cuboid NeuS model, hypercuboid NeRF++ model, and directional multi-layer perceptron model (MLP) are used to process the above three parts respectively. For foreground reconstruction, it is dynamically reconstructed according to the 3D detection box data of moving objects at different times and the foreground asset library. Figure 5 The implicit surface reconstruction result is shown.
[0061] In some embodiments, the trajectory editing and generation module 104 includes a decision-making module and a trajectory planning module; among them, the decision-making module is used to generate a rough and interpretable discrete decision sequence for all autonomous driving vehicles simultaneously to solve the interaction behavior between vehicles in the scene; the trajectory planning module is used to receive the discrete decision sequence and generate a continuous kinematically feasible trajectory for each controlled autonomous driving vehicle.
[0062] Specifically, the trajectory generation uses a two-stage multi-vehicle decision-making and planning algorithm that includes a decision-making module and a trajectory planning module. The input of the algorithm includes the road network topology structure, route information, vehicle state, and trajectory prediction for uncontrolled vehicles. The output is the original ego-vehicle and other-vehicle trajectories and the edited other-vehicle trajectories.
[0063] The first stage is a decision-making module for addressing the interaction behaviors among scenario vehicles. This module generates rough and interpretable decision sequences for all autonomous vehicles simultaneously. The module uses an improved version of Monte-Carlo Tree Search (MCTS), taking into account all vehicles, including human-driven vehicles, in the traffic scenario during the decision-making process, and there is no priority among vehicles during the decision-making process.
[0064] The second stage is a trajectory planning module that receives the discrete decision sequences output by the first stage and generates continuous kinematically feasible trajectories for each controlled autonomous vehicle. The planning module adopts a distributed parallel architecture, that is, a planner runs independently for each vehicle, which is closer to human driving in the real world.
[0065] For the real-time rendering module 105, the inputs to this module are the perception results of the scene obtained above, the generated trajectories of the ego vehicle and other vehicles, and the reconstructed background model and foreground asset library. Different sensors can be selected for real-time simulation rendering according to different simulation requirements. The supported sensors include 16 different models of lidar, ordinary cameras, telephoto cameras, fisheye cameras, 3D Radar, and 4D Radar.
[0066] The front-end display diagram of the final data generation system is as Figure 6 and Figure 7 shown. The right image represents the real camera pictures and lidar point cloud perception results, and the left side represents the multi-sensor data rendered by simulation. The second small box on the upper side represents the rendered wide-angle camera data, and the point clouds of different colors on the left represent the data of different lidars rendered. Different sensor types and models can be selected through the check boxes on the left to obtain their corresponding rendered data. Figure 6 and Figure 7 correspond to the rendered data of the VLP16 and Pandar_qt lidars respectively. Among them, Figure 6 and Figure 7 the right subgraphs are the data collected in reality, and the left subgraphs are the new data generated according to the trajectories.
[0067] In summary, on the one hand, the present application provides a complete autonomous driving data generation system, including five modules: offline object detection and tracking, high-precision map generation, implicit three-dimensional reconstruction, trajectory generation, and new sensor simulation rendering. On the other hand, the present application generates trajectories in combination with traffic flow and performs multi-scene rendering in combination with the reconstruction results. In addition, the present application supports the simulation of 16 different models of lidar, cameras, and millimeter-wave radars.
[0068] Compared with the prior art, the advantages of the present invention are:
[0069] (1) The present invention is a complete data generation system. Only by providing the data collected by cameras and lidars and their corresponding internal and external parameters, the reconstruction of the foreground and background in the autonomous driving scenario can be completed through the implicit surface reconstruction algorithm. It has a high degree of automation, strong scalability, and can be directly deployed on a single machine.
[0070] (2) When generating scenario data, the present invention does not simply delete or insert moving vehicles, but edits the moving vehicles based on road network information. It can not only edit all vehicles simultaneously, but also make the vehicle trajectories edited more conform to the road network constraints and vehicle dynamics laws.
[0071] (3) When rendering sensor data, the present invention can render lidar and camera data of 16 different models, and supports 3D Radar and 4D Radar, meeting the requirements of various sensor configurations in the autonomous driving scenario.
[0072] By conducting experiments and simulations on the method and system for generating autonomous driving scenario data based on implicit neural rendering provided by the present invention to verify its effects, it is proved through experiments that the method and system of this application are feasible and the results meet the expectations. The algorithm in the present invention can generate highly realistic various sensor simulation data in large-scale dynamic autonomous driving scenarios.
[0073] Refer to Figure 8 , another embodiment of this application provides an electronic device, including: at least one processor 110; and a memory 111 communicatively connected to the at least one processor; wherein, the memory 111 stores instructions executable by the at least one processor 110, and the instructions are executed by the at least one processor 110 to enable the at least one processor 110 to execute any of the above method embodiments.
[0074] Wherein, the memory 111 and the processor 110 are connected in a bus manner. The bus can include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors 110 and the memory 111 together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, so they will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be one component or multiple components, such as multiple receivers and transmitters, providing units for communicating with various other devices on the transmission medium. The data processed by the processor 110 is transmitted on the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor 110.
[0075] The processor 110 is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. The memory 111 can be used to store the data used by the processor 110 when executing operations.
[0076] Another embodiment of the present application relates to a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the method embodiments described above are implemented.
[0077] That is, those skilled in the art can understand that all or part of the steps in implementing the above-described method embodiments can be completed by instructing relevant hardware through a program. The program is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the above-described methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
[0078] According to the above technical solutions, the embodiments of the present application provide a method and system for generating autonomous driving scenario data based on implicit neural rendering. The method includes the following steps: First, an offline high-precision map is constructed based on sensor data; the sensor data includes lidar point cloud data and pose data; then, based on the offline high-precision map, the sensor data is subjected to offline object detection and tracking processing to obtain an object detection and tracking result; next, a multi-view implicit surface reconstruction method is used to reconstruct the autonomous driving scenario to obtain a scenario reconstruction result; then, based on the offline high-precision map, the object detection and tracking result, and the road network structure, the vehicle trajectory is edited to generate a vehicle trajectory; finally, based on the offline high-precision map, the object detection and tracking result, the scenario reconstruction result, and the vehicle trajectory, real-time simulation rendering is performed using various different types of sensors to obtain various types of highly realistic sensor data.
[0079] The present application provides a method and system for generating autonomous driving scenario data based on implicit neural rendering. By proposing a complete system for generating autonomous driving data, which includes five modules: offline object detection and tracking, high-precision map generation, implicit three-dimensional reconstruction, trajectory generation, and new sensor simulation rendering. Trajectories are generated in combination with traffic flow and multi-scene rendering is performed in combination with the reconstruction results. In order to generate highly realistic sensor data, on the one hand, the road network information is combined to edit the vehicle trajectories to generate more vehicle dynamics-compliant and more realistic vehicle trajectories; on the other hand, highly realistic scene reconstruction of the autonomous driving scenario is carried out through implicit surface reconstruction technology and in combination with a foreground asset library; finally, the generated trajectories are used to edit and render the scene to generate various types of highly realistic sensor data. In addition, when rendering sensor data, the present application supports the simulation of 16 different models of lidar, cameras, and millimeter-wave radars, and also supports 3D Radar and 4D Radar, meeting the requirements of various sensor configurations in autonomous driving scenarios.
[0080] Those of ordinary skill in the art can understand that the above-described embodiments are specific examples for implementing the present application. In actual applications, various changes can be made in form and details without departing from the spirit and scope of the present application. Any person skilled in the art can make their own changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application should be subject to the scope defined by the claims.
Claims
1. A method for generating autonomous driving scenario data based on implicit neural rendering, characterized in that, Including: Construct an offline high-precision map based on sensor data; the sensor data includes lidar point cloud data and pose data; Based on the offline high-precision map, perform offline object detection and tracking processing on the sensor data to obtain an object detection and tracking result; Reconstruct the autonomous driving scene using the multi-view implicit surface reconstruction method to obtain a scene reconstruction result; Based on the offline high-precision map, the object detection and tracking result, and the road network structure, edit the vehicle's trajectory to generate a vehicle trajectory; Based on the offline high-precision map, the object detection and tracking result, the scene reconstruction result, and the vehicle trajectory, use multiple different types of sensors for real-time simulation rendering to obtain multiple types of highly realistic sensor data; Use the multi-vehicle decision-making and planning method to edit the vehicle's trajectory to generate a vehicle trajectory; The input of the multi-vehicle decision-making and planning method includes the road network topology structure, route information, vehicle state, and trajectory prediction for uncontrolled vehicles; The output of the multi-vehicle decision-making and planning method is the original ego-vehicle and other-vehicle trajectories and the edited other-vehicle trajectories; The multi-vehicle decision-making and planning method includes: Simultaneously generate rough and interpretable discrete decision sequences for all autonomous vehicles to solve the interaction behavior between vehicles in the scene; Based on the discrete decision sequence, generate continuous kinematically feasible trajectories for each controlled autonomous vehicle.
2. The method for generating autonomous driving scenario data based on implicit neural rendering according to claim 1, characterized in that, Construct an offline high-precision map based on sensor data, including: Process each frame of point cloud in the lidar point cloud data, filter non-ground points to obtain processed lidar point cloud; Based on the pose data, splice all processed lidar point clouds to obtain full-scene lidar point cloud data; Convert the full-scene lidar point cloud data into a top-down view image to obtain a road image; wherein, the pixels of the road image correspond one-to-one with the point cloud range; the gray value of the pixel is obtained by calculating the average value of the point cloud intensities within the point cloud range corresponding to the pixel; Based on the road image, draw lanes and the road network structure.
3. The method for generating autonomous driving scenario data based on implicit neural rendering according to claim 1, characterized in that, Based on the offline high-precision map, perform offline object detection and tracking processing on the sensor data to obtain an object detection and tracking result, including: Based on an object detector, use every five frames of point cloud as the input for one detection to perform object detection and obtain a detection result; Refine the detection result using the original point features and voxel features to obtain a refined detection result; Use a two-stage data association method to obtain multiple sets of tracking trajectories, and associate the multiple sets of tracking trajectories together through the position-aware similarity matching score to obtain a tracking result; Predict the geometric shape, position, and confidence attributes of the object, and further optimize the refined detection result and the tracking result to obtain an object detection and tracking result.
4. The method for generating autonomous driving scenario data based on implicit neural rendering according to claim 1, characterized in that, The input of the multi-view implicit surface reconstruction method is camera images, camera internal and external parameter data, lidar data, lidar external parameter data, and 3D detection data of dynamic targets obtained in the offline object detection and tracking section; the multi-view implicit surface reconstruction method is used to reconstruct the autonomous driving scenario, and the reconstruction includes background reconstruction and foreground reconstruction; wherein, the foreground represents moving vehicles and pedestrians, and the background represents other parts excluding the foreground. During the background reconstruction process, the space is divided into three parts: near view, far view, and sky, and the near view, far view, and sky are processed separately to reconstruct the background. During the foreground reconstruction process, the foreground is dynamically reconstructed according to the 3D detection box data of moving targets at different times and the foreground asset library.
5. A system for generating autonomous driving scenario data based on implicit neural rendering, characterized in that, It includes: A high-precision map construction module, an object detection and tracking module, an implicit three-dimensional reconstruction module, a trajectory editing and generation module, and a real-time rendering module connected in sequence. The high-precision map construction module is used to construct an offline high-precision map according to sensor data; the sensor data includes lidar point cloud data and pose data. The object detection and tracking module is used to perform offline object detection and tracking processing on the sensor data according to the offline high-precision map to obtain object detection and tracking results. The implicit three-dimensional reconstruction module is used to reconstruct the autonomous driving scenario by using the multi-view implicit surface reconstruction method to obtain a scene reconstruction result. The trajectory editing and generation module is used to edit the vehicle trajectory according to the offline high-precision map, the object detection and tracking results, and the road network structure to generate a vehicle trajectory. The vehicle trajectory is edited by using a multi-vehicle decision-making and planning method to generate a vehicle trajectory. The input of the multi-vehicle decision-making and planning method includes road network topology structure, route information, vehicle state, and trajectory prediction for uncontrolled vehicles. The output of the multi-vehicle decision-making and planning method is the original ego-vehicle and other-vehicle trajectories and the edited other-vehicle trajectories. The multi-vehicle decision-making and planning method includes: generating rough and interpretable discrete decision sequences for all autonomous vehicles simultaneously to solve the interaction behavior between vehicles in the scenario; based on the discrete decision sequences, generating continuous kinematically feasible trajectories for each controlled autonomous vehicle. The real-time rendering module is used to perform real-time simulation rendering by using multiple different types of sensors according to the offline high-precision map, the object detection and tracking results, the scene reconstruction result, and the vehicle trajectory to obtain highly realistic sensor data of multiple types.
6. The system for generating autonomous driving scenario data based on implicit neural rendering according to claim 5, wherein, The high-precision map construction module includes a data preprocessing module, an image conversion module, and a lane and road network structure drawing module. The data preprocessing module is used to process each frame of point cloud in the lidar point cloud data, filter non-ground points, and obtain processed lidar point cloud. And based on the pose data, all processed lidar point clouds are stitched together to obtain full-scene lidar point cloud data. The image conversion module is used to convert the full-scene radar point cloud data into a top-view image to obtain a road image; wherein, the pixels of the road image correspond one-to-one with the point cloud range; the gray value of the pixel is obtained by calculating the average value of the point cloud intensities within the point cloud range corresponding to the pixel; The lane and road network structure drawing module is used to draw lanes and road network structures according to the road image.
7. The system for generating autonomous driving scenario data based on implicit neural rendering according to claim 5, wherein, The target detection and tracking module includes an object detector, a point intensity perception module, a multi-object tracking module, and a prediction module; The object detector is used to take every five frames of point cloud as the input of one detection, perform object detection, and obtain a detection result; The point intensity perception module is used to refine the detection result by using the original point features and voxel features to obtain a refined detection result; The multi-object tracking module is used to obtain multiple sets of tracking trajectories according to the two-stage data association method, and associate the multiple sets of tracking trajectories together through the position-aware similarity matching score to obtain a tracking result; The prediction module is used to predict the geometric shape, position, and confidence attributes of the object, and further optimize the refined detection result and the tracking result to obtain a target detection and tracking result.
8. The system for generating autonomous driving scenario data based on implicit neural rendering according to claim 5, wherein, The implicit three-dimensional reconstruction module includes a background reconstruction module and a foreground reconstruction module; The background reconstruction module is used to divide the space into the near view, the far view, and the sky, and perform background reconstruction processing on the near view, the far view, and the sky respectively; The foreground reconstruction module is used to perform dynamic reconstruction on the foreground according to the 3D detection box data of moving targets at different times and the foreground asset library.
Citation Information
Patent Citations
Outdoor unbounded scene three-dimensional reconstruction method and system based on neural radiation field
CN116051740A
Slam-based mobile robot mine scene reconstruction method and system
WO2022257801A1