3D Occupancy Grid Based Driving Scenario Simulation Method, System, Device and Medium

Through the driving scene simulation method based on the three-dimensional occupancy grid, a three-dimensional occupancy grid descriptor is constructed, which solves the problem that traditional representation methods are difficult to describe irregular objects and a complete three-dimensional environment, and achieves accurate description of the environment and performance improvement.

CN116452766BActive Publication Date: 2025-07-18SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310450794.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-24
Publication Date
2025-07-18
Estimated Expiration
2043-04-24

AI Technical Summary

Technical Problem

Traditional representation methods such as 3D frames, point cloud segmentation, BEV segmentation, and LiDAR segmentation are difficult to accurately describe the surrounding environment, which affects the performance of the driving system, especially when describing irregular objects and a complete three-dimensional environment.

Method used

A three-dimensional occupancy grid driving scene simulation method is provided. By extracting features from the circumferential image, a BEV encoder is used to obtain features, and decoding in voxel space using voxel decoder is used to construct a three-dimensional occupancy grid descriptor, supporting downstream tasks such as 3D detection, BEV segmentation and motion planning.

Benefits of technology

The precise description of the three-dimensional environment is achieved, which can be better used for downstream tasks, and improves the performance of the driving system, especially when describing irregular objects and moving foreground objects, reducing the collision rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116452766B_ABST
    Figure CN116452766B_ABST
Patent Text Reader

Abstract

The present application relates to the field of computer vision technology, and in particular to a method, system, device and medium for simulating a driving scene based on a three-dimensional occupancy grid, the method comprising the following steps: first, constructing a three-dimensional space occupancy grid descriptor based on visual input; then, implementing downstream tasks based on the occupancy grid descriptor; wherein, constructing a three-dimensional space occupancy grid descriptor based on visual input includes: extracting image features from surround images; obtaining BEV features based on a BEV encoder; using a voxel decoder to decode image features and BEV features in voxel space to obtain a three-dimensional space occupancy grid descriptor. The method for simulating a driving scene based on a three-dimensional occupancy grid provided in the present application proposes a general framework structure based on the three-dimensional occupancy grid representation, which can construct a three-dimensional occupancy grid description of a three-dimensional environment from a surround image, and support multiple downstream perception and regulation tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer vision technology, and particularly to a method, system, device and medium for simulating driving scenarios based on three-dimensional occupancy grids. Background Art

[0002] Human drivers can easily recognize complex traffic scenes through the visual system, and this precise environmental perception ability is very important for vehicle control. We need a driving system that can describe various objects on the road, such as irregular large trucks, static obstacles, moving pedestrians, etc.

[0003] Traditional representation methods such as 3D bounding boxes, point cloud segmentation, BEV segmentation, LiDAR segmentation, etc. are difficult to accurately describe the surrounding environment, thus affecting the performance of the driving system. For example: 3D object detection uses 3D bounding boxes as the perception target in autonomous driving. This representation method only cares about foreground objects and overly simplifies the geometric shape information of the objects, and cannot describe irregular objects. The task of LiDAR segmentation is point-level three-dimensional scene understanding. It requires point clouds as input, which are expensive, and LiDAR is affected by limited sensing range and sparsity in three-dimensional scene description, so it is difficult to completely describe the three-dimensional environment. Summary of the Invention

[0004] The embodiments of the present application provide a method, system, device and medium for simulating driving scenarios based on three-dimensional occupancy grids, and propose a general framework structure based on three-dimensional occupancy grid representation, which can construct a three-dimensional occupancy grid description of a three-dimensional environment from panoramic images and support multiple downstream perception and planning and control tasks.

[0005] To solve the above technical problems, in a first aspect, the embodiments of the present application provide a method for simulating driving scenarios based on three-dimensional occupancy grids, including the following steps: First, based on visual input, construct a three-dimensional space occupancy grid descriptor; then, based on the occupancy grid descriptor, implement downstream tasks; wherein, based on visual input, constructing a three-dimensional space occupancy grid descriptor includes: extracting image features from panoramic images; obtaining BEV features based on a BEV encoder; using a voxel decoder to decode the image features and the BEV features in voxel space to obtain a three-dimensional space occupancy grid descriptor.

[0006] In some exemplary embodiments, the downstream tasks include 3D semantic scene reconstruction, 3D detection, BEV segmentation, and motion planning.

[0007] In some exemplary embodiments, after implementing the downstream tasks based on the occupancy grid descriptor, it further includes: constructing an occupancy grid dataset based on the nuScenes dataset.

[0008] In some exemplary embodiments, based on the nuScenes dataset, an occupancy grid dataset is constructed, including: accumulating the sparse point clouds in the nuScenes dataset into dense point clouds, and processing the moving foreground objects to construct the occupancy grid dataset.

[0009] In some exemplary embodiments, accumulating the sparse point clouds in the nuScenes dataset into dense point clouds and processing the moving foreground objects includes: decomposing the point clouds into static background point clouds and object point clouds based on 3D bounding boxes; accumulating the background point clouds and the object point clouds in the world coordinate system and the object coordinate system respectively to obtain dense point clouds; performing voxelization processing on the three-dimensional space based on the dense point clouds, and obtaining the labels of the voxels based on the voting method to generate occupancy grid data, thereby obtaining the occupancy grid dataset.

[0010] In some exemplary embodiments, the occupancy grid dataset includes 34,149 frames, 700 training scenarios, and 150 validation scenarios.

[0011] In a second aspect, an embodiment of the present application further provides a three-dimensional occupancy grid driving scenario simulation system, including: an occupancy grid construction module and an application module connected in sequence; the occupancy grid construction module is used to construct a three-dimensional space occupancy grid descriptor according to visual input; the application module is used to implement downstream tasks according to the occupancy grid descriptor; wherein, constructing the three-dimensional space occupancy grid descriptor based on visual input includes: extracting image features from the surround-view images; obtaining BEV features based on a BEV encoder; using a voxel decoder to decode the image features and the BEV features in the voxel space to obtain the three-dimensional space occupancy grid descriptor.

[0012] In some exemplary embodiments, the above three-dimensional occupancy grid driving scenario simulation system further includes: an evaluation module; the evaluation module is used to construct an occupancy grid dataset according to the nuScenes dataset and evaluate the simulation system based on the occupancy grid dataset.

[0013] In addition, the present application also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above three-dimensional occupancy grid driving scenario simulation method.

[0014] In addition, the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above three-dimensional occupancy grid driving scenario simulation method is implemented.

[0015] The technical solutions provided by the embodiments of the present application have at least the following advantages:

[0016] An embodiment of the present application provides a method, system, device, and medium for simulating a driving scenario based on a three-dimensional occupancy grid. The method includes the following steps: First, based on visual input, construct a three-dimensional space occupancy grid descriptor; then, based on the occupancy grid descriptor, implement downstream tasks. Among them, constructing a three-dimensional space occupancy grid descriptor based on visual input includes: extracting image features from panoramic images; obtaining BEV features based on a BEV encoder; using a voxel decoder to decode the image features and the BEV features in voxel space to obtain a three-dimensional space occupancy grid descriptor.

[0017] The present application proposes a general framework structure OccNet based on the occupancy grid occupancy representation, constructs a robust occupancy grid feature based on panoramic images, constructs an occupancy grid description of a three-dimensional environment from panoramic images, and supports multiple downstream perception and planning and control tasks. OccNet includes two stages, occupancy reconstruction and occupancy utilization. In the occupancy reconstruction stage, first extract features from panoramic images, obtain BEV features using a BEV encoder, and then use a cascaded voxel decoder to decode the features in voxel space to obtain a high-quality occupancy grid descriptor. In the occupancy utilization stage, based on the occupancy grid descriptor, the present application can implement various downstream tasks, such as 3D detection and 3D semantic scene reconstruction. Based on map segmentation in the BEV space, the present application can implement planning and control tasks. Description of the Drawings

[0018] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings. These exemplary illustrations do not limit the embodiments unless otherwise stated. The figures in the drawings do not constitute a proportional limitation.

[0019] Figure 1 Schematic diagram of three-dimensional occupancy grid representation;

[0020] Figure 2 Flow schematic diagram of a method for simulating a driving scenario based on a three-dimensional occupancy grid provided by an embodiment of the present application;

[0021] Figure 3 Schematic diagram of the overall structure of a general network framework for three-dimensional occupancy provided by an embodiment of the present application;

[0022] Figure 4 Flow schematic diagram of a method for simulating a driving scenario based on a three-dimensional occupancy grid provided by another embodiment of the present application;

[0023] Figure 5 Schematic diagram of the structure of a system for simulating a driving scenario based on a three-dimensional occupancy grid provided by an embodiment of the present application;

[0024] Figure 6 A schematic diagram for visual comparison of the semantic scene completion task provided by an embodiment of the present application;

[0025] Figure 7 A schematic diagram of the trajectory prediction result of the motion planning task provided by an embodiment of the present application;

[0026] Figure 8 A schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0027] As can be seen from the background art, traditional representation methods such as 3D bounding boxes, point cloud segmentation, BEV segmentation, and LiDAR segmentation are difficult to accurately describe the surrounding environment, thus affecting the performance of the driving system.

[0028] In the real environment, a human driver can easily describe the environment and can accurately describe the real world in the form of "occupancy" to achieve safe driving behaviors, such as "there is a car about 10m to my left". However, this is not easy for vision-centered autonomous driving systems because of the various entities in the scene, including vehicles such as cars, SUVs, and engineering trucks, as well as static obstacles, pedestrians, background buildings, and vegetation. Quantifying the three-dimensional scene into a structured unit with semantic labels, called three-dimensional occupancy grid, is an intuitive solution, which is also advocated by the automotive industry such as Mobileye and Tesla, as Figure 1 shown. Compared with the 3D bounding box description that overly simplifies the object shape, the three-dimensional occupancy can accurately depict the geometry of the object, as Figure 1 shown in (c) and (d) of

[0029] Three-dimensional object detection uses 3D bounding boxes as the perception targets in autonomous driving. This representation method only cares about foreground objects and overly simplifies the geometric shape information of the objects, and cannot describe irregular objects. In contrast, the three-dimensional occupancy grid is a fine-grained description of the physical world and can describe the static environment and moving foreground objects in the driving scene. The task of LiDAR segmentation is point-level three-dimensional scene understanding. It requires point clouds as input, which are expensive, and LiDAR is affected by the limited sensing range and sparsity in the three-dimensional scene description, so it is difficult to completely describe the three-dimensional environment. Inferring the three-dimensional structure of the scene from images is common in the field of computer vision but has been challenging for years. Most methods in this field only handle single tasks, such as 3D detection, point cloud segmentation, and scene completion. For the field of autonomous driving applications, a general scene description method needs to be proposed, and a suitable model framework is required to solve a series of autonomous driving tasks.

[0030] Compared with traditional representation methods, the occupancy-based representation can accurately describe the geometric structure and semantic information of 3D scenes. Occupancy grids can capture fine-grained details and distance information of objects in the scene, such as key obstacles, and thus are better used for downstream tasks. Previous studies on occupancy mainly focused on the task of 3D scene completion, ignoring the application potential of occupancy in different driving tasks, such as 3D detection and motion planning.

[0031] This kind of scene geometric description of occupancy shows its potential in enhancing traditional perception tasks and downstream planning tasks. In the field of robotics, occupancy grids are a common representation in mobile navigation, but they are only used as the search space for planning. In 3D Semantic Scene Completion (SSC), occupancy representations of 3D space can be reconstructed for visual and point cloud inputs. Existing technologies usually focus on how to design better network structures to predict 3D occupancy representations based on image information, but ignore the application value of occupancy in other perception tasks. Current research work on occupancy is also limited to the SemanticKitti dataset, which has great limitations in practical applications because it only contains a front-view camera. Therefore, we also need to construct a high-quality surround-view occupancy dataset to better explore occupancy-related tasks.

[0032] To solve the above technical problems such as "being unable to describe irregular objects and difficult to completely describe the 3D environment", the embodiment of this application provides a driving scene simulation method based on 3D occupancy grids, including the following steps: First, based on visual input, construct a 3D space occupancy grid descriptor; then, based on the occupancy grid descriptor, implement downstream tasks; among them, constructing a 3D space occupancy grid descriptor based on visual input includes: extracting image features from surround-view images; obtaining BEV features based on a BEV encoder; using a voxel decoder to decode the image features and the BEV features in voxel space to obtain a 3D space occupancy grid descriptor. This application proposes a general framework structure based on 3D occupancy grid representation, which can construct 3D occupancy grid descriptions of 3D environments from surround-view images and support multiple downstream perception and planning and control tasks.

[0033] The following will elaborate on the embodiments of the present application in conjunction with the accompanying drawings. However, those of ordinary skill in the art can understand that in the embodiments of the present application, many technical details are provided to help readers better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented.

[0034] Referring to Figure 2 , an embodiment of the present application provides a method for simulating a driving scenario based on a three-dimensional occupancy grid, including the following steps:

[0035] Step S1: Based on visual input, construct a three-dimensional spatial occupancy grid descriptor.

[0036] Step S2: Based on the occupancy grid descriptor, implement downstream tasks.

[0037] Among them, in step S1, based on visual input, constructing a three-dimensional spatial occupancy grid descriptor includes the following steps:

[0038] Step S101: Extract image features from the surround-view image.

[0039] Step S102: Based on the BEV encoder, obtain the BEV features.

[0040] Step S103: Use a voxel decoder to decode the image features and the BEV features in the voxel space to obtain a three-dimensional spatial occupancy grid descriptor.

[0041] Since the occupancy grid in three-dimensional space can capture fine-grained details and distance information of objects in the scene, such as key obstacles, the representation of the occupancy grid can accurately describe the geometric structure and semantic information of the three-dimensional scene, and thus be better used for downstream tasks. The present application first constructs a three-dimensional spatial occupancy grid descriptor, and then based on the occupancy grid descriptor, implements downstream tasks. The present application constructs a general network based on three-dimensional occupancy (Occupancy Network, OccNet). The core of the network framework is a three-dimensional spatial occupancy descriptor constructed based on visual input. Based on this descriptor, various downstream tasks can be solved.

[0042] Figure 3 A schematic diagram showing the overall structure of the general network OccNet for three-dimensional occupancy is shown. As Figure 3As shown in the figure, OccNet consists of two stages: occupancy reconstruction and occupancy utilization. In the occupancy reconstruction stage, features are first extracted from the surround-view images, and BEV features are obtained using a BEV encoder. Then, a cascaded voxel decoder is used to decode the features in the voxel space to obtain a high-quality occupancy descriptor.

[0043] As Figure 3 shown, the left half is the Occupancy reconstruction stage, and the right half is the occupancy utilization stage. In the occupancy reconstruction stage, for the current time t, image features F are first extracted from the surround-view images t , combined with the BEV features B at time t-1 t-1 and the BEV query Q at the current time t , and the BEV features B at the current time are obtained using a BEV encoder t . Then, the image features, the BEV features at the historical time and the current time are all decoded in the cascaded voxel decoder to obtain Occupancy features.

[0044] To learn better scene features, a cascaded structure is designed in the decoder in this application. Based on the BEV features, this application gradually recovers the information of the features in the height dimension to obtain a robust voxel feature structure. Through N-step progressive operations, we convert the BEV features B t into Occupancy features V t , where the voxel features in the intermediate steps are defined as V' t,i . As Figure 3 shown, the BEV features B t-1 and B t are first converted into V' t-1,i and V' t,i through a neural network, and an optimized voxel feature V' t,i is obtained through the decoder operation in the i-th step. The subsequent decoding process is similar. Each step of the voxel decoder contains voxel-based temporal self-attention and spatial cross-attention mechanisms. Based on the historical features V' t-1,i and the current image features F t we optimize the voxel feature V' t,i .

[0045] To implement the voxel-based temporal self-attention operation, we first combine the voxel features V' at time t-1 t-1,i with the features V' at the current time t,iFeature alignment is performed, followed by self-attention operation. When performing self-attention, since each query has to operate with keys and values, its computational complexity is positively correlated with the number of features. Performing self-attention operation in the 3D voxel space requires a large computational cost. Therefore, this application proposes a 3D deformable attention operator to efficiently perform self-attention operations.

[0046] In the operation of spatial cross-attention, based on the 2D deformable attention mechanism, we perform the interaction between voxel features V′ t,i and image features Ft. In this operation, in order to maintain the height information of the features, image features are used to optimize voxel features.

[0047] In some exemplary embodiments, the downstream tasks include 3D semantic scene reconstruction, 3D detection, BEV segmentation, and motion planning.

[0048] As Figure 3 shown, in the occupancy utilization stage, Occupancy is applicable to different downstream tasks. OccNet can accurately describe the 3D space scene through the Occupancy descriptor. Therefore, this application can design a variety of lightweight prediction structures to directly input the Occupancy features into various tasks.

[0049] Semantic scene completion: Use a multi-layer perceptron MLP to predict the semantic labels of each voxel. To handle the problem of positive and negative sample imbalance between occupied voxels and empty voxels, this application uses Focal Loss as the loss function.

[0050] 3D detection: This application compresses the Occupancy descriptor into the BEV space, and then applies the detection head of the deformable DETR task to predict the 3D bounding box.

[0051] BEV segmentation: This application compresses the Occupancy descriptor into the BEV space, and then performs the prediction of BEV segmentation. BEV segmentation includes the prediction of drivable areas and lanes for map representation, and the prediction of vehicles and pedestrians for semantic segmentation.

[0052] Motion planning: Whether it is 3D occupancy or 3D bounding boxes, they can first be converted into the results of BEV segmentation and then used for motion planning tasks, such as Figure 3As shown. Based on 3D occupancy, this application converts each BEV cell into a 0-1 format, where 1 indicates that the cell is occupied and 0 indicates an empty area. Under the BEV map, this application designs multiple random speeds, accelerations, and curvatures to sample multiple candidate trajectories, and then filters out an optimal trajectory based on cost functions such as safety and comfort. Then, the image features and GRU cells are used to optimize the trajectory to obtain the final trajectory for the motion planning task.

[0053] See Figure 4 , in some embodiments, after implementing the downstream task based on the occupancy grid descriptor in step S2, it further includes:

[0054] Step S3: Based on the nuScenes dataset, construct an occupancy grid dataset.

[0055] In some embodiments, step S3 constructs an occupancy grid dataset based on the nuScenes dataset, including: accumulating the sparse point clouds in the nuScenes dataset into dense point clouds, and processing the moving foreground objects to construct the occupancy grid dataset.

[0056] In some embodiments, accumulating the sparse point clouds in the nuScenes dataset into dense point clouds and processing the moving foreground objects includes: decomposing the point cloud into static background point clouds and object point clouds based on the 3D bounding box; accumulating the background point clouds and the object point clouds in the world coordinate system and the object coordinate system respectively to obtain dense point clouds; performing voxelization processing on the three-dimensional space based on the dense point clouds, and obtaining the labels of the voxels based on the voting method to generate occupancy grid data, thereby obtaining the occupancy grid dataset.

[0057] To evaluate OccNet, this application constructs the first high-quality dense Occupancy dataset OpenOcc based on the nuScenes dataset. This dataset contains 34,149 frames, 700 training scenes, and 150 validation scenes. This application annotates more than 1.4 billion voxels and 16 classes, including 10 foreground objects and 6 background objects. This application accumulates the sparse point clouds into dense point clouds, and also additionally considers the processing of moving foreground objects. Based on the 3D bounding box, this application first decomposes the point cloud into static background point clouds and object point clouds, and then accumulates the background point clouds and object point clouds in the world coordinate system and the object coordinate system respectively to obtain a dense representation. Based on the dense point clouds, the three-dimensional space is voxelized, and then the labels of the voxels are obtained based on the voting method to generate occupancy data. At the same time, this application also performs a series of optimizations on the generated data to ensure the quality of the dataset.

[0058] See Figure 5, the embodiment of the present application also provides a three-dimensional occupancy grid driving scenario simulation system, including: an occupancy grid construction module 101 and an application module 102 connected in sequence; the occupancy grid construction module 101 is used to construct a three-dimensional space occupancy grid descriptor according to visual input; the application module 102 is used to implement downstream tasks according to the occupancy grid descriptor; wherein, constructing a three-dimensional space occupancy grid descriptor based on visual input includes: extracting image features from a panoramic image; obtaining BEV features based on a BEV encoder; using a voxel decoder to decode the image features and the BEV features in the voxel space to obtain a three-dimensional space occupancy grid descriptor.

[0059] In some exemplary embodiments, the above three-dimensional occupancy grid driving scenario simulation system further includes: an evaluation module 103; the evaluation module is used to construct an occupancy grid dataset according to the nuScenes dataset and evaluate the simulation system based on the occupancy grid dataset.

[0060] To verify this new occupancy scenario representation and the general network framework OccNet, the present application constructs the first high-quality dense Occupancy dataset based on the publicly available autonomous driving dataset nuScenes. The present application conducts sufficient experimental analysis to illustrate the application value of occupancy in multiple tasks. For example, compared with the 3D box-based representation, occupancy-based motion planning can significantly reduce the collision rate of the driving system.

[0061] In summary, the present application provides a three-dimensional occupancy grid driving scenario simulation method and system. Through the general scenario representation framework OccNet, various downstream tasks of autonomous driving can be processed. The core of the network framework is a three-dimensional space occupancy descriptor constructed based on visual input. Based on this descriptor, we can solve various downstream tasks. In addition, the present application also constructs the occupancy dataset OpenOcc in the field of autonomous driving, providing a high-quality dense Occupancy dataset for the autonomous driving community, which can be used to verify the performance of occupancy-related tasks.

[0062] Compared with the prior art, the advantages of the present application are as follows:

[0063] 1. Compared with traditional representation methods such as 3D detection, three-dimensional occupancy can better describe the three-dimensional environment and accurately describe the foreground and background.

[0064] 2. The general network OccNet of three-dimensional occupancy proposed in the present application can be applied to different tasks at the same time, fully demonstrating the application value of occupancy in various different tasks.

[0065] 3. Compared with existing technical solutions, OccNet is a multi-view vision-centered network structure that includes a temporal cascaded voxel decoding structure to obtain a 3D spatial occupancy grid descriptor. The occupancy grid descriptor can be applied to a wide range of driving tasks, including 3D scene completion, detection, segmentation, and planning. OccNet has achieved leading performance in tasks such as 3D detection, 3D scene completion, and motion planning.

[0066] To verify the feasibility of the 3D occupancy grid-based driving scene simulation method and system provided in this application, sufficient experimental analysis and verification were carried out on the method provided in this application in various downstream tasks, demonstrating the feasibility and superiority of the method provided in this application.

[0067] For the semantic scene completion task (SSC), by fully comparing different solutions, the comparison results were obtained, as shown in Table 1. The method provided in this application significantly outperforms other solutions in various indicators. In addition, this application also conducted a visual comparison of different methods for the semantic scene completion task; as Figure 6 shown, this application conducted a visual comparison of the prediction results and found that: compared with the current SOTA solutions such as TPVFormer, the general network OccNet for 3D occupancy provided in this application can more accurately predict foreground objects such as pedestrians. Among them, the performance comparison results of the SSC task are shown in Table 1 below:

[0068] Table 1: Performance comparison of SSC tasks

[0069]

[0070] For the motion planning task, this application compared the impacts of 3D detection and Occupancy on the motion planning task. Compared with 3D detection, occupancy can provide more abundant information for realizing the motion planning task. As shown in Table 2, based on the occupancy task, we can reduce the collision rate by 15%-58% and achieve safer driving. Figure 7 The schematic diagram of trajectory prediction for the motion planning task is shown, where the lower left corner is the trajectory prediction result based on 3D detection, and the lower right corner is the comparison of the trajectory prediction result based on Occupancy. From Figure 7 it can be seen that based on occupancy information, a more reasonable motion trajectory can be planned.

[0071] Table 2: Impacts of different inputs on the planning task

[0072]

[0073] See Figure 8, Another embodiment of the present application provides an electronic device, including: at least one processor 110; and a memory 111 communicatively connected to the at least one processor; wherein, the memory 111 stores instructions executable by the at least one processor 110, and the instructions are executed by the at least one processor 110 to enable the at least one processor 110 to execute any of the above method embodiments.

[0074] Wherein, the memory 111 and the processor 110 are connected in a bus manner. The bus may include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors 110 and the memory 111 together. The bus may also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, and thus will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver may be an element or multiple elements, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium. The data processed by the processor 110 is transmitted on the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor 110.

[0075] The processor 110 is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory 111 can be used to store data used by the processor 110 when performing operations.

[0076] Another embodiment of the present application relates to a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the above method embodiments are implemented.

[0077] That is, those skilled in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by instructing relevant hardware through a program. The program is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the above methods of various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0078] With the above technical solutions, the embodiments of the present application provide a method, system, device, and medium for simulating a driving scenario based on a three-dimensional occupancy grid. The method includes the following steps: First, based on visual input, construct a three-dimensional spatial occupancy grid descriptor; then, based on the occupancy grid descriptor, implement downstream tasks. Among them, constructing a three-dimensional spatial occupancy grid descriptor based on visual input includes: extracting image features from panoramic images; obtaining BEV features based on a BEV encoder; using a voxel decoder to decode the image features and the BEV features in the voxel space to obtain a three-dimensional spatial occupancy grid descriptor.

[0079] Based on the occupancy grid representation, the present application proposes a general framework OccNet, constructs robust occupancy grid features based on panoramic images, constructs an occupancy grid description of a three-dimensional environment from panoramic images, and supports multiple downstream perception and planning and control tasks. OccNet includes two stages, occupancy reconstruction and occupancy utilization. In the occupancy reconstruction stage, first extract features from panoramic images, use a BEV encoder to obtain BEV features, and then use a cascaded voxel decoder to decode the features in the voxel space to obtain a high-quality occupancy grid descriptor. In the occupancy utilization stage, based on the occupancy grid descriptor, the present application can implement various downstream tasks, such as 3D detection and 3D semantic scene reconstruction. Based on map segmentation in the BEV space, the present application can implement planning and control tasks.

[0080] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present application. In actual applications, various changes can be made in form and details without departing from the spirit and scope of the present application. Any person skilled in the art can make their own changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application should be subject to the scope defined by the claims.

Claims

1. A three-dimensional occupancy grid-based driving scenario simulation method, characterized in that, Including: Construct a three-dimensional spatial occupancy grid descriptor based on visual input; Implement downstream tasks based on the occupancy grid descriptor; Among them, the constructing of the three-dimensional spatial occupancy grid descriptor based on visual input includes: For the current moment t, extract the image feature F from the surround-view image t ; Based on the BEV feature B at time t-1 t-1 and the BEV query Q at the current time t , through the BEV encoder, obtain the BEV feature B at the current time t ; Using a voxel decoder, decode the image feature F in the voxel space t and the BEV features at historical and current moments to obtain a three-dimensional spatial occupancy grid descriptor; The voxel decoder has a cascaded structure and converts the BEV feature B into a three-dimensional occupancy grid feature V through N-step progressive operations. Among them, the voxel feature at the intermediate step is defined as; the BEV features B and B are first converted into and through a neural network, and the optimized Voxel feature is obtained through the voxel decoder operation in the i-th step, and then decoding is performed; each step of the voxel decoder includes a voxel-based temporal self-attention and a spatial cross-attention mechanism, which optimizes the voxel feature based on the historical feature and the current image feature F. t into a three-dimensional occupancy grid feature V t , where the voxel feature at the intermediate step is defined as ; the BEV feature B t-1 and B t are first converted into and through a neural network, and the optimized Voxel feature is obtained through the voxel decoder operation in the i-th step, and then decoding is performed; each step of the voxel decoder includes a voxel-based temporal self-attention and a spatial cross-attention mechanism, which optimizes the voxel feature based on the historical feature and the current image feature F t , for the voxel feature ; The optimization process includes: First, align the voxel features at time t-1 with the features at the current time for feature alignment, then perform the operation of spatial cross-attention, and perform the interaction between the voxel features and the image feature F t based on the 2D deformable attention mechanism.

2. The method for simulating a driving scenario based on a three-dimensional occupancy grid according to claim 1, wherein The downstream tasks include 3D semantic scene reconstruction, 3D detection, BEV segmentation, and motion planning.

3. The method for simulating a driving scenario based on a three-dimensional occupancy grid according to claim 1, wherein After implementing the downstream tasks based on the occupancy grid descriptor, it further includes: Construct an occupancy grid dataset based on the nuScenes dataset.

4. The method for simulating a driving scenario based on a three-dimensional occupancy grid according to claim 3, wherein The constructing of the occupancy grid dataset based on the nuScenes dataset includes: Accumulate the sparse point clouds in the nuScenes dataset into dense point clouds, and process the moving foreground objects to construct an occupancy grid dataset.

5. The method for simulating a driving scenario based on a three-dimensional occupancy grid according to claim 4, characterized in that, The accumulating of the sparse point clouds in the nuScenes dataset into dense point clouds and the processing of the moving foreground objects include: Decompose the point cloud into static background point clouds and object point clouds based on 3D bounding boxes; Accumulate the background point clouds and the object point clouds in the world coordinate system and the object coordinate system respectively to obtain dense point clouds; Perform voxelization processing on the three-dimensional space based on the dense point clouds, and obtain the labels of the voxels based on the voting method to generate occupancy grid data and obtain an occupancy grid dataset.

6. The method for simulating a driving scenario based on a three-dimensional occupancy grid according to claim 4, wherein The occupancy grid dataset includes 34,149 frames, 700 training scenes, and 150 validation scenes.

7. A three-dimensional occupancy grid-based driving scenario simulation system, characterized in that, Including: An occupancy grid construction module and an application module connected in sequence; The occupancy grid construction module is used to construct a three-dimensional spatial occupancy grid descriptor according to visual input; The application module is used to implement downstream tasks according to the occupancy grid descriptor; Among them, the constructing of the three-dimensional spatial occupancy grid descriptor according to visual input includes: For the current moment t, extract the image feature F from the surround-view image t ; Based on the BEV feature B at time t-1 t-1 and the BEV query Q at the current time t , through the BEV encoder, obtain the BEV feature B at the current time t ; Using a voxel decoder, decode the image feature F in the voxel space t , the BEV features at the historical moment and the current moment, to obtain a three-dimensional spatial occupancy grid descriptor; The voxel decoder has a cascaded structure and converts the BEV feature B into a 3D occupancy grid feature V through N-step progressive operations. t Among them, the voxel features at intermediate steps are defined as t ; The BEV features B and B t-1 and B t are first transformed into and through a neural network, and the optimized Voxel features are obtained through the voxel decoder operation in the i-th step, and then decoding is performed; Each step of the voxel decoder contains a voxel-based temporal self-attention and spatial cross-attention mechanism, and based on the historical feature and the current image feature F t , the voxel feature is optimized. The optimization process includes: First, align the voxel features at time t-1 with the features at the current time for feature alignment, then perform the operation of spatial cross-attention, and perform the interaction between voxel features and image feature F t based on the 2D deformable attention mechanism.

8. The 3D occupancy grid-based driving scenario simulation system according to claim 7, wherein It further includes: An evaluation module; The evaluation module is used to construct an occupancy grid dataset according to the nuScenes dataset and evaluate the simulation system based on the occupancy grid dataset.

9. An electronic device, characterized in that, Including: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the three-dimensional occupancy grid-based driving scenario simulation method as described in any one of claims 1 to 6.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the three-dimensional occupancy grid-based driving scenario simulation method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Automatic driving track prediction method based on space-time pyramid

    CN115049130A

  • Method and electronic device for 3D object detection using neural networks

    US20230121534A1