Space occupation estimation method based on 2D-3D joint training

By adopting the 2D-3D joint training method and cross-modal query mechanism in 3D space occupation prediction, the problem of dependence and high-quality 3D data annotation in the existing methods is solved, and a more efficient and stable 3D occupancy prediction effect is achieved.

CN119964144APending Publication Date: 2025-05-09DALIAN UNIV OF TECH +2
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510052714.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing 3D space occupancy prediction methods rely on a large number of high-quality 3D data annotations, and the computational volume and memory overhead are relatively large, which is seriously limited by the size of the data set.

Method used

The space occupancy estimation method based on 2D-3D joint training is adopted, and the semantic features of 2D images and 3D voxel features are uniformly processed through a cross-modal query mechanism. A unified 2D-3D feature representation is constructed using pixel-level modules and voxel aggregation modules to realize the joint optimization of 2D semantic segmentation and 3D occupancy prediction.

Benefits of technology

Through the shared query mechanism and unified 2D-3D feature representation, the accuracy and stability of 3D occupancy prediction are improved, redundant calculations are reduced, and the model's inference efficiency and generalization ability are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964144A_ABST
    Figure CN119964144A_ABST
Patent Text Reader

Abstract

The invention belongs to automatic driving, deep learning and 3D occupancy prediction, and provides a space occupancy estimation method based on 2D-3D joint training. According to the method, a cross-modal query mechanism is used, the semantic features of the 2D image and the 3D voxel features are fused, and the precision of 3D occupancy prediction is improved. Unified query parameters are shared between 2D semantic segmentation and 3D occupancy prediction tasks, so that the calculation amount is reduced while the model prediction accuracy is ensured, and the generalization ability of the model to complex scenes is enhanced. According to the invention, effective interaction is carried out between 2D and 3D features through a cross attention mechanism, so that the prediction effect of space occupation is optimized. The method is especially suitable for a real-time scene perception task in an automatic driving environment, and the recognition and prediction capability of target space occupation in a dynamic environment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to autonomous driving, deep learning, and 3D occupancy prediction, and relates to a space occupancy estimation method based on 2D-3D joint training, specifically to joint learning, a space occupancy prediction method Voxformer, a scene layout prediction algorithm, and a space structure generation mechanism. Background Art

[0002] 3D space occupancy prediction is a core task in fields such as autonomous driving, robot navigation, and augmented reality. It aims to predict the occupancy state of voxels in the scene, including both visible areas and occluded spaces. 3D space occupancy prediction usually consists of an encoder backbone network, depth-semantic prediction, and a 3D voxel prediction result generator. It generates 3D space occupancy information in voxel form by combining the depth information and semantic information corresponding to the predicted image.

[0003] Traditional methods rely on voxel-by-voxel classification, such as MonoScene (Monoscene: Monocular 3d semantic scene completion) and other methods. This method regresses the occupancy and category of each occupied area in the space in the form of voxels. This method requires a high amount of computation. In recent years, some 3D occupancy prediction based on bird's-eye view (BEV) feature representation, such as VoxFormer (Voxformer: Sparse voxel transformer for camera-based 3d semantic scene completion), has gradually developed. This type of method introduces spatial occupancy priors through depth information during calculation, and marks all occupied voxels, which greatly reduces the amount of computation and memory overhead. However, 3D occupancy prediction methods are highly dependent on the scale and quality of data, especially the true value annotation of 3D occupancy prediction consumes more labor than 2D data, which makes the current methods severely limited by the scale of the data set.

[0004] To address this problem, the present invention introduces cross-modal learning, aiming to bring new breakthroughs to 3D space occupancy prediction. 2D semantic segmentation and 3D occupancy prediction have a natural correlation, and the prediction accuracy of the model can be improved through joint training and cross-modal interaction. Specifically, the present invention proposes a method for enhancing 3D space occupancy prediction by 2D semantic segmentation, and using the same set of parameters to perform synchronized 2D segmentation and 3D space occupancy prediction to enhance 3D space occupancy prediction. Summary of the invention

[0005] The present invention proposes a space occupancy estimation method based on 2D-3D joint training. By introducing a cross-modal query mechanism, the semantic features of 2D images and 3D voxel features are processed uniformly, solving the problem of insufficient cross-modal information integration in existing methods.

[0006] The technical features of the present invention include:

[0007] 1. Use cross-modal queries to achieve joint optimization of 2D semantic segmentation and 3D occupancy prediction tasks.

[0008] 2. A unified 2D-3D feature representation is constructed using pixel-level modules and voxel aggregation modules, which improves the accuracy and stability of 3D occupancy prediction.

[0009] 3. Utilize a joint training strategy to improve the prediction effect of 2D semantic segmentation by sharing query information, thereby further enhancing the 3D occupancy prediction performance.

[0010] The technical solution of the present invention:

[0011] A space occupancy estimation method based on 2D-3D joint training, the steps are as follows:

[0012] Step 1: Image 2D pixel feature extraction

[0013] The RGB image is fed into the backbone network ResNet, which is an image feature extractor composed of several residual blocks. After large-scale pre-training on ImageNet, feature maps at multiple downsampling levels are extracted. The feature maps are fed into the pixel extraction module together with the initialized cross-modal query for 2D feature extraction.

[0014] Among them, cross-modal query is a set of learnable, randomly initialized tensor arrays. The cross-modal query is randomly initialized as a learnable parameter with a dimension of (n, c), where n represents the number of cross-modal queries and c represents the number of channels of each cross-modal query. The pixel extraction module imitates the structure of MaskFormer (Per-Pixel Classification is Not All You Need for Semantic Segmentation), which is mainly composed of a Transformer Decoder module and a linear layer. The Transformer Decoder module consists of 6 stacked Transformer blocks, and the linear layer consists of several fully connected layers and BN layers.

[0015] Specifically, the cross-modal query and the feature map output by the backbone network are input into the Transformer Decoder module to interact through the cross-attention operation. The output results of the Transformer Decoder module are respectively sent to two linear layers for channel dimension transformation, and finally generate category prediction and mask labeling; at the same time, the feature pyramid method is used to perform multi-scale feature fusion on the feature map output by the backbone network to generate a 2D pixel embedding of size (c, h, w), where h and w are set to one eighth of H and W; finally, the above operation transforms the input original RGB image into a 2D pixel embedding of size (n, n class ) category prediction, mask tags of size (n,c) and 2D pixel embeddings of size (c,h,w), where n class is the number of categories in 3D space occupancy prediction;

[0016] Step 2: 3D polymerization;

[0017] The 3D aggregation module consists of depth-point cloud-voxel projection, deformable cross attention and deformable self-attention. The 3D aggregation module inputs a depth map of size (H, W) and feature maps of multiple down-sampling levels generated in step 1, and promotes the feature maps of multiple down-sampling levels obtained in step 1 to 3D space through interaction.

[0018] Specifically, the depth map is first projected into the 3D space into a point cloud using the pre-known camera parameters, and a voxel space is defined with a size of (h 3d ,w 3d ,z 3d ), respectively representing the length, width and height of the voxel in 3D space, h 3d ,w 3d is one eighth of the input image, z 3d is 16; then the occupied grids in the voxel space are set to an occupied state according to the projected point cloud; then the marked voxel space is initialized to a size of (c, h 3d ,w 3d ,z 3d ), where the occupied grid fill length is c, and the size is (c,h 3d ,w 3d ,z 3d ), and then send the obtained 3D voxel query and the feature map obtained in step 1 to the deformable cross attention defined in Deformable DETR for interaction, and then perform the deformable self-attention used in GroundingDino; fill the unoccupied areas in the 3D voxel query, and finally get a size of (c, h 3d ,w 3d ,z 3d)’s 3D voxel embedding;

[0019] Step 3: Joint reasoning;

[0020] The joint reasoning module does not contain any learnable parts and consists of several matrix multiplications. It accepts the 2D pixel embedding, mask label, 3D voxel embedding and category prediction generated in step 1 and step 2 to generate the final 2D semantic segmentation map and 3D voxel prediction map. Specifically, the 2D pixel embedding is first multiplied with the mask label to obtain a 2D mask prediction of size (n,h,w), which is then multiplied with the category prediction to obtain a 2D mask prediction of size (n class ,h,w) 2D semantic segmentation result; at the same time, the 3D voxel embedding is multiplied with the mask label to obtain a size of (n,h 3d ,w 3d ,z 3d ) is multiplied by the category prediction to obtain a 3D mask prediction of size (n class ,h 3d ,w 3d ,z 3d )’s 3D space occupancy prediction result; after this inference, the results of 2D semantic segmentation and three-dimensional space prediction are output at the same time.

[0021] Beneficial effects of the present invention:

[0022] (1) The joint optimization of 2D semantic segmentation and 3D occupancy prediction is achieved through a shared query mechanism, which fully explores the potential relationship between the two and improves model performance.

[0023] (2) The cross-modal query mechanism reduces the redundant calculations caused by independent feature extractors and improves the reasoning efficiency of the model.

[0024] (3) The introduction of a unified 2D-3D feature representation effectively enhances the model's generalization ability for complex scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 Based on the overall algorithm flow chart.

[0026] Figure 2 Schematic diagram of the pixel extraction module structure.

[0027] Figure 3 Schematic diagram of the 3D aggregation module.

[0028] Figure 4 Schematic diagram of the joint reasoning module. DETAILED DESCRIPTION

[0029] The specific implementation of the present invention is further described below in conjunction with the accompanying drawings and technical solutions.

[0030] Figure 1The flow chart of the volume algorithm is shown. In the upper branch of the method, the input single-frame RGB image is subjected to feature extraction through a pixel-level module to generate multi-scale image features. These features are then interacted with the initialized cross-modal query, which contains category tags and mask embeddings to characterize the semantic category and geometric occupancy state of each voxel in the scene. Through the cross-attention mechanism, these 2D features are combined with the cross-modal query to generate a unified 2D-3D feature representation, thereby providing support for subsequent 3D occupancy prediction.

[0031] In the lower branch, based on the dense depth map generated by the depth estimation module, the pixel information of the RGB image is projected into the 3D point cloud space, and the generated point cloud is voxelized. Subsequently, the voxel query is embedded into the 3D voxel space through the voxel aggregation module. These embeddings interact with the pixel-level features, and the deformable attention mechanism is used to complete the spatial diffusion of the features to form a complete 3D voxel feature.

[0032] In the joint reasoning stage, the results of the two branches, 2D semantic segmentation and 3D occupancy prediction, are combined. The prediction results of 2D semantic segmentation are used as auxiliary information to further improve the accuracy of 3D occupancy prediction. 3D occupancy prediction generates the geometric state and semantic category of each voxel through mask level classification, and forms the final 3D occupancy prediction result.

[0033] During the training process, we used the SemanticKITTI dataset as the experimental benchmark. The input data was preprocessed using a specific data augmentation strategy, including random cropping, flipping, and color jittering, to enhance the model's generalization ability for different scenarios. The AdamW optimizer was used during the optimization process, and the initial learning rate was set to 2×10 -4 The training cycle is 30 epochs. The warm-up strategy is adopted throughout the training, and the learning rate increases linearly in the early stage and then decreases gradually.

[0034] During the training and inference phases, the network’s input image size is 1220×310, and the final output 3D occupancy prediction result is represented by the state of each voxel and its semantic category.

Claims

1. A space occupancy estimation method based on 2D-3D joint training, characterized in that: Here are the steps: Step 1: Image 2D pixel feature extraction The RGB image is fed into the backbone network ResNet, which is an image feature extractor composed of several residual blocks. After large-scale pre-training on ImageNet, feature maps at multiple downsampling levels are extracted. The feature maps are fed into the pixel extraction module together with the initialized cross-modal query for 2D feature extraction. Among them, the cross-modal query is a set of learnable, randomly initialized tensor arrays. The cross-modal query is randomly initialized as a learnable parameter with a dimension of (n, c), where n represents the number of cross-modal queries and c represents the number of channels of each cross-modal query. The pixel extraction module is mainly composed of a Transformer Decoder module and a linear layer. The Transformer Decoder module consists of 6 stacked Transformer blocks, and the linear layer consists of several fully connected layers and BN layers. Specifically, the cross-modal query and the feature map output by the backbone network are input into the Transformer Decoder module to interact through the cross-attention operation. The output results of the Transformer Decoder module are respectively sent to two linear layers for channel dimension transformation, and finally generate category prediction and mask labeling; at the same time, the feature pyramid method is used to perform multi-scale feature fusion on the feature map output by the backbone network to generate a 2D pixel embedding of size (c, h, w), where h and w are set to one eighth of H and W; the final input original RGB image is transformed into a size of (n, n class ) category prediction, mask tags of size (n,c) and 2D pixel embeddings of size (c,h,w), where n class is the number of categories in 3D space occupancy prediction; Step 2: 3D polymerization; The 3D aggregation module consists of depth-point cloud-voxel projection, deformable cross attention and deformable self-attention. The 3D aggregation module inputs a depth map of size (H, W) and feature maps of multiple down-sampling levels generated in step 1, and promotes the feature maps of multiple down-sampling levels obtained in step 1 to 3D space through interaction. Specifically, the depth map is first projected into the 3D space into a point cloud using the pre-known camera parameters, and a voxel space is defined with a size of (h 3d ,w 3d ,z 3d ), respectively representing the length, width and height of the voxel in 3D space, h 3d ,w 3d is one eighth of the input image, z 3d is 16; then the occupied grids in the voxel space are set to an occupied state according to the projected point cloud; then the marked voxel space is initialized to a size of (c, h 3d ,w 3d ,z 3d ), where the occupied grid fill length is c, and the size is (c,h 3d ,w 3d ,z 3d ), and then send the obtained 3D voxel query and the feature map obtained in step 1 to the deformable cross attention defined in Deformable DETR for interaction, and then perform the deformable self-attention used in GroundingDino; fill the unoccupied areas in the 3D voxel query, and finally get a size of (c, h 3d ,w 3d ,z 3d )’s 3D voxel embedding; Step 3: Joint reasoning; The joint reasoning module does not contain any learnable parts and consists of several matrix multiplications. It accepts the 2D pixel embedding, mask label, 3D voxel embedding and category prediction generated in step 1 and step 2 to generate the final 2D semantic segmentation map and 3D voxel prediction map. Specifically, the 2D pixel embedding is first multiplied with the mask label to obtain a 2D mask prediction of size (n,h,w), which is then multiplied with the category prediction to obtain a 2D mask prediction of size (n class ,h,w) 2D semantic segmentation result; at the same time, the 3D voxel embedding is multiplied with the mask label to obtain a size of (n,h 3d ,w 3d ,z 3d ) is multiplied by the category prediction to obtain a 3D mask prediction of size (n class ,h 3d ,w 3d ,z 3d )’s 3D space occupancy prediction result; after this inference, the results of 2D semantic segmentation and three-dimensional space prediction are output at the same time.

Citation Information

Cited By

  • Semantic occupancy prediction method and system based on hierarchical semantic supervision and collaborative loss

    CN122200256A