A high-dimension-high-density three-dimensional semantic scene completion method
By employing multi-view image processing and cross-attention mechanisms, a high-dimensional, high-density 3D semantic scene completion method is generated, which solves the problem of dimensionality and density differences in autonomous driving road scenes and achieves fine-grained voxel semantic separation under a 3D stereoscopic perspective.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU GONGSHU DISTRICT HOLOGRAPHIC INTELLIGENT TECHNOLOGY RESEARCH INSTITUTE
- Filing Date
- 2026-04-30
- Publication Date
- 2026-05-29
Smart Images

Figure CN122116365A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of scene perception technology, and in particular relates to a high-dimensional and high-density three-dimensional semantic scene completion method. Background Technology
[0002] In recent years, with the development of autonomous driving and intelligent space perception technologies, 3D semantic scene completion has become a key task for understanding and perceiving complex environments. Its goal is to reconstruct a complete 3D semantic scene from multi-view image inputs and predict the geometric occupancy state and semantic category of each voxel in space. It has important research and application value in scenarios such as autonomous driving and robot navigation.
[0003] Existing 3D semantic scene completion methods often ignore the dimensional and density differences in autonomous driving road scenes. (1) Dimensional difference: The input image data is captured from a 2D plane perspective, resulting in coarse pixel semantics that are confused by object occlusion; however, the semantic scene completion task requires outputting fine-grained voxel semantics from a 3D stereo perspective to accurately separate occluded objects in the input image. (2) Density difference: Manual annotation based on LiDAR sensors results in sparse labels with gaps due to the limited resolution of point clouds; however, real road scenes present dense and coherent spatial occupancy, with a semantic information density much higher than that of manual annotation.
[0004] Patent document CN120543989A discloses a method, device, and medium for 3D semantic scene completion based on geometric and temporal modeling and hierarchical context alignment. It acquires depth features and contextual features, and uses a depth confidence-aware cross-attention mechanism to supplement information in low-depth confidence regions using contextual features to generate current frame-related features. It also acquires historical frame-related features from a pose network, calculates cross-frame feature affinity, and dynamically optimizes historical frame-related features to obtain multi-layered historical frame aggregated features. In a unified space, it uses the depth assumption of temporal feature voxels as the distance axis and projects volumetric feature voxels onto the unified space, performing global alignment and combination of geometric and temporal feature voxels to obtain the final aggregated features.
[0005] Patent document CN119379562A discloses a method, apparatus, storage medium, and computer device for three-dimensional semantic scene completion. The method includes inputting a two-dimensional image and a depth image into a three-dimensional semantic scene completion model; extracting features through a feature extraction module to extract two-dimensional semantic features from the two-dimensional image and depth features from the depth image; fusing the two-dimensional semantic features and depth features to obtain a fused feature map; inputting the fused feature map into a feature enhancement module for feature enhancement to obtain an enhanced fused feature map; inputting the enhanced fused feature map into a downsampling module for downsampling to obtain a downsampled feature map; inputting the downsampled feature map into an attention enhancement module for attention enhancement processing; and outputting a three-dimensional image of the scene to be completed. Summary of the Invention
[0006] The purpose of this invention is to provide a high-dimensional and high-density three-dimensional semantic scene completion method, which can solve the problems of dimensional and density differences in complex road scenes.
[0007] To achieve the objectives of this invention, the following technical solution is provided: a high-dimensional, high-density three-dimensional semantic scene completion method, comprising the following steps: Acquire multi-view images and extract multi-scale representations from the multi-view images. Aggregate the multi-scale representations to obtain the corresponding two-dimensional image features. The two-dimensional image features are subjected to pseudo-semantic expansion and random noise is added to generate corresponding pseudo-voxelated features. Pixel query is introduced, and global semantic extraction is performed on pseudo-voxelated features through cross-attention mechanism. Semantic clustering is then performed based on the extracted global semantics to generate multiple semantic feature clusters. Localization is performed based on the similarity between pseudo-voxalization features and multiple semantic feature clusters to generate corresponding aggregated semantic features; For multi-view images, view transformation mapping is performed, transforming similar queries into voxel queries and aggregated semantic features into 3D voxel features; Based on voxel query and 3D voxel features, corresponding voxel features are generated. The generated voxel features are input into a heuristic binary classification head to detect occupied voxels, and the voxel with the highest score is taken as the geometric key voxel. Based on voxel queries and voxel features, perform category-by-category predictions to generate initial semantic scene completion results; The semantic key voxels are obtained by filtering based on the classification confidence of the initial semantic scene completion results; Data augmentation is performed based on geometric and semantic key voxels to obtain augmented key voxels. The augmented geometric key voxels are then used to optimize the initial semantic scene completion result to obtain the final semantic scene completion result.
[0008] This invention expands planar image features into pseudo-voxelated features and aligns them with 3D scene features through high-dimensional semantic aggregation; then, it enhances the discriminative information density of the 3D scene representation through high-density geometric optimization, completes and corrects erroneous voxel predictions, and achieves accurate semantic scene completion.
[0009] Specifically, the two-dimensional image features are extracted using ResNet50 for multi-scale representation and then aggregated through an FPN network.
[0010] Specifically, the expression for the pseudo-voxelation feature is as follows: ; in, This represents the expanded voxelized features. This represents a two-dimensional convolutional network layer used for dimensional expansion. This represents the number of channels in the extended dimension. This represents random noise.
[0011] Specifically, the expression for semantic clustering is as follows: ; in, Represents the corresponding extended dimension Clusters of semantic features Indicates the first Cluster centers of each cluster. Indicates a cross-attention layer. This refers to a semantic clustering algorithm.
[0012] Specifically, the expression for the aggregated semantic features is as follows: ; in, This represents the final aggregate semantic features. The first voxelization feature extension Each feature channel The function that calculates feature similarity scores.
[0013] Specifically, the data augmentation expression is as follows: ; in, Representing geometric key voxels, Represents semantic key voxels, This represents the distribution consistency loss based on KL divergence.
[0014] Specifically, the expression for the semantic scene completion result is as follows: ; in, This represents the optimized semantic scene completion result. This indicates a feature concatenation operation. Indicates a fully connected layer. This indicates the initial semantic scene completion result.
[0015] Specifically, the expression for the initial semantic scene completion result is as follows: ; in, This represents the initial semantic scene completion result. This indicates a category header. Indicates a voxel query. Represents 3D voxel features.
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: By extending pixel semantics to fine-grained voxel semantic features, the alignment of pixel and voxel semantics is achieved, alleviating the problem of dimensional differences. At the same time, by identifying and aligning geometric and semantic key voxels, the density of local discriminative information is enhanced, further alleviating the problem of density differences, thereby achieving accurate semantic scene completion. Attached Figure Description
[0017] Figure 1 This is a flowchart of a high-dimensional, high-density three-dimensional semantic scene completion method provided in this embodiment. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0019] like Figure 1As shown, this embodiment provides a high-dimensional, high-density 3D semantic scene completion method, which includes the following steps: Acquire multi-view images and extract multi-scale representations from the multi-view images. Aggregate the multi-scale representations to obtain the corresponding two-dimensional image features. The two-dimensional image features are subjected to pseudo-semantic expansion and random noise is added to generate corresponding pseudo-voxelated features. Pixel query is introduced, and global semantic extraction is performed on pseudo-voxelated features through cross-attention mechanism. Semantic clustering is then performed based on the extracted global semantics to generate multiple semantic feature clusters. Localization is performed based on the similarity between pseudo-voxalization features and multiple semantic feature clusters to generate corresponding aggregated semantic features; For multi-view images, view transformation mapping is performed, transforming similar queries into voxel queries and aggregated semantic features into 3D voxel features; Based on voxel query and 3D voxel features, corresponding voxel features are generated. The generated voxel features are input into a heuristic binary classification head to detect occupied voxels, and the voxel with the highest score is taken as the geometric key voxel. Based on voxel queries and voxel features, perform category-by-category predictions to generate initial semantic scene completion results; The semantic key voxels are obtained by filtering based on the classification confidence of the initial semantic scene completion results; Data augmentation is performed based on geometric and semantic key voxels to obtain augmented key voxels. The augmented geometric key voxels are then used to optimize the initial semantic scene completion result to obtain the final semantic scene completion result.
[0020] More specifically, data preparation and feature extraction: For multi-view image input, ResNet50 is used to extract multi-scale representations, and the multi-scale representations are aggregated through an FPN network to obtain two-dimensional image features. .
[0021] High-dimensional semantic aggregation: First, based on the two-dimensional image features obtained in step (1), a dimension expansion layer is used to expand them into voxelized features along a pseudo-"semantic dimension," and random noise is added. This process can be expressed as the following formula: ; in, This represents the expanded voxelized features. This represents a two-dimensional convolutional network layer used for dimensional expansion. This represents the number of channels in the extended dimension. This represents random noise.
[0022] Then, pixel query is introduced. Global semantics are extracted from pseudo-voxelated features through a cross-attention mechanism, and semantic clustering is then performed. ; in, Represents the corresponding extended dimension Clusters of semantic features Indicates the first Cluster centers of each cluster. Indicates a cross-attention layer. This represents a semantic clustering algorithm. Finally, the identifiable regions are located based on the similarity between the pseudo-voxelated features and each semantic cluster, achieving the aggregation of global high-dimensional fine-grained semantic features. ; in, This represents the final aggregate semantic features. The first voxelization feature extension Each feature channel The function that calculates feature similarity scores.
[0023] Feature perspective shift: Pixel queries are performed through viewpoint transformation mapping. Mapping to obtain voxel query The aggregated semantic features are mapped to obtain 3D voxel features. .
[0024] High-density geometry optimization: A coarse-to-fine detection architecture is adopted. First, the voxel queries obtained from pixel query mapping and the voxel features obtained through viewpoint transformation are fed into a heuristic binary classification head to detect occupied voxels, and the voxel with the highest score is designated as the geometric key voxel. ; in, Representing geometric key voxels, Represents semantic key voxels, This represents the distribution consistency loss based on KL divergence.
[0025] In the optimization phase, voxel queries and voxel features are first used to perform category-by-category predictions to generate initial semantic scene completion results: ; in, This represents the initial semantic scene completion result. This represents the category-by-category classification header. Then, semantic key voxels are selected based on the classification confidence of the initial semantic scene completion results: ; in, This represents semantic key voxels. Furthermore, a key loss is employed to align the overall distribution of geometric and semantic key voxels, enhancing the density of context-discriminative information. ; in, This represents the distribution consistency loss based on KL divergence. Finally, the initial semantic scene completion result is optimized using aligned key voxels: ; in, This represents the optimized semantic scene completion result; Indicates a fully connected layer. This indicates a feature splicing operation.
[0026] To better illustrate the technical effects of the method provided in this embodiment, the following experimental testing process is provided.
[0027] This embodiment is based on experiments using the SemanticKITTI autonomous driving dataset, proposed in the paper "SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDARSequences" (authors Jens Behley, Martin Garbade, Andres Milioto, Jan Quenzel, SvenBehnke, Cyrill Stachniss, Juergen Gall, published in 2019), which includes target objects of 20 different semantic categories. We compare this invention with the following three existing 3D semantic scene completion methods in experiments: Existing Method 1: The method in the paper "Not All Voxels Are Equal: Hardness-Aware SemanticScene Completion with Self-Distillation" (authors Song Wang, Jiawei Yu, Wentong Li, Wenyu Liu, Xiaolu Liu, Junbo Chen, Jianke Zhu, published at the 2024 IEEE Conference on Computer Vision and Pattern Recognition) mines the local and global hardness scores of each voxel and proposes a self-distillation strategy for a more stable and consistent training process.
[0028] The second existing method is the one in the paper "Camera-based 3d Semantic Scene Completion with SparseGuidance Network" (authors Jianbiao Mei, Yu Yang, Mengmeng Wang, Junyu Zhu, Jongwon Ra, Yukai Ma, Laijian Li and Yong Liu, published in IEEE Transactions on Image Processing in 2024). This method designs a dense-sparse-dense architecture to enhance the clarity of object boundaries in the semantic scene completion results.
[0029] The third existing method is the one in the paper "Context and Geometry Aware Voxel Transformer for Semantic Scene Completion" (authors Zhu Yu, Runmin Zhang, Jiacheng Ying, Junchen Yu, Xiaohai Hu, Lun Luo, Si-Yuan Cao and Hui-Liang Shen, published at the 2024 Conference on Neural Information Processing Systems). This method introduces context-dependent semantic queries and multiple 3D scene representations to enhance the performance of semantic scene completion.
[0030] This embodiment uses the Intersection over Union (IoU) and Mean Intersection over Union (mIoU) metrics to measure the predictive performance of different semantic scene completion methods. Specifically, the IoU metric is used to evaluate the accuracy of the model in semantic scene completion in 3D space. It is calculated as the ratio of the intersection to the union of the voxels predicted by the model as occupied and the actual occupied voxels. The mIoU metric measures the classification prediction performance of the model across all semantic categories. The IoU is calculated for each category and then averaged. The larger the values of IoU and mIoU, the better the performance of the semantic scene completion method.
[0031]
[0032] As can be seen from Table 1, the present invention achieves better 3D semantic scene completion results. Existing semantic scene completion methods ignore the dimensional and density differences in the semantic scene completion task; while the high-dimensional and high-density 3D semantic scene completion method of the present invention can align pixel and voxel semantic features, enhance local discriminative information density, effectively alleviate the problems of dimensional and density differences, and thus generate more accurate and realistic semantic scene completion results.
[0033] Furthermore, the terms "upper," "lower," "inner," "outer," "front," and "rear" are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Unless otherwise specifically stated, the relative steps, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the invention.
[0034] Of course, the above description is only a specific embodiment of the present invention and is not intended to limit the scope of the present invention. All equivalent changes or modifications made to the structure, features and principles described in the claims of the present invention should be included in the scope of the claims of the present invention.
[0035] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A high-dimensional, high-density three-dimensional semantic scene completion method, characterized in that, Includes the following steps: Acquire multi-view images and extract multi-scale representations from the multi-view images. Aggregate the multi-scale representations to obtain the corresponding two-dimensional image features. The two-dimensional image features are subjected to pseudo-semantic expansion and random noise is added to generate corresponding pseudo-voxelated features. Pixel query is introduced, and global semantic extraction is performed on pseudo-voxelated features through cross-attention mechanism. Semantic clustering is then performed based on the extracted global semantics to generate multiple semantic feature clusters. Localization is performed based on the similarity between pseudo-voxalization features and multiple semantic feature clusters to generate corresponding aggregated semantic features; For multi-view images, view transformation mapping is performed, transforming similar queries into voxel queries and aggregated semantic features into 3D voxel features; Based on voxel query and 3D voxel features, corresponding voxel features are generated. The generated voxel features are input into a heuristic binary classification head to detect occupied voxels, and the voxel with the highest score is taken as the geometric key voxel. Based on voxel queries and voxel features, perform category-by-category predictions to generate initial semantic scene completion results; The semantic key voxels are obtained by filtering based on the classification confidence of the initial semantic scene completion results; Data augmentation is performed based on geometric and semantic key voxels to obtain augmented key voxels. The augmented geometric key voxels are then used to optimize the initial semantic scene completion result to obtain the final semantic scene completion result.
2. The high-dimensional, high-density three-dimensional semantic scene completion method according to claim 1, characterized in that, The two-dimensional image features are extracted using ResNet50 for multi-scale representation and then aggregated through an FPN network.
3. The high-dimensional, high-density three-dimensional semantic scene completion method according to claim 1, characterized in that, The expression for the pseudo-voxometry feature is as follows: ;in, This represents the expanded voxelized features. This represents a two-dimensional convolutional network layer used for dimensional expansion. This represents the number of channels in the extended dimension. This represents random noise.
4. The high-dimensional, high-density three-dimensional semantic scene completion method according to claim 1, characterized in that, The expression for semantic clustering is as follows: ;in, Represents the corresponding extended dimension Clusters of semantic features Indicates the first Cluster centers of each cluster. Indicates a cross-attention layer. This refers to a semantic clustering algorithm.
5. The high-dimensional, high-density three-dimensional semantic scene completion method according to claim 1, characterized in that, The expression for the aggregated semantic feature is as follows: ;in, This represents the final aggregate semantic features. The first voxelization feature extension Each feature channel The function that calculates feature similarity scores.
6. The high-dimensional, high-density three-dimensional semantic scene completion method according to claim 1, characterized in that, The data augmentation expression is as follows: ;in, Representing geometric key voxels, Represents semantic key voxels, This represents the distribution consistency loss based on KL divergence.
7. The high-dimensional, high-density three-dimensional semantic scene completion method according to claim 1, characterized in that, The expression for the semantic scene completion result is as follows: ;in, This represents the optimized semantic scene completion result. This indicates a feature concatenation operation. Indicates a fully connected layer. This indicates the initial semantic scene completion result.
8. The high-dimensional, high-density three-dimensional semantic scene completion method according to claim 1 or 7, characterized in that, The expression for the initial semantic scene completion result is as follows: ;in, This represents the initial semantic scene completion result. This indicates a category header. Indicates a voxel query. Represents 3D voxel features.
Citation Information
Patent Citations
Three-dimensional semantic scene completion method and device, storage medium and computer equipment
CN119379562A
Three-dimensional semantic scene completion method and device based on geometric and time sequence modeling hierarchical context alignment and medium
CN120543989A
Three-dimensional semantic scene completion method and system based on semantic-physical joint representation
CN121616841A
Three-dimensional semantic scene completion method based on camera enhancement, medium and equipment
CN121837560A