Point cloud semantic scene completion system
By designing a point cloud semantic scene completion system, using the improved Cylinder3D encoder and 2D convolutional network to extract multi-scale features and fusion, the problems of poor scene completion integrity and low semantic segmentation accuracy in the existing technology are solved, and more accurate scene reconstruction is achieved.
Patent Information
- Application Number
- CN202510559486.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-30
AI Technical Summary
The existing point cloud semantic scene completion network has problems such as poor scene completion integrity and low semantic segmentation accuracy.
A point cloud semantic scene completion system is designed, including a voxelization processing module, a multi-scale voxel feature extraction module and a bird's-eye view feature processing module. Through the improved Cylinder3D encoder and 2D convolutional network, multi-scale voxel features and bird's-eye view features are extracted and fused for more accurate scene reconstruction.
It realizes more accurate and complete scene reconstruction, improves semantic segmentation accuracy, and performs best in the evaluation on the SemanticKITTI dataset, solving the problems of poor scene completion integrity and low semantic segmentation accuracy.
Smart Images

Figure CN120070902A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a point cloud semantic scene completion system. Background Art
[0002] LiDAR technology has changed the way of perceiving three-dimensional environments by providing detailed geometric data, and is widely used in fields such as robotics, autonomous driving, urban planning, and environmental monitoring. However, although LiDAR point clouds provide rich geometric information, their incompleteness and sparsity limit their effectiveness in practical applications.
[0003] To address these challenges, semantic scene completion (SSC) of point clouds aims to predict the occupancy of all voxels in three-dimensional space and their semantic labels from partial point clouds, thereby not only inferring spatial occupancy but also assigning specific semantic meanings to each voxel (such as buildings, trees, roads, etc.).
[0004] However, existing point cloud semantic scene completion networks still have problems such as poor scene completion integrity and low semantic segmentation accuracy. Summary of the Invention
[0005] In view of this, the present invention provides a point cloud semantic scene completion system to solve the problems of poor scene completion integrity and low semantic segmentation accuracy.
[0006] A point cloud semantic scene completion system includes a voxelization processing module, a multi-scale voxel feature extraction module, and a bird's-eye view feature processing module; The voxelization processing module takes a single-frame LiDAR point cloud as input, discretizes the single-frame LiDAR point cloud to obtain a structured 3D voxel grid, then performs feature extraction operations on each voxel in the structured 3D voxel grid to obtain point-level features, maps the point-level features through a multi-layer perceptron, and then aggregates them using a max pooling operation to obtain initial voxel features; The multi-scale voxel feature extraction module processes the initial voxel features to obtain multi-scale voxel features. The multi-scale voxel feature extraction module uses an improved Cylinder3D encoder, and the improved Cylinder3D encoder includes an asymmetric residual block and an asymmetric downsampling block; The bird's-eye view feature processing module includes an encoder and a decoder. The encoder processes the initial voxel features to extract multi-scale bird's-eye view features, fuses the multi-scale bird's-eye view features with the multi-scale voxel features, and sends the fused features to the decoder. The decoder finally outputs the prediction result of semantic scene completion based on skip connections, feature stitching, and content-aware feature recombination.
[0007] According to the point cloud semantic scene completion system provided by the present invention, the voxelization processing module discretizes a single-frame lidar point cloud to obtain initial voxel features. The initial voxel features are respectively sent to the multi-scale voxel feature extraction module and the bird's-eye view feature processing module. The multi-scale voxel feature extraction module uses an improved Cylinder3D encoder, which focuses on extracting fine-grained 3D spatial features. The bird's-eye view feature processing module projects voxel data onto a two-dimensional plane and effectively captures rich context information through a 2D convolutional network of an encoder and a decoder. By fusing multi-scale bird's-eye view features with multi-scale voxel features in the encoder, the advantages of the two features can be complementary, enabling the system to take into account both detailed geometric information and global context, and thus achieving more accurate and complete scene reconstruction, solving the problems of poor scene completion integrity and low semantic segmentation accuracy. The results on the evaluation website of the public dataset SemanticKITTI show that the system of the present invention performs best in the scene completion task compared with other existing algorithms. Brief Description of the Drawings
[0008] Figure 1 is a structural block diagram of the point cloud semantic scene completion system provided by an embodiment of the present invention; Figure 2 is a qualitative visualization result comparison diagram. Detailed Embodiment
[0009] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the drawings are exemplary and are intended to explain the embodiments of the present invention and should not be construed as limiting the present invention.
[0010] Please refer to Figure 1 , an embodiment of the present invention provides a point cloud semantic scene completion system, including a voxelization processing module, a multi-scale voxel feature extraction module, and a bird's-eye view feature processing module.
[0011] The voxelization processing module takes a single-frame lidar point cloud as input, discretizes the single-frame lidar point cloud to obtain a structured 3D voxel grid, then performs feature extraction operations on each voxel in the structured 3D voxel grid to obtain point-level features, maps the point-level features through a multi-layer perceptron, and then performs aggregation using a max pooling operation to obtain initial voxel features.
[0012] In the voxelization processing module, first, the points located within a predefined range are filtered out from the input single-frame lidar point cloud to obtain a filtered point cloud set , , Denote the th point cloud, where represents the number of point clouds, and based on the point cloud set obtain a structured 3D voxel grid , , Denote the th voxel, , , are respectively the numbers of voxels in the three dimensions of length, width, and height corresponding to the voxelized point cloud in 3D space. The voxel is represented by its centroid coordinate . To ensure the structuring of the scene representation, any point in the point cloud is assigned to the voxel closest to it.
[0013] After voxelization, perform a feature extraction operation on each voxel in the structured 3D voxel grid to obtain each point-level feature. The point-level feature can capture the spatial and reflection characteristics of the point. The th point-level feature has the following expression: ; where, represents the offset between and the centroid of the corresponding voxel, represents the absolute spatial coordinate of, is the reflection intensity of; Then map the point-level features through a multi-layer perceptron (MLP) for feature mapping. To enhance the feature representation, the multi-layer perceptron maps the 7-dimensional point features to a high-dimensional space through successive non-linear transformations. In this embodiment, the multi-layer perceptron consists of stacked fully connected layers, batch normalization layers, and ReLU activation functions, and can extract richer features; then perform a max pooling operation for aggregation to obtain the initial voxel feature. This operation can not only capture the most critical spatial and semantic information but also effectively reduce data redundancy and computational complexity. The expression of the initial voxel feature is: ; ; where, represents the initial voxel feature, represents the aggregated feature of the th voxel, is a linear layer with a ReLU activation function for further compressing the aggregated high-dimensional features; represents the max pooling operation, represents the denote the th point cloud.
[0014] Each voxel is encoded as a D -dimensional feature vector. For the entire scene, the size of the initial voxel features is , which constitutes a structured and semantically rich scene representation and serves as the input for subsequent tasks.
[0015] The multi-scale voxel feature extraction module processes the initial voxel features to obtain multi-scale voxel features. The multi-scale voxel feature extraction module uses an improved Cylinder3D encoder, and the improved Cylinder3D encoder includes an asymmetric residual block and an asymmetric downsampling block.
[0016] Among them, the multi-scale voxel feature extraction module uses an improved Cylinder3D encoder to process the initial voxel features , capturing detailed spatial information at different scales within the 3D volume. The multi-scale voxel feature extraction module effectively encodes the local and global context from the voxelized point cloud data.
[0017] In the multi-scale voxel feature extraction module, the asymmetric residual block processes the initial voxel features through two different branches : One branch applies consecutive 3D submanifold sparse convolutions with kernel sizes of 1×3×3 and 3×1×3 in sequence; the other branch applies consecutive 3D submanifold sparse convolutions with kernel sizes of 3×1×3 and 1×3×3 in sequence; finally, the output tensors from the two branches are added element-wise; this combination of two different branches reduces the computational amount compared to using a 3×3×3 convolution and also significantly enhances the expressive power of the convolutional kernels in the horizontal and vertical directions, facilitating the extraction of features of cuboid-shaped objects in the 3D point cloud scene. This fits well with the point cloud distributions of some objects in the autonomous driving scenario, such as cars, trucks, bicycles, motorcycles, etc., which often approximate the cuboid shape.
[0018] The result of the addition is input to the asymmetric downsampling block. The asymmetric downsampling block is a 3D sparse convolutional layer with a stride of 2 set after the asymmetric residual block. The asymmetric downsampling block applies 2-fold downsampling three times, and the multi-scale voxel features are obtained after being processed by the asymmetric downsampling block , , where represents the set of real numbers, denotes the number of channels of the th scale voxel features, = 1, 2, 3. The multi-scale voxel features effectively represent the hierarchical structure of the 3D scene.
[0019] The Bird's Eye View (BEV) feature processing module includes an encoder and a decoder. The encoder processes the initial voxel features, extracts multi-scale BEV features, fuses the multi-scale BEV features with the multi-scale voxel features, and sends the fused features to the decoder. The decoder, based on skip connections, feature concatenation, and content-aware feature recombination, finally outputs the prediction results for semantic scene completion.
[0020] Specifically, the BEV feature processing module projects the 3D voxel grid onto the 2D xy plane and then aggregates voxel features within each 2D grid cell using max pooling operations.
[0021] Specifically, in the encoder, the encoder first processes the initial voxel features by projection, ignoring the z coordinate of the voxels, thus generating a 2D grid with a resolution of . Then, feature fusion is performed through max pooling. The BEV feature of the th 2D grid cell is expressed as: ; where is a processing unit composed of a linear layer and a ReLU activation function, represents the th voxel, represents the aggregated feature of the th voxel; Furthermore, the initial BEV feature is obtained. Each 2D grid cell is encoded as a D -dimensional vector.
[0022] Then, the encoder performs 4 times of 2x downsampling operations on the initial BEV feature in sequence to extract multi-scale BEV features represents the number of channels of the th scale voxel feature, = 1, 2, 3, 4; where the initial BEV feature is first downsampled for the first time to obtain the BEV feature after the first downsampling. Then, is fused with the th scale voxel feature. After fusion, it is input into the downsampling module of the encoder. After convolution operations, it is downsampled for the second time to obtain the BEV feature after the second downsampling. Then, With the second-scale voxel feature Fuse them, and after the fusion, input them into the downsampling module of the encoder. After convolution operations, perform the third downsampling to obtain the bird's-eye view feature after the third downsampling , and then Fuse it with the third-scale voxel feature Fuse them, and after the fusion, input them into the downsampling module of the encoder. After convolution operations, perform the fourth downsampling to finally obtain the fused feature.
[0023] Through three times of fusion, Stack them along the z-axis and reduce the number of channels through convolution to match 's channel dimension. By fusing the multi-scale bird's-eye view features and the multi-scale voxel features, the 3D scene representation is enhanced, and the complementary advantages of voxels and BEV views are fully utilized, thus achieving a more comprehensive scene understanding.
[0024] In this embodiment, an Atrous Spatial Pyramid Pooling (ASPP) module is provided between the encoder and the decoder. In the Atrous Spatial Pyramid Pooling module, considering the size of the input features, 1, 2, and 4 are selected as the dilation rate combination to perform dilated convolution. This setting enables each convolutional kernel to operate efficiently and avoids some convolutional kernel elements exceeding the feature range of the feature map due to too high dilation rate. The Atrous Spatial Pyramid Pooling module captures multi-scale context information by parallelly using convolutional kernels with different dilation rates while maintaining the input resolution unchanged. With the advantage of the Atrous Spatial Pyramid Pooling module, the system proposed in this application can effectively integrate and refine features at each scale before upsampling, significantly enhancing the feature expression ability.
[0025] Then, the decoder finally outputs the prediction result of semantic scene completion based on skip connections, feature concatenation, and content-aware feature recombination.
[0026] Specifically, in the decoder, a content-aware feature recombination module is used to upsample the fused features. The content-aware feature recombination module can process the features of the low-resolution feature map and construct a high-resolution output by generating an adaptive kernel and recombining the features.
[0027] For any position in the feature map feature of the final upsampled output, there is a corresponding position in the original input feature map feature is the upsampling magnification factor, and here double upsampling is adopted represents the floor operation. Use denotes the size of the upsampling convolution kernel, and the feature at each target position of the output feature is determined by the region centered at the source position and the predicted recombination kernel.
[0028] Specifically, the content-aware feature recombination module includes a recombination kernel prediction unit and a feature recombination unit.
[0029] The recombination kernel prediction unit is used to dynamically predict a recombination kernel according to the input fused feature and in combination with each position of the final output feature. The recombination kernel is a set of weight parameters used to assign corresponding weights to different regions during the upsampling process according to the input fused feature, so as to achieve content-aware feature recombination. Subsequently, these predicted recombination kernels are applied to the input features to achieve content-aware feature recombination.
[0030] The feature recombination unit is used to perform content-aware feature recombination based on the recombination kernel, and then obtain the prediction result of semantic scene completion; In the recombination kernel prediction unit, the fused feature first undergoes feature channel reduction through a 1×1 convolution kernel, and then passes through a convolution operation with a size of to generate a recombination kernel. The channel size of the recombination kernel is , is the upsampling magnification, is the size of the upsampling convolution kernel. Then, pixel-level rearrangement is performed on the recombination kernel. After rearrangement, the size of the recombination kernel is . Finally, the normalization exponential function is used to normalize the recombination kernel in the channel dimension; In the feature recombination unit, based on the upsampling weight , feature recombination is performed within the neighborhood of the target position. The expression for feature recombination is: ; where, denotes the output feature at the target position , denotes the radius of the local sampling neighborhood, , denotes the floor operation, denotes at the position value, denotes the feature matrix of the region centered at the source position corresponding to the target position, denotes at the position feature.
[0031] The total loss function of the above point cloud semantic scene completion system is: ; ; ; ; ; ; ; ; Among them, represents the weighted cross-entropy loss function, represents the Lovasz loss function, represents the focal loss function, represents the scene-category affinity loss function, is the true label corresponding to the class of the ith voxel, is the weight of class ; is the probability output value that the ith voxel is predicted to be class represents the total number of classes, represents a piecewise linear function with a global minimum, is the error vector of class ; is the balance factor of class ; is the focusing parameter, , and represent precision, recall, and specificity respectively; represents the binary value indicating whether the ith voxel belongs to class . If the ith voxel belongs to class , the value is 1; if the ith voxel does not belong to class , the value is 0.
[0032] The present invention constructs a total loss function specifically for semantic scene completion, which further improves the performance of point cloud semantic scene completion. Among them, the weighted cross-entropy loss function is the standard loss function in the classification task and is used to assist the model in predicting the correct label of each voxel. The Lovasz loss function solves the class imbalance problem in the semantic segmentation task by optimizing the intersection over union score, thereby improving the model's detection ability for foreground objects. The focal loss function Reduces the contribution of well-classified samples, enabling the model to focus more on samples that are difficult to classify and misclassified, thereby further improving the model's performance on challenging parts of the scene. Scene-class affinity loss function Extended from the binary affinity loss, it helps to enhance the spatial structure and semantic consistency of scene completion.
[0033] The system of the present invention is tested below.
[0034] Test dataset: SemanticKITTI is a large-scale outdoor lidar point cloud dataset extended from the KITTI visual odometry dataset, mainly used for research in the field of autonomous driving, especially in semantic segmentation and semantic scene completion. The dataset is obtained by scanning with a single Velodyne HDL-64E model lidar sensor in various complex scenes such as inner cities, residential areas, highways, and rural roads around Karlsruhe, Germany, and dense semantic annotations are made on the collected point cloud data. The entire dataset contains 22 sequences, among which sequences 00 - 10 are divided into the training set, sequences 11 - 21 are used as the test set, and sequence 08 in the training set is used as the validation set. The training set provides 23,201 complete 3D point cloud scans, while the test set provides 20,351 scans. It should be noted that only the training set provides label information, and the test set does not directly provide labels. Researchers need to submit the prediction results of the test set to the CodaLab competition platform, and the test set results are obtained through the evaluation service of this platform.
[0035] The dataset covers 28 categories, including categories that distinguish moving and non-moving objects. These categories are divided into types such as ground, buildings, transportation, nature, humans, objects, and others. Ground categories (such as roads, sidewalks, parking lots) occupy most of the point cloud in the dataset, followed by natural and transportation categories. In contrast, categories such as pedestrians, cyclists, and motorcyclists have fewer point clouds, showing a typical long-tail distribution. For the SSC task, 25 out of the 28 categories are used, excluding outlier, other structures, and other object categories because these categories are not considered during the evaluation process. In addition, moving categories are merged with their corresponding non-moving categories, resulting in a total of 19 evaluation categories.
[0036] Experimental settings: The valid range of the point cloud data is set to [0 to 51.2 meters, -25.6 to 25.6 meters, -2 to 4.4 meters]. The scene is divided into 256x256x32 voxel units, and each voxel unit represents a spatial resolution of 0.2m. In the data preprocessing stage, this application adopts a data augmentation strategy. Specifically, the augmentation strategy includes the following three types: mirror flipping along the x-axis, mirror flipping along the y-axis, and mirror flipping along both the x-axis and y-axis simultaneously. During the training process, one of the augmentation methods will be randomly selected to process the point cloud data to enhance the robustness of the model. Network training uses the PyTorch framework and is carried out on an NVIDIA GeForce RTX 3090 GPU with 24GB of video memory. The Adam optimizer is adopted. Except for the default parameters, the momentum term is set to 0.9, and the weight decay coefficient is set to 0.0001. The initial learning rate is set to 0.001 and decays to 98% of the original value after each epoch. The training process is carried out for 80 epochs. Due to the GPU video memory limitation, the batch size is set to 2 to ensure that there is no video memory overflow during the training process.
[0037] Comparative experiments: This application presents a detailed comparison of various methods on the official SemanticKITTI benchmark dataset. Among them, LMSCNet processes large-scale missing scenes through multi-scale 3D convolution and context aggregation modules, and uses voxelization expression to achieve indoor and outdoor scene completion; Local-DIFs divides the point cloud into neighborhoods and learns continuous implicit functions, and achieves high-precision detail reconstruction through the signed distance field (SDF); JS3C jointly optimizes the semantic segmentation and geometric completion tasks, and uses cross-task feature interaction to improve the semantic consistency of complex scenes; UDNet adopts a symmetric encoder-decoder structure, combines skip connections to achieve multi-level geometric feature fusion, and is good at processing noisy point clouds; SSA-SC introduces a non-local self-attention mechanism and restores a globally reasonable structure through long-range dependence modeling; SSC-RS designs a residual skip architecture for the sparse characteristics of lidar and progressively refines the missing areas. The comparison results are shown in Tables 1, 2, and 3. The experiments focus on the evaluation metrics of the intersection over union (IoU) for the scene completion task, the mean intersection over union (mIoU) for the semantic segmentation task, and the frames per second (FPS).
[0038] Table 1 Comparison of test results on the SemanticKITTI dataset for the completion task
[0039] Table 2 Comparison of test results on the SemanticKITTI dataset for the semantic segmentation task
[0040] Table 3 Comparison of Processing Speed Results on the SemanticKITTI Dataset
[0041] As can be seen from Table 1, for the scene completion task, the system of this application achieved an IoU of 62.6%, exceeding all baseline models. In particular, this application observed significant improvements compared to the better-performing methods such as SSC-RS (59.7%), UDNet (59.4%), and SSA-SC (58.8%). This demonstrates the effectiveness of this application in reconstructing a complete 3D scene.
[0042] As can be seen from Table 2, in the semantic segmentation task, measured by mIoU, this application also performed excellently, with an mIoU reaching 26.5%. This is significantly better than other methods, including SSC-RS (24.2%), SSA-SC (23.5%), and Local-DIFS (22.7%). It is worth noting that the system of this application achieved excellent segmentation performance in various categories, especially in challenging categories such as buildings, cars, trucks, and vegetation, exceeding existing methods in all cases.
[0043] As can be seen from Table 3, in terms of FPS, LMSCNet performed the best, reaching 21.3. However, its IoU and mIoU performance was poor. Compared with most algorithms, this application achieved a comparable FPS while ensuring accuracy, achieving a good balance between speed and accuracy. Figure 2 is a qualitative visualization result comparison, demonstrating excellent geometric reconstruction and semantic annotation capabilities in areas with complex scene structures, Figure 2 The circled parts in indicate the key comparison objects. From Figure 2 it can be seen that compared with SSA-SC and SSC-RS, the system of this application shows better details in geometric reconstruction and semantic segmentation, especially in areas with complex scene structures.
[0044] The above results verify the effectiveness of the system of this application in semantic scene completion, demonstrating its ability to capture local details and global context to enhance scene understanding.
[0045] In summary, for the point cloud semantic scene completion system according to the above embodiments, the voxelization processing module discretizes the single-frame lidar point cloud to obtain the initial voxel features. The initial voxel features are respectively sent to the multi-scale voxel feature extraction module and the bird's-eye view feature processing module. The multi-scale voxel feature extraction module uses an improved Cylinder3D encoder, which focuses on extracting fine-grained 3D spatial features. The bird's-eye view feature processing module projects the voxel data onto a two-dimensional plane and effectively captures rich context information through the 2D convolutional networks of the encoder and the decoder. By fusing the multi-scale bird's-eye view features with the multi-scale voxel features in the encoder, the advantages of the two types of features can be complementary, enabling the system to take into account both detailed geometric information and global context, and thus achieving more accurate and complete scene reconstruction, and solving the problems of poor scene completion integrity and low semantic segmentation accuracy. The results on the evaluation website of the public dataset SemanticKITTI show that the system of the present invention performs best in the scene completion task compared with other existing algorithms.
[0046] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.
Claims
1. A point cloud semantic scene completion system, characterized in that: It includes a voxel processing module, a multi-scale voxel feature extraction module and a bird's-eye view feature processing module; The voxel processing module takes a single-frame lidar point cloud as input, discretizes the single-frame lidar point cloud to obtain a structured 3D voxel grid, and then performs feature extraction operations on each voxel in the structured 3D voxel grid to obtain each point-level feature. The point-level features are feature mapped by a multi-layer perceptron, and then aggregated by a maximum pooling operation to obtain the initial voxel features. The multi-scale voxel feature extraction module processes the initial voxel feature to obtain a multi-scale voxel feature, wherein the multi-scale voxel feature extraction module adopts an improved Cylinder3D encoder, and the improved Cylinder3D encoder includes an asymmetric residual block and an asymmetric downsampling block; The bird's-eye view feature processing module includes an encoder and a decoder. The encoder processes the initial voxel features, extracts multi-scale bird's-eye view features, fuses the multi-scale bird's-eye view features with the multi-scale voxel features, and transmits the fused features to the decoder. The decoder reorganizes the features based on skip connections, feature splicing and content-aware features, and finally outputs the prediction results of semantic scene completion.
2. The point cloud semantic scene completion system according to claim 1, characterized in that: In the voxel processing module, points within a predefined range are first filtered out from the input single-frame lidar point cloud to obtain a filtered point cloud set. , , Indicates Point cloud, Indicates the number of point clouds, based on point cloud sets Get a structured 3D voxel grid , , Indicates Individual elements, , , They are the number of voxels in the length, width and height dimensions in 3D space after the point cloud is voxelized; Then, feature extraction is performed on each voxel in the structured 3D voxel grid to obtain each point-level feature. Point-level features The expression is: ; in, express The offset between the corresponding physical centroid, express The absolute space coordinates of for The reflection intensity; The point-level features are then mapped using a multi-layer perceptron, which consists of stacked fully connected layers, batch normalization layers, and ReLU activation functions. The maximum pooling operation is then used for aggregation to obtain the initial voxel features. The expression of the initial voxel features is: ; ; in, represents the initial voxel features, Indicates Aggregate features of individual voxels, is a linear layer with a ReLU activation function, represents the maximum pooling operation, Indicates Point-level features, Indicates A point cloud.
3. The point cloud semantic scene completion system according to claim 2, characterized in that: In the multi-scale voxel feature extraction module, the asymmetric residual block processes the initial voxel features through two different branches. :One branch applies continuous 3D submanifold sparse convolutions with kernel sizes of 1×3×3 and 3×1×3 respectively; the other branch applies continuous 3D submanifold sparse convolutions with kernel sizes of 3×1×3 and 1×3×3 respectively; finally, the output tensors from the two branches are added element by element; the result of the addition is input to the asymmetric downsampling block, which is a 3D sparse convolution layer with a stride of 2 set after the asymmetric residual block. The multi-scale voxel features are obtained after the asymmetric downsampling block. , ,in, represents the set of real numbers, Indicates The number of channels of voxel features at each scale, =1,2,3.
4. The point cloud semantic scene completion system according to claim 3, characterized in that: In the encoder, the encoder first performs the initial voxel features Projection, generating a resolution of 2D grid, and then perform feature fusion through maximum pooling. 2D grid cells Bird's-eye view features The expression is: ; in, It is a processing unit consisting of a linear layer and a ReLU activation function. Indicates Individual elements, Indicates Aggregate characteristics of individual voxels; Then we get the initial bird's-eye view features , ; The encoder then performs the initial bird's-eye view features Perform 4 times of 2-fold downsampling operations in sequence to extract multi-scale bird's-eye view features , , Indicates The number of channels of voxel features at each scale, =1,2,3,4; among them, the initial bird's-eye view features are first Perform the first downsampling to obtain the bird's-eye view features after the first downsampling , then With the first scale voxel feature Fusion: After fusion, it is input into the downsampling module of the encoder, and then after the convolution operation, it is downsampled for the second time to obtain the bird's-eye view features after the second downsampling. , and then and the second scale voxel feature Fusion: After fusion, it is input into the downsampling module of the encoder, and then after the convolution operation, it is downsampled for the third time to obtain the bird's-eye view features after the third downsampling. , and then and the third scale voxel features After fusion, the fusion is input into the downsampling module of the encoder, and then after the convolution operation, the fourth downsampling is performed to finally obtain the fused features.
5. The point cloud semantic scene completion system according to claim 1, characterized in that: A dilated spatial pyramid pooling module is provided between the encoder and the decoder. In the dilated spatial pyramid pooling module, 1, 2, and 4 are selected as dilation rate combinations to perform dilated convolution.
6. The point cloud semantic scene completion system according to claim 4, characterized in that: In the decoder, the fused features are upsampled using a content-aware feature reorganization module, which includes a reorganization kernel prediction unit and a feature reorganization unit; The reorganization kernel prediction unit is used to dynamically predict a reorganization kernel based on the input fused features and combined with each position of the final output features. The reorganization kernel is a set of weight parameters used to assign corresponding weights to different regions in the upsampling process based on the input fused features, thereby achieving content-aware feature reorganization. The feature recombination unit is used to perform content-aware feature recombination based on the recombination kernel, and then obtain the prediction result of semantic scene completion; In the recombinant kernel prediction unit, the fused features are first reduced by a convolution kernel of size 1×1, and then passed through a convolution kernel of size The convolution operation generates a recombinant kernel, and the channel size of the recombinant kernel is , is the upsampling factor, is the size of the up-sampling convolution kernel, and then the reorganized kernel is rearranged at the pixel level. After rearrangement, the size of the reorganized kernel is ,Finally, the normalized exponential function is used to normalize the recombined kernel in the channel dimension; In the feature recombination unit, based on the upsampling weights , feature reorganization is performed in the neighborhood of the target position, and the expression of feature reorganization is: ; in, Indicates the target location The output features of represents the radius of the local sampling neighborhood, , Indicates a round-down operation. express In Location The value of represents the feature matrix of the area centered at the source position corresponding to the target position, express In Location characteristics.
7. The point cloud semantic scene completion system according to claim 6, characterized in that: The total loss function of the system for: ; ; ; ; ; ; ; ; in, represents the weighted cross entropy loss function, represents the Lovaas loss function, represents the focal loss function, represents the scene-category affinity loss function, It is Voxel corresponding category The real label, Yes Category The weight of It is Voxels are predicted to be of class The probability output value of represents the total number of categories, represents a piecewise linear function with a global minimum, Yes Category The error vector, Yes Category The balance factor, is the focus parameter, , and They represent precision, recall, and specificity, respectively; Indicates Whether the voxel belongs to the class The binary value of Voxels belong to the category , then the value is 1, if the Voxel does not belong to class , the value is 0.
Citation Information
Patent Citations
Laser radar 3D real-time target detection method fusing multi-frame time sequence point cloud
CN111429514A
Point cloud 3D target detection method based on symmetric point generation
CN112598635A
Semantic scene completion method based on feature representation decomposition and aerial view fusion
CN116630975A
Implicit scene completion method based on dense pixel representation
CN118135519A
Cylinder-voxel fusion-based airborne hyperspectral point cloud three-dimensional target detection system and method
CN119516170A
Cited By
Scene completion method and device, electronic equipment and storage medium
CN120726243A
Scene completion methods, devices, electronic devices and storage media
CN120726243B
Indoor semantic scene completion method and system based on aerial view layered interactive perception
CN121526928A