A point cloud semantic scene completion system

Through the voxelization processing module and multi-scale feature fusion technology, the problems of scene completion integrity and semantic segmentation accuracy in point cloud semantic scene completion are solved, and more accurate and complete scene reconstruction is achieved.

CN120070902BActive Publication Date: 2025-07-08EAST CHINA JIAOTONG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510559486.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-07-08
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The existing point cloud semantic scene completion network has problems such as poor scene completion integrity and low semantic segmentation accuracy.

Method used

The voxelization processing module, multi-scale voxel feature extraction module and bird's-eye view feature processing module are adopted to realize multi-scale feature fusion and context information capture of point clouds through improved Cylinder3D encoder and 2D convolutional network, combining jump connection, feature stitching and content-aware feature recombination.

Benefits of technology

It improves the completeness of scene completion and the accuracy of semantic segmentation, significantly improving the accuracy of scene reconstruction and detail capture capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070902B_ABST
    Figure CN120070902B_ABST
Patent Text Reader

Abstract

The present invention provides a point cloud semantic scene completion system, including a voxelization processing module, a multi-scale voxel feature extraction module, and a bird's-eye view feature processing module; the voxelization processing module takes a single-frame lidar point cloud as input to obtain initial voxel features; the multi-scale voxel feature extraction module processes the initial voxel features to obtain multi-scale voxel features; the bird's-eye view feature processing module includes an encoder and a decoder. The encoder processes the initial voxel features to extract multi-scale bird's-eye view features, fuses the multi-scale bird's-eye view features with the multi-scale voxel features, and sends the fused features to the decoder. The decoder finally outputs the prediction result of semantic scene completion based on skip connections, feature splicing, and content-aware feature recombination. The present invention can solve the problems of poor scene completion integrity and low semantic segmentation accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a point cloud semantic scene completion system. Background Art

[0002] LiDAR technology has changed the way of perceiving three-dimensional environments by providing detailed geometric data, and is widely used in fields such as robotics, autonomous driving, urban planning, and environmental monitoring. However, although LiDAR point clouds provide rich geometric information, their incompleteness and sparsity limit their effectiveness in practical applications.

[0003] To address these challenges, semantic scene completion (SSC) of point clouds aims to predict the occupancy of all voxels in three-dimensional space and their semantic labels from partial point clouds, thereby not only inferring spatial occupancy but also assigning specific semantic meanings (such as buildings, trees, roads, etc.) to each voxel.

[0004] However, existing point cloud semantic scene completion networks still have problems of poor scene completion integrity and low semantic segmentation accuracy. Summary of the Invention

[0005] In view of this, the present invention provides a point cloud semantic scene completion system to solve the problems of poor scene completion integrity and low semantic segmentation accuracy.

[0006] A point cloud semantic scene completion system includes a voxelization processing module, a multi-scale voxel feature extraction module, and a bird's-eye view feature processing module;

[0007] The voxelization processing module takes a single-frame LiDAR point cloud as input, discretizes the single-frame LiDAR point cloud to obtain a structured 3D voxel grid, then performs feature extraction operations on each voxel in the structured 3D voxel grid to obtain each point-level feature, maps the point-level features through a multi-layer perceptron, and then performs aggregation using a max-pooling operation to obtain initial voxel features;

[0008] The multi-scale voxel feature extraction module processes the initial voxel features to obtain multi-scale voxel features. The multi-scale voxel feature extraction module uses an improved Cylinder3D encoder, and the improved Cylinder3D encoder includes an asymmetric residual block and an asymmetric downsampling block;

[0009] The bird's-eye view feature processing module includes an encoder and a decoder. The encoder processes the initial voxel features, extracts multi-scale bird's-eye view features, fuses the multi-scale bird's-eye view features with the multi-scale voxel features, and sends the fused features to the decoder. The decoder finally outputs the prediction results of semantic scene completion based on skip connections, feature concatenation, and content-aware feature recombination.

[0010] According to the point cloud semantic scene completion system provided by the present invention, the single-frame lidar point cloud is discretized by the voxelization processing module to obtain initial voxel features. The initial voxel features are respectively sent to the multi-scale voxel feature extraction module and the bird's-eye view feature processing module. The multi-scale voxel feature extraction module uses an improved Cylinder3D encoder to focus on extracting fine-grained 3D spatial features. The bird's-eye view feature processing module projects the voxel data onto a two-dimensional plane and effectively captures rich context information through the 2D convolutional networks of the encoder and the decoder. Fusing the multi-scale bird's-eye view features with the multi-scale voxel features in the encoder can make the advantages of the two features complementary, enabling the system to take into account both detailed geometric information and global context, and thus achieving more accurate and complete scene reconstruction, solving the problems of poor scene completion integrity and low semantic segmentation accuracy. The results on the evaluation website of the public dataset SemanticKITTI show that the system of the present invention performs best in the scene completion task compared with other existing algorithms. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 is a structural block diagram of the point cloud semantic scene completion system provided by an embodiment of the present invention;

[0012] Figure 2 is a comparative diagram of qualitative visualization results. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0013] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals indicate the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the embodiments of the present invention and should not be construed as limiting the present invention.

[0014] Please refer to Figure 1 , an embodiment of the present invention provides a point cloud semantic scene completion system, including a voxelization processing module, a multi-scale voxel feature extraction module, and a bird's-eye view feature processing module.

[0015] The voxelization processing module takes a single-frame LiDAR point cloud as input, discretizes the single-frame LiDAR point cloud to obtain a structured 3D voxel grid, then performs feature extraction operations on each voxel in the structured 3D voxel grid to obtain individual point-level features, maps the point-level features through a multi-layer perceptron, and then uses max pooling operations for aggregation to obtain initial voxel features.

[0016] In the voxelization processing module, first, points within a predefined range are selected from the input single-frame LiDAR point cloud to obtain a filtered point cloud set , , denotes the th point cloud, denotes the number of point clouds. Based on the point cloud set , a structured 3D voxel grid , , denotes the th voxel. , , are the numbers of voxels corresponding to the length, width, and height dimensions in 3D space after point cloud voxelization respectively. The voxel is represented by its centroid coordinate . To ensure the structuring of the scene representation, any point in the point cloud is assigned to the voxel closest to it.

[0017] After voxelization is completed, feature extraction operations are performed on each voxel in the structured 3D voxel grid to obtain individual point-level features. The point-level features can capture the spatial and reflection characteristics of points. The expression for the th point-level feature is:

[0018] ;

[0019] where denotes the offset between and the corresponding voxel centroid, denotes the absolute spatial coordinate of , is the reflection intensity of

[0020] Then, the point-level features are subjected to feature mapping through a multi-layer perceptron (MLP). To enhance the feature representation, the multi-layer perceptron maps the 7-dimensional point features to a high-dimensional space through successive non-linear transformations. In this embodiment, the multi-layer perceptron consists of stacked fully-connected layers, batch normalization layers, and ReLU activation functions, capable of extracting richer features. Then, a max-pooling operation is adopted for aggregation to obtain the initial voxel features. This operation can not only capture the most crucial spatial and semantic information but also effectively reduce data redundancy and computational complexity. The expression of the initial voxel features is:

[0021] ;

[0022] ;

[0023] Among them, represents the initial voxel features, represents the aggregated features of the th voxel, is a linear layer with a ReLU activation function, used to further compress the aggregated high-dimensional features; represents the max-pooling operation, represents the th point-level feature, represents the th point cloud.

[0024] Each voxel is encoded as a D -dimensional feature vector. For the entire scene, the size of the initial voxel features is , which constitutes a structured and semantically rich scene representation and serves as the input for subsequent tasks.

[0025] The multi-scale voxel feature extraction module processes the initial voxel features to obtain multi-scale voxel features. The multi-scale voxel feature extraction module uses an improved Cylinder3D encoder, and the improved Cylinder3D encoder includes asymmetric residual blocks and asymmetric downsampling blocks.

[0026] Among them, the multi-scale voxel feature extraction module uses an improved Cylinder3D encoder to process the initial voxel features , capturing detailed spatial information at different scales within the 3D volume. The multi-scale voxel feature extraction module effectively encodes the local and global context from the voxelized point cloud data.

[0027] In the multi-scale voxel feature extraction module, the asymmetric residual blocks process the initial voxel features through two different branches One branch applies consecutive 3D submanifold sparse convolutions with convolutional kernel sizes of 1×3×3 and 3×1×3 in sequence; the other branch applies consecutive 3D submanifold sparse convolutions with convolutional kernel sizes of 3×1×3 and 1×3×3 in sequence; finally, the output tensors from the two branches are added element-wise; this combination of two different branches reduces the computational amount compared to using a 3×3×3 convolution and also significantly enhances the expressive power of convolutional kernels in the horizontal and vertical directions, which is beneficial for extracting the features of cuboid-shaped objects in the 3D point cloud scene. This is in line with the fact that the point cloud distributions of some objects in the autonomous driving scenario, such as cars, trucks, bicycles, motorcycles, etc., often approximate the cuboid shape.

[0028] The result of the addition is input to the asymmetric downsampling block, which is a 3D sparse convolutional layer with a stride of 2 set after the asymmetric residual block. The asymmetric downsampling block applies 2-fold downsampling three times, and multi-scale voxel features are obtained after being processed by the asymmetric downsampling block. , , where, represents the set of real numbers, represents the th number of channels of the scale voxel feature, = 1, 2, 3. The multi-scale voxel features effectively represent the hierarchical structure of the 3D scene.

[0029] The bird's eye view feature processing module includes an encoder and a decoder. The encoder processes the initial voxel features, extracts multi-scale bird's eye view (BEV) features, fuses the multi-scale bird's eye view features with the multi-scale voxel features, and conveys the fused features to the decoder. The decoder finally outputs the prediction results for semantic scene completion based on skip connections, feature concatenation, and content-aware feature recombination.

[0030] Among them, the bird's eye view feature processing module projects the 3D voxel grid onto the 2D xy plane, and then applies the max pooling operation to aggregate voxel features within each 2D grid cell.

[0031] Specifically, in the encoder, the encoder first projects the initial voxel features , and the z coordinate of the voxel is ignored during the projection, thereby generating a 2D grid with a resolution of . Then, feature fusion is performed through max pooling. The expression of the bird's eye view feature of the th 2D grid cell is:

[0032] ;

[0033] Among them, It is a processing unit composed of a linear layer and a ReLU activation function. represents the th voxel, represents the aggregated feature of the th voxel;

[0034] Furthermore, the initial bird's-eye view feature is obtained. , where each 2D grid cell is encoded as a D -dimensional vector.

[0035] Then, the encoder performs 4 two-fold downsampling operations on the initial bird's-eye view feature in sequence to extract the multi-scale bird's-eye view feature . . represents the number of channels of the th scale voxel feature, = 1, 2, 3, 4; where the initial bird's-eye view feature is first downsampled to obtain the bird's-eye view feature after the first downsampling , then is fused with the first scale voxel feature , and after fusion, it is input into the downsampling module of the encoder. After convolution operation, the second downsampling is performed to obtain the bird's-eye view feature after the second downsampling , then is fused with the second scale voxel feature , and after fusion, it is input into the downsampling module of the encoder. After convolution operation, the third downsampling is performed to obtain the bird's-eye view feature after the third downsampling , then is fused with the third scale voxel feature , and after fusion, it is input into the downsampling module of the encoder. After convolution operation, the fourth downsampling is performed, and finally the fused feature is obtained.

[0036] Through three times of fusion, is stacked along the z-axis and the number of channels is reduced by convolution to match the channel dimension. By fusing the multi-scale bird's-eye view feature and the multi-scale voxel feature, the 3D scene representation is enhanced, and the complementary advantages of voxels and BEV views are fully utilized, thereby achieving a more comprehensive scene understanding.

[0037] In this embodiment, an Atrous Spatial Pyramid Pooling (ASPP) module is provided between the encoder and the decoder. In the Atrous Spatial Pyramid Pooling module, considering the size of the input features, 1, 2, and 4 are selected as the dilation rate combination to perform dilated convolution. This setting enables each convolutional kernel to operate efficiently and avoids some convolutional kernel elements exceeding the feature range of the feature map due to too high dilation rate. The Atrous Spatial Pyramid Pooling module captures multi-scale context information by parallelly using convolutional kernels with different dilation rates while maintaining the input resolution unchanged. Relying on this advantage of the Atrous Spatial Pyramid Pooling module, the system proposed in this application can effectively integrate and refine features at each scale before upsampling, significantly enhancing the feature expression ability.

[0038] Then, based on skip connections, feature concatenation, and content-aware feature recombination, the decoder finally outputs the prediction result of semantic scene completion.

[0039] Specifically, in the decoder, a content-aware feature recombination module is used to upsample the fused features. The content-aware feature recombination module can process the features of the low-resolution feature map and construct a high-resolution output by generating an adaptive kernel and recombining features.

[0040] For any position of the final upsampled output feature map features , there is a corresponding position in the original input feature map features . is the upsampling ratio. Here, two-fold upsampling is adopted. denotes the floor operation. Let denote the size of the upsampling convolutional kernel. Then, the features at each target position of the output features are determined by the region centered on the source position and the predicted recombination kernel.

[0041] Specifically, the content-aware feature recombination module includes a recombination kernel prediction unit and a feature recombination unit.

[0042] The recombination kernel prediction unit is used to dynamically predict a recombination kernel according to the input fused features and in combination with each position of the final output features. The recombination kernel is a set of weight parameters used to allocate corresponding weights to different regions in the upsampling process according to the input fused features, so as to achieve content-aware feature recombination. Subsequently, these predicted recombination kernels are applied to the input features to achieve content-aware feature recombination.

[0043] The feature recombination unit is used to perform content-aware feature recombination based on the recombination kernel, and then obtain the prediction result of semantic scene completion;

[0044] In the reconstructed kernel prediction unit, the fused features are first reduced in the number of feature channels through a 1×1 convolutional kernel, and then passed through a convolution operation with a size of to generate a reconstructed kernel. The channel size of the reconstructed kernel is , is the upsampling magnification, is the size of the upsampling convolutional kernel. Then, pixel-level rearrangement is performed on the reconstructed kernel. After rearrangement, the size of the reconstructed kernel is . Finally, the reconstructed kernel is normalized in the channel dimension using the softmax function;

[0045] In the feature reconstruction unit, based on the upsampling weight , feature reconstruction is performed within the neighborhood of the target position. The expression for feature reconstruction is:

[0046] ;

[0047] where, represents the output feature at the target position , represents the radius of the local sampling neighborhood, , represents the floor operation, represents at the position value, represents the feature matrix of the region centered on the source position corresponding to the target position, represents at the position feature.

[0048] The total loss function of the above point cloud semantic scene completion system is:

[0049] ;

[0050] ;

[0051] ;

[0052] ;

[0053] ;

[0054] ;

[0055] ;

[0056] ;

[0057] Among them, represents the weighted cross - entropy loss function, represents the Lovasz loss function, represents the focal loss function, represents the scene - class affinity loss function, is the true label corresponding to the ith voxel for a certain class, is the weight of class ; is the probability output value that the ith voxel is predicted to be a certain class, represents the total number of classes, represents a piece - wise linear function with a global minimum, is the error vector of class ; is the balance factor of class ; is the focusing parameter, , and represent precision, recall, and specificity respectively; represents the binary value indicating whether the ith voxel belongs to a certain class . If the ith voxel belongs to a certain class , the value is 1; if the ith voxel does not belong to a certain class , the value is 0.

[0058] The present invention constructs a total loss function specifically for semantic scene completion, further improving the performance of point - cloud semantic scene completion. Among them, the weighted cross - entropy loss function is a standard loss function in classification tasks, which is used to assist the model in predicting the correct label for each voxel. The Lovasz loss function solves the class imbalance problem in semantic segmentation tasks by optimizing the intersection - over - union score, thereby enhancing the model's detection ability for foreground objects. The focal loss function reduces the contribution of well - classified samples, enabling the model to pay more attention to difficult - to - classify and misclassified samples, thus further improving the model's performance on challenging parts of the scene. The scene - class affinity loss function is extended from the binary affinity loss and helps to enhance the spatial structure and semantic consistency of scene completion.

[0059] The system of the present invention is tested below.

[0060] Test dataset:

[0061] SemanticKITTI is an outdoor large-scale LiDAR point cloud dataset extended from the KITTI visual odometry dataset, mainly used for research in the field of autonomous driving, especially in semantic segmentation and semantic scene completion. This dataset is obtained by scanning with a single Velodyne HDL-64E model LiDAR sensor in various complex scenarios such as the inner city, residential areas, highways, and rural roads around Karlsruhe, Germany, and dense semantic annotations are made on the collected point cloud data. The entire dataset contains 22 sequences, among which sequences 00 - 10 are divided into the training set, sequences 11 - 21 are used as the test set, and sequence 08 in the training set is used as the validation set. The training set provides 23,201 complete 3D point cloud scans, while the test set provides 20,351 scans. It should be noted that only the training set provides label information, and the test set does not directly provide labels. Researchers need to submit the prediction results of the test set to the CodaLab competition platform, and the test set results can be obtained through the evaluation service of this platform.

[0062] This dataset covers 28 categories, including categories that distinguish between moving and non-moving objects. These categories are divided into types such as ground, buildings, transportation, nature, humans, objects, and others. Ground categories (such as roads, sidewalks, parking lots) occupy most of the point cloud in the dataset, followed by the nature category and the transportation category. In contrast, the point cloud of categories such as pedestrians, cyclists, and motorcyclists is relatively small, showing a typical long-tail distribution. For the SSC task, 25 out of the 28 categories are used, excluding the outlier, other structures, and other object categories because these categories are not considered during the evaluation process. In addition, the moving categories are merged with their corresponding non-moving categories, resulting in a total of 19 evaluation categories.

[0063] Experimental setup:

[0064] The valid range of the point cloud data is set to [0~51.2 m, -25.6~25.6 m, -2~4.4 m], and the scene is divided into 256x256x32 voxel units, with each voxel unit representing a spatial resolution of 0.2 m. In the data preprocessing stage, this application adopts a data augmentation strategy. Specifically, the augmentation strategy includes the following three types: mirror flipping along the x-axis, mirror flipping along the y-axis, and mirror flipping along both the x-axis and the y-axis simultaneously. During the training process, one of the augmentation methods will be randomly selected to process the point cloud data to enhance the robustness of the model. The network training uses the PyTorch framework and is carried out on an NVIDIA GeForce RTX 3090 GPU with 24GB of video memory. The Adam optimizer is adopted. In addition to the default parameters, the momentum term is set to 0.9, and the weight decay coefficient is set to 0.0001. The initial learning rate is set to 0.001 and decays to 98% of the original value after each epoch. The training process is carried out for 80 epochs. Due to the GPU video memory limitation, the batch size is set to 2 to ensure that there is no video memory overflow during the training process.

[0065] Comparative experiment:

[0066] This application demonstrates a detailed comparison of various methods on the official SemanticKITTI benchmark dataset. Among them, LMSCNet processes large-scale missing scenes through multi-scale 3D convolution and context aggregation modules, and uses voxelization representation to achieve indoor and outdoor scene completion; Local-DIFs divides the point cloud into neighborhoods and learns continuous implicit functions, and achieves high-precision detail reconstruction through the signed distance field (SDF); JS3C jointly optimizes the semantic segmentation and geometric completion tasks, and uses cross-task feature interaction to improve the semantic consistency of complex scenes; UDNet adopts a symmetric encoder-decoder structure, combines skip connections to achieve multi-level geometric feature fusion, and is good at processing noisy point clouds; SSA-SC introduces a non-local self-attention mechanism and restores a globally reasonable structure through long-range dependence modeling; SSC-RS designs a residual skip architecture for the sparse characteristics of lidar and progressively refines the missing areas. The comparison results are shown in Tables 1, 2, and 3. The experiments focus on the evaluation metrics of the intersection over union (IoU) for the scene completion task, the mean intersection over union (mIoU) for the semantic segmentation task, and the frames per second (FPS).

[0067] Table 1 Comparison of test results on the SemanticKITTI dataset for the completion task

[0068]

[0069] Table 2 Comparison of test results on the SemanticKITTI dataset for the semantic segmentation task

[0070]

[0071] Table 3 Comparison of processing speed results on the SemanticKITTI dataset

[0072]

[0073] As can be seen from Table 1, for the scene completion task, the system of the present application achieved an IoU of 62.6%, exceeding all baseline models. In particular, the present application observed significant improvements compared to methods with better performance such as SSC-RS (59.7%), UDNet (59.4%), and SSA-SC (58.8%). This proves the effectiveness of the present application in reconstructing a complete 3D scene.

[0074] As can be seen from Table 2, in the semantic segmentation task, measured by mIoU, the present application also performed excellently, with an mIoU reaching 26.5%. This is significantly better than other methods, including SSC-RS (24.2%), SSA-SC (23.5%), and Local-DIFS (22.7%). It is worth noting that the system of the present application achieved excellent segmentation performance in various categories, especially in challenging categories such as buildings, cars, trucks, and vegetation, exceeding existing methods in all cases.

[0075] As can be seen from Table 3, in terms of FPS, LMSCNet performed best, reaching 21.3. However, its IoU and mIoU performance was poor. Compared with most algorithms, the present application achieved a comparable FPS while ensuring accuracy, achieving a good balance between speed and accuracy.

[0076] Figure 2 is a qualitative visualization result comparison, demonstrating excellent geometric reconstruction and semantic annotation capabilities in areas with complex scene structures, Figure 2 The circled parts in indicate the key comparison objects. From Figure 2 it can be seen that compared with SSA-SC and SSC-RS, the system of the present application showed better details in geometric reconstruction and semantic segmentation, especially in areas with complex scene structures.

[0077] The above results verify the effectiveness of the system of the present application in semantic scene completion, demonstrating its ability to capture local details and global context to enhance scene understanding.

[0078] In summary, for the point cloud semantic scene completion system according to the above embodiments, the voxelization processing module discretizes the single-frame lidar point cloud to obtain the initial voxel features. The initial voxel features are respectively sent to the multi-scale voxel feature extraction module and the bird's-eye view feature processing module. The multi-scale voxel feature extraction module uses an improved Cylinder3D encoder, which focuses on extracting fine-grained 3D spatial features. The bird's-eye view feature processing module projects the voxel data onto a two-dimensional plane and effectively captures rich context information with the 2D convolutional networks of the encoder and decoder. The multi-scale bird's-eye view features and the multi-scale voxel features are fused in the encoder, enabling the complementary advantages of the two features, allowing the system to take into account both detailed geometric information and global context, and thus achieving more accurate and complete scene reconstruction, solving the problems of poor scene completion integrity and low semantic segmentation accuracy. The results on the evaluation official website of the public dataset SemanticKITTI show that the system of the present invention performs best in the scene completion task compared with other existing algorithms.

[0079] The above embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the invention patent shall be subject to the appended claims.

Claims

1. A point cloud semantic scene completion system, characterized in that, It includes a voxelization processing module, a multi-scale voxel feature extraction module, and a bird's-eye view feature processing module; The voxelization processing module takes a single-frame lidar point cloud as input, discretizes the single-frame lidar point cloud to obtain a structured 3D voxel grid, then performs feature extraction operations on each voxel in the structured 3D voxel grid to obtain each point-level feature, maps the point-level features through a multi-layer perceptron, and then aggregates them using a max pooling operation to obtain initial voxel features; The multi-scale voxel feature extraction module processes the initial voxel features to obtain multi-scale voxel features. The multi-scale voxel feature extraction module uses an improved Cylinder3D encoder, and the improved Cylinder3D encoder includes an asymmetric residual block and an asymmetric downsampling block; The bird's-eye view feature processing module includes an encoder and a decoder. The encoder processes the initial voxel features to extract multi-scale bird's-eye view features, fuses the multi-scale bird's-eye view features with the multi-scale voxel features, and sends the fused features to the decoder. The decoder is based on skip connections, feature concatenation, and content-aware feature recombination, and finally outputs the prediction result of semantic scene completion.

2. The point cloud semantic scene completion system according to claim 1, wherein In the voxelization processing module, first, points within a predefined range are filtered out from the input single-frame lidar point cloud to obtain a filtered point cloud set. , , denotes the th point cloud, denotes the number of point clouds. Based on the point cloud set , a structured 3D voxel grid , , is obtained. denotes the th voxel, , are the numbers of voxels corresponding to the length, width, and height dimensions in the 3D space after point cloud voxelization, respectively. Then, perform a feature extraction operation on each voxel in the structured 3D voxel grid to obtain each point-level feature. The expression of the point-level feature is as follows: ; Among them, denotes the offset between the corresponding body centroid, denotes the absolute spatial coordinates of, is the reflection intensity of; Then, the point-level features are mapped through a multi-layer perceptron. The multi-layer perceptron is composed of stacked fully connected layers, batch normalization layers, and ReLU activation functions, and then aggregated using a max pooling operation to obtain initial voxel features. The expression of the initial voxel features is: ; ; Among them, represents the initial voxel feature, represents the aggregation feature of the th voxel, is a linear layer with a ReLU activation function, represents the max pooling operation, represents the th point-level feature, represents the th point cloud.

3. The point cloud semantic scene completion system according to claim 2, wherein In the multi-scale voxel feature extraction module, the asymmetric residual block processes the initial voxel features through two different branches : one branch applies consecutive 3D sub-manifold sparse convolutions with kernel sizes of 1×3×3 and 3×1×3 in sequence; the other branch applies consecutive 3D sub-manifold sparse convolutions with kernel sizes of 3×1×3 and 1×3×3 in sequence; finally, the output tensors from the two branches are added element-wise; the result of the addition is input to the asymmetric downsampling block, which is a 3D sparse convolutional layer with a stride of 2 set after the asymmetric residual block, and the multi-scale voxel features are obtained after being processed by the asymmetric downsampling block , , where represents the set of real numbers, represents the number of channels of the -th scale voxel feature, = 1, 2, 3 4. The point cloud semantic scene completion system according to claim 3, wherein In the encoder, the encoder first projects the initial voxel features to generate a 2D grid with a resolution of . Then, feature fusion is performed through max pooling. The expression for the bird's-eye view feature of the th 2D grid cell is as follows: ​ ; Among them, is a processing unit composed of a linear layer and a ReLU activation function, represents the ith voxel, represents the aggregated feature of the ith voxel; Furthermore, the initial bird's-eye view features are obtained. , ; Then the encoder processes the initial bird's-eye view features and successively performs 4 downsampling operations by a factor of 2 to extract multi-scale bird's-eye view features , , denotes the number of channels of the th scale voxel feature, = 1, 2, 3, 4; among which, first the initial bird's-eye view features are downsampled for the first time to obtain the bird's-eye view features after the first downsampling , and then is fused with the first scale voxel feature . After fusion, it is input into the downsampling module of the encoder, and after convolution operation, it is downsampled for the second time to obtain the bird's-eye view features after the second downsampling , and then is fused with the second scale voxel feature . After fusion, it is input into the downsampling module of the encoder, and after convolution operation, it is downsampled for the third time to obtain the bird's-eye view features after the third downsampling , and then is fused with the third scale voxel feature . After fusion, it is input into the downsampling module of the encoder, and after convolution operation, it is downsampled for the fourth time to finally obtain the fused features.

5. The point cloud semantic scene completion system according to claim 1, characterized in that, A dilated spatial pyramid pooling module is provided between the encoder and the decoder. In the dilated spatial pyramid pooling module, 1, 2, and 4 are selected as the dilation rate combination to perform dilated convolution.

6. The point cloud semantic scene completion system according to claim 4, characterized in that In the decoder, a content-aware feature recombination module is used to upsample the fused features. The content-aware feature recombination module includes a recombination kernel prediction unit and a feature recombination unit; The recombination kernel prediction unit is used to dynamically predict a recombination kernel according to the input fused features and in combination with each position of the final output features. The recombination kernel is a set of weight parameters used to allocate corresponding weights to different regions in the upsampling process according to the input fused features, so as to achieve content-aware feature recombination; The feature recombination unit is used to perform content-aware feature recombination based on the recombination kernel, and then obtain the prediction result of semantic scene completion; In the recombinant kernel prediction unit, the fused features are first reduced in the feature channels through a 1×1 convolutional kernel, and then passed through a convolutional operation with a size of to generate a recombinant kernel. The channel size of the recombinant kernel is , is the upsampling magnification, is the size of the upsampling convolutional kernel. Then, pixel-level rearrangement is performed on the recombinant kernel. After rearrangement, the size of the recombinant kernel is . Finally, the recombinant kernel is normalized in the channel dimension using the normalized exponential function; In the feature recombination unit, based on the upsampling weights , feature recombination is performed within the neighborhood of the target position, and the expression for feature recombination is: ; Among them, represents the output feature of the target position . represents the radius of the local sampling neighborhood , represents the floor operation represents at the position value represents the feature matrix of the area centered on the source position corresponding to the target position represents at the position feature 7. The point cloud semantic scene completion system according to claim 6, characterized in that The total loss function of the system is as follows: ; ; ; ; ; ; ; ; Among them, represents the weighted cross-entropy loss function, represents the Lovász loss function, represents the focal loss function, represents the scene-class affinity loss function, is the true label corresponding to the ith voxel's corresponding category, is the weight of category ; is the probability output value that the ith voxel is predicted to be category represents the total number of categories, represents a piecewise linear function with a global minimum, is the error vector of category is the balance factor of category is the focusing parameter, , and respectively represent precision, recall, and specificity; represents the binary value indicating whether the ith voxel belongs to category ; if the ith voxel belongs to category , the value is 1; if the ith voxel does not belong to category , the value is 0.

Citation Information

Patent Citations

  • Laser radar 3D real-time target detection method fusing multi-frame time sequence point cloud

    CN111429514A

  • Point cloud 3D target detection method based on symmetric point generation

    CN112598635A