A semantic scene completion method based on feature representation decomposition and bird's-eye view fusion

By using the technology of fusion of feature representation decomposition and bird's eye view in the semantic scene completion method, the problems of large amount of calculation and low recognition accuracy in the existing technology are solved, efficient and real-time semantic scene completion are achieved, and good robustness and detailed recovery capabilities are provided.

CN116630975BActive Publication Date: 2025-06-17ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310562753.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-18
Publication Date
2025-06-17
Estimated Expiration
2043-05-18

AI Technical Summary

Technical Problem

In the prior art, semantic scene completion methods have problems such as large amount of computation and low recognition accuracy, especially when dealing with large-scale outdoor scenes, it is difficult to effectively restore semantics and geometric shapes.

Method used

A point cloud semantic scene completion method based on feature representation decomposition and bird's eye view fusion is proposed. By designing separate semantic branches and complementary branches, semantic features and geometric features are extracted respectively, and feature fusion from the aerial view perspective is performed in the semantic completion branch.

Benefits of technology

It realizes efficient calculation and high accuracy of semantic scene completion, can complete semantic scenes on the input point cloud in real time, has good robustness, can better restore semantic scene details, and make the completed scene more in line with reality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116630975B_ABST
    Figure CN116630975B_ABST
Patent Text Reader

Abstract

A semantic scene completion method based on feature representation decomposition and bird's-eye view fusion, comprising the following steps: Step S1: Obtain the target point cloud data to be completed; Step S2: Use the pre-trained semantic branch to extract the semantic features of the target point cloud data to be completed; Step S3: Use the pre-trained completion branch to extract the geometric features of the target point cloud data to be completed; Step S4: The semantic features and geometric features are respectively mapped to the bird's-eye view perspective and then input into the pre-trained semantic completion branch for feature fusion to obtain the semantic scene completion result; The method of the present invention has a fast calculation speed and can perform semantic scene completion on the input point cloud in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of 3D vision technology, and particularly relates to a point cloud semantic scene completion method based on feature representation decomposition and bird's-eye view fusion. Background Art

[0002] In recent years, 3D scene understanding, as one of the most important functions of the perception system in autonomous driving, has attracted extensive research and achieved rapid progress. When dealing with large-scale outdoor scene understanding, semantic scene completion (SSC) aims to predict the semantic occupancy of each voxel in the entire 3D scene from sparse LiDAR scans, including the completion of certain regions. Since it can restore the geometric structure, SSC can facilitate further applications such as 3D object detection, which are usually affected by the sparsity and incompleteness of LiDAR point clouds. However, due to complex outdoor scenes with various shapes / sizes and occlusions, it is challenging to accurately estimate the semantics and geometry of the entire 3D real-world scene from partial observations.

[0003] Following the pioneering work of SSCNet, some existing outdoor SSC methods use a single U-Net network, such as a dense 3D convolutional network, to jointly predict semantics and geometry. However, dense 3D CNNs usually involve unnecessary computations and bring additional memory and computational overhead, especially when the input voxel resolution is large because there are a large number of empty voxels in the 3D scene. On the other hand, some methods combine the semantic completion network with the segmentation network and use the semantic information in segmentation to assist outdoor SSC, but this method of scene completion is not accurate enough. Therefore, the SSC methods in the prior art have the problems of large computational amount and low recognition accuracy.

[0004] In order to better achieve semantic scene completion, the semantic scene completion method needs to meet the conditions of good restoration of semantic scene details, a more realistic scene after completion, etc., and the algorithm should be simple and efficient. Summary of the Invention

[0005] In view of the above problems, the present invention proposes a point cloud semantic scene completion method based on feature representation decomposition and bird's-eye view fusion, designs a separate branch for semantic / geometric feature representation, and also designs a BEV fusion network, namely a semantic completion branch, to fuse two types of features from the two branches, which has a strong feature expression ability, fast calculation speed, and can perform semantic scene completion on the input point cloud in real time.

[0006] To achieve the above object, the present invention provides a semantic scene completion method based on feature representation decomposition and bird's-eye view fusion, including the following steps:

[0007] Step S1: Obtain the target point cloud data to be completed, and construct a semantic branch, a completion branch, and a semantic completion branch;

[0008] Step S2: Use the pre-trained semantic branch to extract the semantic features of the target point cloud data to be completed;

[0009] Step S3: Use the pre-trained completion branch to extract the geometric features of the target point cloud data to be completed;

[0010] Step S4: After the semantic features and geometric features are respectively mapped to the bird's-eye view perspective, they are input into the pre-trained semantic completion branch for feature fusion to obtain the semantic scene completion result.

[0011] Preferably, in step S4, the semantic completion branch includes an adaptive feature fusion module, and the adaptive feature fusion module is used to fuse the semantic features in the bird's-eye view perspective and the geometric features in the bird's-eye view perspective.

[0012] Preferably, in step S2, the semantic branch includes N semantic feature extraction modules;

[0013] In step S3, the completion branch includes N geometric feature extraction modules; the N geometric feature extraction modules correspond to the N semantic feature extraction modules respectively;

[0014] In step S4, the semantic completion branch includes N adaptive feature fusion modules; the N adaptive feature fusion modules correspond to the N semantic feature extraction modules respectively; the N adaptive feature fusion modules correspond to the N geometric feature extraction modules respectively; the semantic features and geometric features extracted by the semantic feature extraction module and the corresponding geometric feature extraction module are respectively mapped to the bird's-eye view perspective and then input into the corresponding adaptive feature fusion module for feature fusion;

[0015] The N adaptive feature fusion modules are connected in series. Each adaptive feature fusion module outputs a stage fusion feature. The first adaptive feature fusion module outputs the first stage fusion feature, and then the first stage fusion feature is passed to the next adaptive feature fusion module in series with it, and is fused with the semantic features in the bird's-eye view perspective and the geometric features in the bird's-eye view perspective input in the next adaptive feature fusion module to output the next stage fusion feature, and is passed one by one until the last adaptive feature fusion module performs the final feature fusion.

[0016] Preferably, in step S2, the semantic branch further includes a voxelization layer, and the N semantic feature extraction modules include a first sparse coding block, a second sparse coding block, and a third sparse coding block; specifically, there are 3 semantic feature extraction modules, and each semantic feature extraction module includes 1 sparse coding block, specifically the first sparse coding block, the second sparse coding block, and the third sparse coding block; preferably, there is 1 voxelization layer;

[0017] Step S2 includes the following steps:

[0018] Step S21: Input the target point cloud data to be completed into the voxelization layer for voxelization to obtain voxel features;

[0019] Step S22: First input the voxel features into the first sparse coding block, and then output the first semantic feature and the first voxel feature. The first voxel feature is then input into the second sparse coding block, and then output the second semantic feature and the second voxel feature. The second voxel feature is finally input into the third sparse coding block, and then output the third semantic feature; Preferably, each sparse coding block consists of a sparse convolutional residual block and a sparse geometric feature extraction module. The sparse convolutional residual block halves the resolution of the input voxel features, and the sparse geometric feature extraction module uses different-scale sparse mapping and attention selection mechanisms to enhance the geometric characteristics of the voxel features;

[0020] In step S3, the completion branch further includes an input layer. The N geometric feature extraction modules include a first dense residual block, a second dense residual block, and a third dense residual block; Specifically, there are 3 geometric feature extraction modules, and each geometric feature extraction module includes 1 dense residual block, specifically the first dense residual block, the second dense residual block, and the third dense residual block;

[0021] Step S3 includes the following steps:

[0022] Step S31: Calculate the occupied voxels of the target point cloud data to be completed; The occupied voxels are 0 / 1 binary values, generated from the point cloud. If a voxel contains points, it is 1, otherwise it is 0;

[0023] Step S32: Input the occupied voxels into the input layer. After processing by the input layer to increase the receptive field, then input into the first dense residual block to output the first geometric feature and the first occupied voxels. The first occupied voxels are then input into the second dense residual block to output the second geometric feature and the second occupied voxels. The second occupied voxels are finally input into the third dense residual block to output the third geometric feature; The input layer is a 7×7×7 3D dense convolution, and each dense residual block consists of a 3×3×3 3D dense convolution;

[0024] In step S4, the semantic completion branch includes a first adaptive feature fusion module, a second adaptive feature fusion module, and a third adaptive feature fusion module;

[0025] Step S4 includes the following steps:

[0026] Step S41: Map the target point cloud data to be completed to the bird's-eye view perspective to obtain the target bird's-eye view features to be completed;

[0027] Step S42: After the first semantic feature and the first geometric feature are respectively mapped to the bird's-eye view perspective, they are input into the first adaptive feature fusion module together with the target bird's-eye view feature to be completed for feature fusion to obtain the first-stage fusion feature. After the second semantic feature and the second geometric feature are respectively mapped to the bird's-eye view perspective, they are input into the second adaptive feature fusion module together with the first-stage fusion feature for feature fusion to obtain the second-stage fusion feature. After the third semantic feature and the third geometric feature are respectively mapped to the bird's-eye view perspective, they are input into the third adaptive feature fusion module together with the second-stage fusion feature for feature fusion to obtain the third-stage fusion feature. The fusion features of the three stages are passed through the decoder to obtain the semantic scene completion result.

[0028] Preferably, the specific steps of voxelization are as follows:

[0029] Let P represent the target point cloud data to be completed, and p i =(x i , y i , z i ) represent a point in the target point cloud data to be completed. Its voxel index where s is the resolution of voxelization, is the floor operation;

[0030]

[0031] f Vm represents the voxel feature of the m-th non-empty voxel with voxel index V m ; R f represents a fully connected layer for reducing the dimension of the input feature to 64; MLP represents a multi-layer perceptron for encoding the input point cloud feature; A f is an aggregation function, usually an average pooling function, for aggregating the features of all points belonging to the same voxel; f p represents a point feature with a dimension of 7, including the coordinates of the point cloud (3D), the offset vector of the coordinates of each point in the point cloud from the center of the voxel where it is located (3D), and the radar point cloud reflection intensity (1D); V p represents the voxel index where point p is located;

[0032] Preferably, step S4 includes the following steps:

[0033] Step S4a: Calculate the sparse voxel index corresponding to the semantic feature, calculate the bird's-eye view index through the sparse voxel index, use the aggregation function to map the semantic feature to the sparse bird's-eye view feature, and finally generate the semantic feature in the bird's-eye view perspective according to the sparse bird's-eye view feature and the corresponding bird's-eye view index;

[0034] Step S4b: Perform max pooling on the geometric feature in the height dimension to obtain the geometric feature in the bird's-eye view perspective;

[0035] Step S4c: Input the semantic features and geometric features in the bird's-eye view into the adaptive feature fusion module for feature fusion to obtain the features output by the adaptive feature fusion module, and then perform multiple upsampling operations on the features output by the adaptive feature fusion module through the decoder to obtain the semantic scene completion result, the semantic scene completion result (L, H, W) is the size of the 3D voxelization space, and C is the semantic category. Preferably, the decoder upsampling step is specifically: the decoder gradually upsamples the features from the encoder 3 times through skip connections, and the upsampling multiple each time is 2.

[0036] Preferably, in step S4, each adaptive feature fusion module outputs a stage fusion feature, denoted by F prev represents the stage fusion feature output by the previous adaptive feature fusion module, and input F prev into the next adaptive feature fusion module connected in series with the previous adaptive feature fusion module. The semantic features and geometric features mapped to the bird's-eye view input into the next adaptive feature module are denoted by F sem and F com respectively; first calculate their respective channel attention weights, then multiply their respective channel attention weights by their corresponding features, and then add the results of their multiplication and pass through a 1×1 convolution to obtain the fused result.

[0037] The stage fusion feature output by the next adaptive feature module can be expressed as:

[0038] F f = φ{σ[MLP(AvgPool(F prev ))]·F prev + σ[MLP(AvgPool(F sem ))]·F sem + [MLP(AvgPool(F com ))]·F com}

[0039] F f represents the stage fusion feature output by the next adaptive feature module; σ is the Sigmoid function, which maps the input to the range (0, 1); AvgPool is the global average pooling; MLP is the multi-layer perceptron; Φ is the 1×1 convolution layer;

[0040] Preferably, pre-train the semantic branch in step S2, the completion branch in step S3, and the semantic completion branch in step S4. The pre-training steps include:

[0041] Step S51: Use the server to obtain the target point cloud training data to be completed;

[0042] Step S52: Use the semantic branch to extract the semantic features of the target point cloud training data to be completed, and perform semantic segmentation supervision to promote the learning of semantic context. The semantic supervision loss function consists of Lovasz loss and multi-class cross-entropy loss; The semantic segmentation supervision is specifically to use a lightweight multi-layer perceptron as an auxiliary head, output the prediction results of this stage, and calculate the loss with the labels of the corresponding scale;

[0043] Step S53: Use the completion branch to extract the geometric features of the target point cloud training data to be completed, and perform scene completion supervision to promote the learning of geometric information. The geometric supervision loss function consists of Lovasz loss and binary cross-entropy loss; The scene completion supervision is specifically to use a lightweight multi-layer perceptron as an auxiliary head, output the prediction results of this stage, and calculate the loss with the labels of the corresponding scale;

[0044] Step S54: The semantic features and geometric features are respectively mapped to the bird's-eye view perspective and then input into the semantic completion branch for feature fusion to obtain the semantic scene completion result. Perform main supervision on the semantic scene completion result. The semantic completion supervision loss function consists of Lovasz loss and multi-class cross-entropy loss;

[0045] Step S55: Use the server to perform network training and adopt an end-to-end method for multi-task training; The loss function L total is the weighted sum of the supervision loss functions in Step S52, Step S53, and Step S54;

[0046] Step S56: Use the server to optimize the loss function, obtain the locally optimal network parameters, and obtain the pre-trained semantic branch, completion branch, and semantic completion branch.

[0047] Preferably, in Step S52, the semantic supervision loss function can be expressed as:

[0048]

[0049] L lovasz,i and L ce,i respectively represent the i-th Lovasz loss and the i-th multi-class cross-entropy loss of semantic supervision;

[0050] In Step S53, the geometric supervision loss function can be expressed as:

[0051]

[0052] L lovasz,i and L bce,i respectively represent the i-th Lovasz loss and the i-th binary cross-entropy loss of geometric supervision;

[0053] In step S54, the semantic completion supervision loss function can be expressed as:

[0054] L bev = L lovasz + L ce

[0055] L lovasz and L ce respectively represent the Lovasz loss and multi-class cross-entropy loss of semantic completion supervision;

[0056] In step S55, the loss function L total can be expressed as:

[0057] L total = 3·L bev + L s + L c

[0058] Preferably, the Lovasz loss in steps S52, S53 and S54 is specifically:

[0059]

[0060] where J is the Lovasz extended version of IoU, and e(c) is the error vector of class c; the cross-entropy loss is specifically: where y i is the predicted value, is the true value.

[0061] Compared with the prior art, the beneficial effects of the present invention are:

[0062] The semantic scene completion method based on feature representation decomposition and bird's-eye view fusion provided by the present invention uses a pre-trained semantic branch and completion branch to extract the semantic features and geometric features of the target point cloud data to be completed respectively; in addition, a pre-trained semantic completion branch is used to effectively fuse the semantic features and geometric features from the bird's-eye view perspective. Based on the feature representation decomposition and then fusing the features, the semantic features and geometric features complement each other, with strong feature representation ability while the calculation is simple, which is beneficial to restoring the details of the semantic scene. Moreover, compared with the dense feature fusion in 3D space, the bird's-eye view fusion is more convenient and efficient; the semantic scene completion method proposed by the present invention is simple, efficient, with a fast calculation speed, can realize real-time semantic scene completion of the input point cloud, and has good robustness to problems such as target occlusion and fast movement in the actual 3D environment, and has the advantages of good restoration of semantic scene details and a more realistic completed scene.

[0063] Semantic context and geometric structure complement each other and are crucial for the SSC task. It is easy to restore geometric details based on semantics, while the complete geometric shape helps identify semantic categories. Explicitly separating feature representations can promote and accelerate the learning process of semantic context and geometric structure. We propose an adaptive feature fusion module in the semantic completion branch, which can obtain effective clues from semantic / geometric features and fully fuse semantic context and geometric details, greatly enhancing the feature expression ability.

[0064] N semantic feature extraction modules and geometric feature extraction modules are designed to extract semantic features and geometric features multiple times respectively, and an equal number of adaptive feature fusion modules are designed to perform multiple feature fusions on the semantically and geometrically features extracted in stages, further enhancing the feature expression ability. The semantic scene detail ability is excellent. When compared with the visualization results of SSC-SA and JS3CNet, the method proposed in the present invention has a better completion effect on moving objects and flat objects. The completed objects are clearer, more complete, and more in line with reality.

[0065] The design of the method of the present invention is positioned as lightweight. 3 sparse coding blocks are used to encode semantic features, and 3 lightweight dense residual blocks are used to obtain geometric features. The semantic completion branch (BEV fusion network) is used to fuse semantic / geometric features; making the method of the present invention have both a lightweight design and a powerful expression ability. When in use, the method has low latency, can run in real time, and has good generalization ability, achieving state-of-the-art performance on the SemanticKITTI test set. Brief Description of the Drawings

[0066] Figure 1 It is a schematic diagram of the overall algorithm framework of a semantic scene completion method based on feature representation decomposition and bird's-eye view fusion of the present invention;

[0067] Figure 2 It is a schematic diagram of the specific structures of the semantic branch and the completion branch designed by the present invention;

[0068] Figure 3 It is a schematic diagram of the specific structure of the adaptive feature fusion module designed by the present invention;

[0069] Figure 4 It is a comparison diagram of the visualization results of the method proposed by the present invention and other advanced methods (SSC-SA, JS3CNet);

[0070] Figure 5 It is a comparison of the method proposed by the present invention and other methods on the SemanticKITTI test set. Detailed Description of the Invention

[0071] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0072] A semantic scene completion method based on feature representation decomposition and bird's-eye view fusion, comprising the following steps:

[0073] Step S1: Obtain the target point cloud data to be completed; construct a semantic branch, a completion branch, and a semantic completion branch;

[0074] Step S2: Use the pre-trained semantic branch to extract the semantic features of the target point cloud data to be completed;

[0075] Step S3: Use the pre-trained completion branch to extract the geometric features of the target point cloud data to be completed;

[0076] Step S4: The semantic features and geometric features are respectively mapped to the bird's-eye view (BEV) perspective and then input into the pre-trained semantic completion branch for feature fusion to obtain the semantic scene completion result.

[0077] The semantic scene completion method based on feature representation decomposition and bird's-eye view fusion provided by the present invention uses the pre-trained semantic branch and completion branch to extract the semantic features and geometric features of the target point cloud data to be completed respectively, realizing decoupled learning of features; in addition, the pre-trained semantic completion branch is used to effectively fuse the semantic features and geometric features in the bird's-eye view perspective. After feature representation decomposition and then feature fusion, the semantic features and geometric features complement each other. It has strong feature representation ability while having simple calculation, which is beneficial to restoring semantic scene details. Moreover, compared with the dense feature fusion in 3D space, the bird's-eye view fusion is more convenient and efficient; the semantic scene completion method proposed by the present invention is simple, efficient, and has a fast calculation speed. It can realize real-time semantic scene completion of the input point cloud, and has good robustness to problems such as target occlusion and fast movement in the actual 3D environment, and has advantages such as good restoration of semantic scene details and a more realistic scene after completion.

[0078] In this embodiment, the semantic completion branch in Step S4 includes an adaptive feature fusion module (ARF), and the adaptive feature fusion module is used to fuse the semantic features in the bird's-eye view perspective and the geometric features in the bird's-eye view perspective. The said Step S4 includes the following steps:

[0079] Step S4a: Calculate the sparse voxel indices corresponding to the semantic features, generate bird's-eye view indices through the sparse voxel indices, use an aggregation function to map the semantic features to sparse bird's-eye view features, and finally generate the semantic features in the bird's-eye view according to the sparse bird's-eye view features and the corresponding bird's-eye view indices;

[0080] Step S4b: Perform max pooling on the geometric features in the height dimension to obtain the geometric features in the bird's-eye view;

[0081] Step S4c: Input the semantic features in the bird's-eye view and the geometric features in the bird's-eye view into the adaptive feature fusion module for feature fusion to obtain the features output by the adaptive feature fusion module, and then perform multiple upsampling operations on the features output by the adaptive feature fusion module through the decoder to obtain the semantic scene completion result, the semantic scene completion result (L, H, W) is the size of the 3D voxelization space, and C is the semantic category.

[0082] Semantic context and geometric structure complement each other and are crucial for the SSC task. It is easy to restore geometric details according to semantics, while the complete geometric shape helps to identify semantic categories. Explicitly separating feature representations can promote and accelerate the learning process of semantic context and geometric structure. We propose an adaptive feature fusion module in the semantic completion branch, which can obtain effective clues from semantic / geometric features and fully fuse semantic context and geometric details, greatly enhancing the feature representation ability.

[0083] In this embodiment, the semantic branch in step S2 includes N semantic feature extraction modules;

[0084] The completion branch in step S3 includes N geometric feature extraction modules; the N geometric feature extraction modules correspond to the N semantic feature extraction modules respectively;

[0085] The semantic completion branch in step S4 includes N adaptive feature fusion modules; the N adaptive feature fusion modules correspond to the N semantic feature extraction modules respectively; the N adaptive feature fusion modules correspond to the N geometric feature extraction modules respectively; the semantic features and geometric features extracted by the semantic feature extraction module and the corresponding geometric feature extraction module are mapped to the bird's-eye view and then input into the corresponding adaptive feature fusion module for feature fusion; after being mapped to the bird's-eye view, the semantic features and geometric features become 2D semantic features and geometric features.

[0086] The N adaptive feature fusion modules are connected in series. Each adaptive feature fusion module outputs a stage fusion feature. The first adaptive feature fusion module outputs the first stage fusion feature, and then transfers the first stage fusion feature to the next adaptive feature fusion module connected in series with it, where it is fused with the semantic feature and geometric feature in the bird's-eye view input to the next adaptive feature fusion module to output the next stage fusion feature. This transfer is carried out one by one until the last adaptive feature fusion module performs the final feature fusion.

[0087] Each adaptive feature fusion module outputs a stage fusion feature, denoted by F prev which represents the stage fusion feature output by the previous adaptive feature fusion module, and transfers F prev to the next adaptive feature fusion module connected in series with the previous one. The semantic feature and geometric feature mapped to the bird's-eye view input to the next adaptive feature module are denoted by F sem and F com respectively. First, their respective channel attention weights are calculated, then their respective channel attention weights are multiplied by their corresponding features, and then the results of their multiplication are added and passed through a 1×1 convolution to obtain the fused result.

[0088] As Figure 3 shown, the stage fusion feature output by the next adaptive feature module can be expressed as:

[0089] F f =φ{σ[MLP(AvgPool(F prev )]·F prev

[0090] +σ[MLP(AvgPool(F sem )]·F sem

[0091] +σ[MLP(AvgPool(F com )]·F com}

[0092] F f represents the stage fusion feature output by the next adaptive feature module; σ is the Sigmoid function that maps the input to the range (0,1); AvgPool is global average pooling; MLP is a multi-layer perceptron; Φ is a 1×1 convolutional layer;

[0093] N semantic feature extraction modules and geometric feature extraction modules are designed to extract semantic features and geometric features multiple times respectively, and an equal number of adaptive feature fusion modules are designed to perform multiple feature fusions on the semantically and geometrically features extracted in batches, further enhancing the feature expression ability. The semantic scene detail ability is excellent. When compared with the visualization results of SSC-SA and JS3CNet, the method proposed by the present invention has better completion effects on moving objects and flat objects, and the complemented objects are clearer, more complete and more in line with reality.

[0094] Specifically, the semantic branch in step S2 further includes a voxelization layer. There are 3 semantic feature extraction modules, and each semantic feature extraction module includes 1 sparse coding block, as Figure 2 shown, specifically the first sparse coding block, the second sparse coding block and the third sparse coding block; there is 1 voxelization layer;

[0095] Step S2 includes the following steps:

[0096] Step S21: Input the point cloud data of the target to be completed into the voxelization layer for voxelization to obtain voxel features;

[0097] The specific steps of voxelization are as follows:

[0098] Let P represent the point cloud data of the target to be completed, and p i =(x i , y i , z i ) represent a point in the point cloud data of the target to be completed. Its voxel index where s is the resolution of voxelization, is the floor operation;

[0099]

[0100] represents the voxel feature of the m-th non-empty voxel with voxel index V m ; R f represents a fully connected layer for reducing the dimension of the input feature to 64; MLP represents a multi-layer perceptron for encoding the input point cloud feature; A f is an aggregation function, usually an average pooling function, for aggregating the features of all points belonging to the same voxel; f p represents a point feature with a dimension of 7, including the coordinates of the point cloud (3 dimensions), the offset vector of the coordinates of each point in the point cloud from the center of the voxel where it is located (3 dimensions), and the radar point cloud reflection intensity (1 dimension); V p represents the index of the voxel where point p is located;

[0101] Step S22: Input the voxel features into the first sparse coding block first, and then output the first semantic feature and the first voxel feature. The first voxel feature is then input into the second sparse coding block, and the second semantic feature and the second voxel feature are output. The second voxel feature is finally input into the third sparse coding block, and the third semantic feature is output. In this embodiment, each sparse coding block consists of a sparse convolutional residual block and a sparse geometric feature extraction module. The sparse convolutional residual block halves the resolution of the input voxel features, and the sparse geometric feature extraction module uses different-scale sparse mapping and attention selection mechanisms to enhance the geometric characteristics of the voxel features. Since different scales are set in the sparse geometric feature extraction module, the semantic features extracted by the semantic branch are multi-scale sparse semantic features, and the multi-scale sparse semantic features can be expressed as (F V, F s,1, F s,2, F s,3 ). F V represents the voxel features, F s,1 is the first semantic feature, F s,2 is the second semantic feature, and F s,3 is the third semantic feature; the scales between the features F s,1 , F s,2 , and F s,3 are not the same.

[0102] In step S3, the completion branch further includes an input layer, and there are 3 geometric feature extraction modules. Each geometric feature extraction module includes 1 dense residual block, as shown in Figure 2 , specifically the first dense residual block, the second dense residual block, and the third dense residual block;

[0103] Step S3 includes the following steps:

[0104] Step S31: Calculate the occupancy voxels O V of the target point cloud data to be completed; the occupancy voxels are 0 / 1 binary values, generated from the point cloud. If a voxel contains points, it is 1, otherwise it is 0;

[0105] Step S32: Input the occupancy voxels into the input layer. After processing by the input layer to increase the receptive field, input them into the first dense residual block to output the first geometric feature and the first occupancy voxels. The first occupancy voxels are then input into the second dense residual block to output the second geometric feature and the second occupancy voxels. The second occupancy voxels are finally input into the third dense residual block to output the third geometric feature; the input layer is a 7×7×7 3D dense convolution, and each dense residual block consists of a 3×3×3 3D dense convolution. Compared with the multi-scale sparse semantic features, a corresponding multi-scale dense geometric feature (O V, F c,1, F c,2, F c,3 ) is designed for subsequent feature fusion, Fc,1 is the first geometric feature, F c,2 is the second geometric feature, F c,3 is the third geometric feature; where feature F s,1 and F c,1 have the same scale, F s,2 and F c,2 have the same scale, F s,3 and F c,3 have the same scale.

[0106] The design of multi-scale sparse semantic features and multi-scale dense geometric features enables higher extraction accuracy of semantic and geometric features, stronger feature expression ability, and good robustness to problems such as target occlusion and fast movement in the actual 3D environment.

[0107] In step S4, the semantic completion branch includes a first adaptive feature fusion module, a second adaptive feature fusion module, and a third adaptive feature fusion module; in this embodiment, the semantic completion branch is a 2D U-Net network, which includes an encoder and a decoder. The encoder consists of adaptive feature fusion modules, and the encoder hierarchically fuses 2D semantic features and geometric features.

[0108] Step S4 includes the following steps:

[0109] Step S41: Map the target point cloud data to be completed to the bird's-eye view perspective to obtain the bird's-eye view feature of the target to be completed;

[0110] Step S42: After the first semantic feature and the first geometric feature are respectively mapped to the bird's-eye view perspective, they are input into the first adaptive feature fusion module together with the bird's-eye view feature of the target to be completed for feature fusion to obtain the first-stage fusion feature. After the second semantic feature and the second geometric feature are respectively mapped to the bird's-eye view perspective, they are input into the second adaptive feature fusion module together with the first-stage fusion feature for feature fusion to obtain the second-stage fusion feature. After the third semantic feature and the third geometric feature are respectively mapped to the bird's-eye view perspective, they are input into the third adaptive feature fusion module together with the second-stage fusion feature for feature fusion to obtain the third-stage fusion feature; the fusion features of the three stages are passed through the decoder to obtain the semantic scene completion result, as Figure 1 shown.

[0111] The design of the method of the present invention is positioned as lightweight. Three sparse coding blocks are used to encode semantic features, and three lightweight dense residual blocks are used to obtain geometric features. A semantic completion branch (BEV fusion network) is used to fuse semantic / geometric features; this method has both a lightweight design and strong expression ability. When used, this method has low latency, can run in real time, and has good generalization, achieving state-of-the-art performance on the SemanticKITTI test set.

[0112] In this embodiment, pre-training is performed on the semantic branch in step S2, the completion branch in step S3, and the semantic completion branch in step S4. The pre-training steps include:

[0113] Step S51: Use the server to obtain the training data of the target point cloud to be completed;

[0114] Step S52: Use the semantic branch to extract the semantic features of the training data of the target point cloud to be completed, and perform semantic segmentation supervision to promote the learning of semantic context. The semantic supervision loss function consists of Lovasz loss and multi-class cross-entropy loss; specifically, semantic segmentation supervision uses a lightweight multi-layer perceptron as an auxiliary head, outputs the prediction results of this stage, and calculates the loss with the labels of the corresponding scale;

[0115] The semantic supervision loss function can be expressed as:

[0116]

[0117] L lovasz,i and L ce,i respectively represent the Lovasz loss and multi-class cross-entropy loss of the i-th semantic stage of semantic supervision;

[0118] Step S53: Use the completion branch to extract the geometric features of the training data of the target point cloud to be completed, and perform scene completion supervision to promote the learning of geometric information. The geometric supervision loss function consists of Lovasz loss and binary cross-entropy loss; specifically, scene completion supervision uses a lightweight multi-layer perceptron as an auxiliary head, outputs the prediction results of this stage, and calculates the loss with the labels of the corresponding scale;

[0119] The geometric supervision loss function can be expressed as:

[0120]

[0121] L lovasz,i and L bce,i respectively represent the Lovasz loss and binary cross-entropy loss of the i-th completion stage of geometric supervision;

[0122] Step S54: After mapping the semantic features and geometric features to the bird's-eye view perspective respectively, input them into the semantic completion branch for feature fusion to obtain the semantic scene completion result, and perform main supervision on the semantic scene completion result. The semantic completion supervision loss function consists of Lovasz loss and multi-class cross-entropy loss;

[0123] The semantic completion supervision loss function can be expressed as:

[0124] L bev = L lovasz + Lce

[0125] L lovasz and L ce respectively represent the Lovasz loss and the multi-class cross-entropy loss for semantic completion supervision;

[0126] The Lovasz loss in step S52, step S53 and step S54 is specifically:

[0127]

[0128] where J is the Lovasz extended version of IoU, and e(c) is the error vector of class c; the cross-entropy loss is specifically: where y i is the predicted value, is the true value;

[0129] Step S55: Use the server for network training and perform multi-task training in an end-to-end manner; the loss function L total is the weighted sum of the supervision loss functions in step S52, step S53 and step S54;

[0130] The loss function L total can be expressed as:

[0131] L total = 3·L bev + L s + L c

[0132] Step S56: Use the server to optimize the loss function, obtain the locally optimal network parameters, and obtain the pre-trained semantic branch, completion branch and semantic completion branch.

[0133] For the two independent branches of the semantic branch and the completion branch, and for the application layer, we apply hierarchical supervision to promote the representation learning process. The point cloud semantic scene completion method based on feature representation decomposition and bird's-eye view fusion provided by the present invention designs a semantic branch and a completion branch based on feature representation decomposition to extract semantic features and geometric features respectively to accelerate network convergence and the feature characterization ability of the algorithm. In addition, an adaptive fusion module and a semantic completion branch are designed based on bird's-eye view fusion to effectively fuse semantic features and geometric features. The method (SSC-RS) proposed by the present invention is fast, can perform semantic scene completion on the input point cloud in real time, and has good robustness to problems such as target occlusion and fast movement in the actual 3D environment, and achieves state-of-the-art performance on the large-scale dataset SemanticKITTI test set.

[0134] Figure 4Comparison chart of visualization results of the method proposed in this embodiment and other advanced methods (SSC-SA, JS3CNet); Input the radar point cloud into the pre-trained semantic scene completion network (SSC-SA, JS3CNet and SSC-RS proposed in the present invention), output the semantic completion result, and visualize it; As Figure 4 shown, the algorithm proposed in the present invention has better completion effects on moving objects and flat objects.

[0135] Figure 5 Comparison results of the method proposed in this embodiment and other advanced methods on the SemanticKITTI test set; Input the radar point cloud in the SemanticKITTI test set into the pre-trained semantic scene completion network (LMCNet, Local-DIFs, JS3CNet, S3CNet, UDNet, SSA-SC and SSC-RS proposed in the present invention), output the semantic completion result, and upload the result to the server to calculate various semantic completion metrics, including the completion result (intersection over union IoU) and the semantic completion result (average IoU of each category, mIoU); As Figure 5 shown, the algorithm proposed in the present invention has the most advanced performance in terms of the completion result, and can run in real time (16.7fps), ranking first in terms of the completion metric IoU and second in terms of the semantic completion metric mIoU among the published works.

[0136] Although the present invention has been described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the present invention. Therefore, it should be understood that many modifications can be made to the exemplary embodiments, and other arrangements can be designed, as long as they do not depart from the spirit and scope of the present invention as defined by the appended claims. It should be understood that different dependent claims and features in the present application can be combined in a manner different from that described in the original claims. It should also be understood that the features described in connection with a single embodiment can be used in other embodiments.

Claims

1. A semantic scene completion method based on feature representation decomposition and bird's-eye view fusion, comprising the following steps: Step S1: Obtain the target point cloud data to be completed; Step S2: Use the pre-trained semantic branch to extract the semantic features of the target point cloud data to be completed; Step S3: Use the pre-trained completion branch to extract the geometric features of the target point cloud data to be completed; Step S4: The semantic features and geometric features are respectively mapped to the bird's-eye view perspective and then input into the pre-trained semantic completion branch for feature fusion to obtain the semantic scene completion result; The pre-trained semantic branch, pre-trained completion branch, and pre-trained semantic completion branch use the server to optimize the loss function L total obtained; Loss function L total can be expressed as: L total = 3·L bev + L s + L c L bev is the semantic completion supervision loss function, which can be expressed as: L bev = L lovasz + L ce L lovasz and L ce respectively represent the Lovasz loss and the multi-class cross-entropy loss for semantic completion supervision; L s is the semantic supervision loss function, which can be expressed as: L lovasz,i and L ce,i respectively represent the i-th Lovasz loss and the i-th multi-class cross-entropy loss for semantic supervision; L c is the geometric supervision loss function and can be expressed as: L lovasz,i and L bce,i respectively represent the i-th Lovasz loss and the i-th binary cross-entropy loss for geometric supervision.

2. The semantic scene completion method based on feature representation decomposition and bird's-eye view fusion according to claim 1, wherein: In Step S4, the semantic completion branch includes an adaptive feature fusion module, which is used to fuse the semantic features in the bird's-eye view perspective and the geometric features in the bird's-eye view perspective.

3. The semantic scene completion method based on feature representation decomposition and bird's-eye view fusion according to claim 2, wherein: In Step S2, the semantic branch includes N semantic feature extraction modules; In Step S3, the completion branch includes N geometric feature extraction modules; The N geometric feature extraction modules respectively correspond to the N semantic feature extraction modules; In Step S4, the semantic completion branch includes N adaptive feature fusion modules; The N adaptive feature fusion modules respectively correspond to the N semantic feature extraction modules; the N adaptive feature fusion modules respectively correspond to the N geometric feature extraction modules; the semantic features and geometric features extracted by the semantic feature extraction module and its corresponding geometric feature extraction module are respectively mapped to the bird's-eye view perspective and then input into the corresponding adaptive feature fusion module for feature fusion; The N adaptive feature fusion modules are connected in series. Each adaptive feature fusion module outputs a stage fusion feature. The first adaptive feature fusion module outputs the first stage fusion feature, and then the first stage fusion feature is passed to the next adaptive feature fusion module in series with it, and is fused with the semantic features in the bird's-eye view perspective and the geometric features in the bird's-eye view perspective input in the next adaptive feature fusion module to output the next stage fusion feature, and is passed one by one until the last adaptive feature fusion module performs the final feature fusion.

4. The semantic scene completion method based on feature representation decomposition and bird's-eye view fusion according to claim 3, wherein: In Step S2, the semantic branch includes a voxelization layer, and the N semantic feature extraction modules include a first sparse coding block, a second sparse coding block, and a third sparse coding block; Step S2 includes the following steps: Step S21: Input the target point cloud data to be completed into the voxelization layer for voxelization to obtain voxel features; Step S22: Input the voxel features into the first sparse coding block first and then output the first semantic feature and the first voxel feature. The first voxel feature is then input into the second sparse coding block and then output the second semantic feature and the second voxel feature. The second voxel feature is finally input into the third sparse coding block and then output the third semantic feature; In Step S3, the completion branch includes an input layer, and the N geometric feature extraction modules include a first dense residual block, a second dense residual block, and a third dense residual block; Step S3 includes the following steps: Step S31: Calculate the occupied voxels of the target point cloud data to be completed; Step S32: Input the occupied voxel into the input layer. After processing by the input layer to increase the receptive field, it is then input into the first dense residual block to output the first geometric feature and the first occupied voxel. The first occupied voxel is then input into the second dense residual block to output the second geometric feature and the second occupied voxel. The second occupied voxel is finally input into the third dense residual block to output the third geometric feature; In step S4, the semantic completion branch includes a first adaptive feature fusion module, a second adaptive feature fusion module, and a third adaptive feature fusion module; Step S4 includes the following steps: Step S41: Map the target point cloud data to be completed to the bird's-eye view perspective to obtain the target bird's-eye view feature to be completed; Step S42: The first semantic feature and the first geometric feature are respectively mapped to the bird's-eye view perspective and then input into the first adaptive feature fusion module together with the target bird's-eye view feature to be completed for feature fusion to obtain the first-stage fusion feature. The second semantic feature and the second geometric feature are respectively mapped to the bird's-eye view perspective and then input into the second adaptive feature fusion module together with the first-stage fusion feature for feature fusion to obtain the second-stage fusion feature. The third semantic feature and the third geometric feature are respectively mapped to the bird's-eye view perspective and then input into the third adaptive feature fusion module together with the second-stage fusion feature for feature fusion to obtain the third-stage fusion feature. The fusion features of the three stages are passed through the decoder to obtain the semantic scene completion result.

5. A semantic scene completion method based on feature representation decomposition and bird's-eye view fusion according to claim 4, characterized in that, The specific steps of voxelization are as follows: Let \(P\) represent the target point cloud data to be completed, and \(p\) i \(=\)(x i , y i , z i ) represents a point in the target point cloud data to be completed, and its voxel index where \(s\) is the resolution of voxelization, is the floor operation; Denote the voxel feature of the m-th non-empty voxel with voxel index V m ; R f Denote the fully connected layer; MLP denotes the multi-layer perceptron; A f is an aggregation function; f p Denote the point feature, V p denotes the voxel index where the point p is located.

6. A semantic scene completion method based on feature representation decomposition and bird's-eye view fusion according to claim 2, characterized in that: The said step S4 includes the following steps: Step S4a: Calculate the sparse voxel index corresponding to the semantic feature, calculate the bird's-eye view index through the sparse voxel index, use the aggregation function to map the semantic feature to the sparse bird's-eye view feature, and finally generate the semantic feature in the bird's-eye view perspective according to the sparse bird's-eye view feature and the corresponding bird's-eye view index; Step S4b: Perform max pooling on the geometric feature in the height dimension to obtain the geometric feature in the bird's-eye view perspective; Step S4c: Input the semantic features and geometric features in the bird's-eye view into the adaptive feature fusion module for feature fusion to obtain the features output by the adaptive feature fusion module, and then perform multiple upsampling operations on the features output by the adaptive feature fusion module through the decoder to obtain the semantic scene completion result, where the semantic scene completion result (L, H, W) is the size of the 3D voxelization space, and C is the semantic category.

7. A semantic scene completion method based on feature representation decomposition and bird's-eye view fusion according to claim 3, characterized in that, In the step S4, each adaptive feature fusion module outputs a stage fusion feature, denoted as F prev which represents the stage fusion feature output by the previous adaptive feature fusion module, and inputs F prev to the next adaptive feature fusion module connected in series with the previous adaptive feature fusion module. The semantic feature and geometric feature mapped to the bird's-eye view perspective and input to the next adaptive feature module are denoted as F sem and F com respectively; The stage fusion feature output by the next adaptive feature module can be expressed as: F f = φ{σ[MLP(AvgPool(F prev )]·F prev +σ[MLP(AvgPool(F sem )]·F sem +σ[MLP(AvgPool(F com )]·F com} F f represents the stage fusion feature output by the next adaptive feature module; σ is the Sigmoid function that maps the input to the range between (0, 1); AvgPool is the global average pooling; MLP is the multi-layer perceptron; Φ is the 1×1 convolutional layer.

8. A semantic scene completion method based on feature representation decomposition and bird's-eye view fusion according to any one of claims 1-7, characterized in that, Pre-train the semantic branch in step S2, the completion branch in step S3, and the semantic completion branch in step S4. The pre-training steps include: Step S51: Use the server to obtain the target point cloud training data to be completed; Step S52: Use the semantic branch to extract the semantic features of the target point cloud training data to be completed, and perform semantic segmentation supervision to promote the learning of semantic context; Step S53: Use the completion branch to extract the geometric features of the target point cloud training data to be completed, and perform scene completion supervision to promote the learning of geometric information; Step S54: The semantic feature and the geometric feature are respectively mapped to the bird's-eye view perspective and then input into the semantic completion branch for feature fusion to obtain the semantic scene completion result, and perform main supervision on the semantic scene completion result; Step S55: Use the server to perform network training, and perform multi-task training in an end-to-end manner; Step S56: Use the server to optimize the loss function, obtain the local optimal network parameters, and obtain the pre-trained semantic branch, completion branch, and semantic completion branch.

9. A semantic scene completion method based on feature representation decomposition and bird's-eye view fusion according to claim 1, characterized in that: The specific Lovasz loss is: J is the Lovasz extension version of IoU, and e(c) is the error vector for class c.

Citation Information

Patent Citations

  • 3D point cloud semantic segmentation method under bird's-eye view coding view angle

    CN111862101A

  • Semantic scene completion method and system based on point cloud-voxel aggregation network model

    CN113850270A