A 3D Scene Semantic Completion Method Based on Implicit Representation

Through the three-dimensional scene semantic completion method based on implicit representation, an implicit codec network is built and the network is trained, which solves the problem of difficulty in adapting to different resolution requirements in the existing technology, and efficient semantic completion of arbitrary resolutions is achieved, avoiding the generation of artifacts.

CN114782603BActive Publication Date: 2025-06-13NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210481019.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-05
Publication Date
2025-06-13
Estimated Expiration
2042-05-05

AI Technical Summary

Technical Problem

The existing three-dimensional scene semantic completion technology is difficult to dynamically adapt to the demand for geometric model variable particle size in different occasions, and the existing methods are inconvenient when adjusting the network structure and retraining the network, and artifacts are easily generated when high-resolution voxel completion.

Method used

The three-dimensional scene semantic completion method based on implicit representation is adopted. By building an implicit codec network, the semantic label of implicit representation is generated, and the network is trained through the stochastic gradient descent algorithm to achieve voxel semantic completion of any resolution.

Benefits of technology

It realizes semantic completion of three-dimensional scenes that meet different resolution requirements without adjusting the network structure and retraining, avoids the generation of artifacts and improves the accuracy and flexibility of the completion results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114782603B_ABST
    Figure CN114782603B_ABST
Patent Text Reader

Abstract

The present invention discloses a three-dimensional scene semantic completion method based on implicit representation, which includes the following steps: Step 1, construction of an implicit encoding and decoding network: construct an implicit representation encoding network for generating implicit representation and an implicit representation decoding network for decoding the implicit representation; Step 2, generation of implicit representation semantic labels; Step 3, training of the implicit encoding and decoding network: according to the stochastic gradient descent algorithm, use the implicit representation semantic labels to train the implicit encoding and decoding network; Step 4, generation of voxel semantic completion results. The present invention decouples the direct connection between the granularity and the network structure, achieving the effect of meeting the variable granularity requirements without adjusting the network and retraining. The present invention can realize scene completion with arbitrary granularity and can meet the requirements of different application scenarios for the geometric variable granularity of the scene model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for semantic completion of 3D scenes, in particular to a method for semantic completion of 3D scenes based on implicit representation. Background Art

[0002] Semantic completion of 3D scenes aims to recover complete scene geometry from incomplete scene geometry and use semantic information to guide the completion process. Currently, this technology has been widely applied to various 3D scene modeling occasions such as intelligent robots, virtual reality, and augmented reality. However, since different occasions have different requirements for the granularity of the model geometry, how to dynamically adapt to the requirements of different occasions for variable granularity of geometric models has become a difficult problem in semantic completion of 3D scenes.

[0003] According to the representation methods of 3D models of scenes, scene semantic completion techniques can be divided into completion methods based on voxel representation and completion methods based on point cloud representation. The methods based on voxel representation meet the requirements of variable granularity by outputting voxel results at different resolutions, such as in Document 1. Song S, Yu F, Zeng A, et al. Semantic scene completion from a single depth image[C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2017:1746-1754.; Document 2 Wang Y, Tan D J, Navab N, et al. Forknet: Multi-branch volumetric semantic completion from a single depth image[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2019:8608-8617.; Document 3 Li J, Liu Y, Gong D, et al. Rgbd based dimensional decomposition residual network for 3d semantic scene completion[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2019:7693-7702.; Document 4 Li J, Han K, Wang P, et al. Anisotropic convolutional networks for 3d semantic scene completion[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2020:3351-3359.; Document 5 Cai Y, Chen X, Zhang C, et al.Semantic scene completion via integrating instances and scene in-the-loop[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2021: 324-333. By encoding the scene into feature voxels and using 3D channel convolution to decode the per-voxel class scores, the voxel completion result is finally obtained by finding the index value of the maximum per-voxel class score. Due to the high memory occupancy of voxel representation, these methods can only obtain voxel completion results with relatively low resolution and cannot predict high-resolution voxel completion results. Literature [6] Zhang P, Liu W, Lei Y, et al. Cascaded context pyramid for full-resolution 3D semantic scene completion[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2019: 7801-7810. adopts a grouped convolution method to obtain higher-resolution voxel completion results by reducing network parameters. However, this method cannot directly adapt to the voxel completion requirements of low resolution. On the one hand, the final resolution of voxel completion directly depends on the resolution of the feature voxels, and the resolution of the feature voxels depends on the network structure. Therefore, adapting to different resolutions means adjusting the network structure and retraining the network, which is very inconvenient. On the other hand, directly downsampling high-resolution voxels will produce serious artifacts because certain geometric features are inevitably lost during the downsampling process.

[0004] The completion method based on point cloud representation meets the variable granularity requirements by generating different numbers of point clouds. For example, in reference 7. Zhang S, Li S, Hao A, et al. Point Cloud Semantic Scene Completion from RGB-D Images[C] / / Proceedings of the AAAI Conference on Artificial Intelligence. 2021, 35(4): 3385-3393. It restores a certain number of point clouds by estimating the geometric features of the missing point clouds, and stitches the predicted point clouds and the original point clouds to obtain the point cloud completion result. However, this method also requires adjusting the network structure in advance to determine the number of final output points. Therefore, to meet the variable granularity requirements, it is necessary to readjust the network structure and retrain each time, which is extremely inconvenient for some downstream applications. Summary of the Invention

[0005] Object of the Invention: The technical problem to be solved by the present invention is to provide a three-dimensional scene semantic completion method based on implicit representation in view of the deficiencies of the prior art.

[0006] To solve the above technical problem, the present invention discloses a three-dimensional scene semantic completion method based on implicit representation, including the following steps:

[0007] Step 1, constructing an implicit encoding and decoding network: constructing an implicit representation encoding network for generating implicit representation and an implicit representation decoding network for decoding the implicit representation;

[0008] Step 2, generating implicit representation semantic labels: For each training sample, randomly sample a certain number of points in three-dimensional space, calculate the class label of each point according to the given voxel-level label, and generate implicit representation semantic labels;

[0009] Step 3, training the implicit encoding and decoding network: According to the stochastic gradient descent algorithm, use the implicit representation semantic labels to train the implicit encoding and decoding network;

[0010] Step 4, generating the voxel semantic completion result: Input the RGB image and the depth image, use the implicit representation encoding network to calculate the implicit representation feature encoding of the scene, generate a sampling space according to the specified resolution, uniformly sample points from this sampling space, and use the implicit representation decoding network to calculate the semantic label of each point. Finally, obtain the final voxel completion result through the voxel result generation module.

[0011] The construction of the implicit encoding and decoding network described in Step 1 of the present invention includes the following steps:

[0012] Step 1-1, constructing an implicit representation encoding network

[0013] Step 1-2, construct the implicit representation decoding network

[0014] The implicit representation encoding network constructed in Step 1-1 of the present invention includes the following steps:

[0015] Step 1-1-1, two-dimensional feature extraction: construct the RGB image feature extraction network F rgb and the depth image feature extraction network F dep ;

[0016] Among them, the RGB image feature extraction network consists of a two-dimensional convolutional layer and two two-dimensional channel decomposition residual networks (Reference: Li J, Liu Y, Gong D, et al. Rgbd based dimensional decomposition residual network for 3d semantic scene completion[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2019:7693-7702.); the input channel number of the two-dimensional convolutional layer in the RGB image feature extraction network is 3;

[0017] The structure of the depth image feature extraction network is the same as that of the RGB image feature extraction network, and the input channel number of the two-dimensional convolutional layer in the depth image feature extraction network is 1;

[0018] Input the RGB image I rgb into F rgb to obtain the two-dimensional RGB feature Input the depth image I depth into F dep to obtain the two-dimensional depth feature

[0019] Step 1-1-2, 2D to 3D back-projection: Given a box-shaped space S centered at the origin of the coordinate system with length, width and height D s , W s , H s and a given feature image f with a shape of H×W×ch, where H is the height of the feature image, W is the width of the feature image, and ch is the number of channels of the feature image. Given the camera internal parameter K and the external parameter P, for a point coordinate x=(u, v) on the feature image, the feature vector of this point on the feature image is f x , calculate the coordinate X of x in the three-dimensional space X = P -1K -1 x. If X falls outside the box-shaped space S, the point is ignored. Otherwise, the feature vector at X is f X and f X = f x ;

[0020] Voxelize the three-dimensional space into voxels V of h×w×d. h, w, and d represent the number of grids in the height, width, and length directions of the voxels respectively; the length, width, and height of each grid are and calculate the index (i, j, k) of the grid to which X belongs. i, j, and k represent the index values in the height, width, and length directions respectively; for each grid, if no points fall within the grid, set the feature vector of the grid to a zero vector. If p points fall within the grid, calculate the feature vector of the grid through max pooling. The calculation method is as follows:

[0021]

[0022] where V i,j,k,m represents the value of the m-th dimension of the feature vector of the grid with index (i, j, k), f X,m represents the value of the m-th dimension of the feature vector at point X, and Φ X represents the set of all points that fall within the grid;

[0023] Two-dimensional RGB features Obtain RGB voxel features through 2D to 3D backprojection Two-dimensional depth features Obtain depth voxel features through 2D to 3D backprojection

[0024] Step 1-1-3, multi-scale three-dimensional feature completion: Construct a downsampling network F down and an upsampling network F up , the downsampling network includes a three-dimensional convolutional layer F d2 with a kernel size of 3, a stride of 2, and a padding of 1, a max pooling layer F mp , and a three-dimensional convolutional layer F fill with a kernel size of 3, a stride of 1, and a padding of 1. For a voxel feature f vox of shape 2h′×2w′×2d′×ch′, where h′, w′, d′, and ch′ are the height, width, length, and number of channels of the voxel feature respectively, the calculation process of the downsampling network is as follows:

[0025] f′ vox = F fill (cat(F d2 (f vox ),F mp (fvox )))

[0026] Among them, f' vox is the calculated voxel feature, with a shape of h'×w'×d'×ch', and cat(·,·) represents concatenating two voxel features in the channel dimension;

[0027] The upsampling network includes a 3D transposed convolution layer F with a kernel size of 2, a stride of 2, and a padding of 0 u2 , a trilinear upsampling layer F up , with an upsampling size of 2 times, and a 3D convolution layer F with a kernel size of 3, a stride of 1, and a padding of 1 fill For a voxel feature f with a shape of h'×w'×d'×ch' vox , the calculation process of the upsampling network is as follows:

[0028] f' vox = F fill (cat(F u2 (f vox ), F up (f vox )))

[0029] Among them, f' vox is the calculated voxel feature, with a shape of 2h'×2w'×2d'×ch';

[0030] Construct a feature completion network F u and calculate the feature pyramid. The feature completion network includes two cascaded downsampling networks and two cascaded upsampling networks. For a voxel feature f with a shape of 4h'×4w'×4d'×ch' vox , the calculation method of the feature completion network is as follows:

[0031]

[0032]

[0033]

[0034] Among them, the shape of the low-resolution voxel feature f low is h'×w'×d'×ch', the shape of the medium-resolution voxel feature f mid is 2h'×2w'×2d'×ch', the shape of the high-resolution voxel feature f high is 4h'×4w'×4d'×ch'. Upsample f low and f mid to f by the nearest neighbor upsampling methodhigh The shape of, finally, by splicing f on the channel low , f mid , f high to obtain the feature pyramid f pyr ;

[0035] RGB voxel features After passing through the feature completion network, the RGB feature pyramid is obtained Depth voxel features After passing through the feature completion network, the depth feature pyramid is obtained

[0036] Step 1-1-4, fusion: fuse the RGB feature pyramid and the depth feature pyramid by element-wise addition to obtain the implicit representation feature encoding where, f imp has a shape of h×w×d×ch p , ch p is the number of channels of f imp .

[0037] The construction of the implicit representation decoding network described in step 1-2 of the present invention includes the following steps:

[0038] Step 1-2-1, feature sampling;

[0039] Step 1-2-2, vector splicing;

[0040] Step 1-2-3, implicit decoding.

[0041] The method of feature sampling described in step 1-2-1 of the present invention includes:

[0042] For any point X in the box-shaped space S and the implicit representation feature encoding f imp , calculate the grid index (i, j, k) to which the point belongs, and then use the following formula to calculate the three-dimensional coordinate values of the grid vertices with index values (i, j, k), (i, j, k+1), (i, j+1, k), (i+1, j, k), (i+1, j+1, k), (i+1, j, k+1), (i, j+1, k+1), (i+1, j+1, k+1):

[0043]

[0044] where, a∈{i, i+1}, b∈{j, j+1}, c∈{k, k+1}, X a,b,c represents the three-dimensional coordinate values of the grid vertices with index values a, b, c; take out the feature vector with index value (a, b, c) from the implicit representation feature encoding Then the feature sampling calculation method is:

[0045]

[0046] d a,b,c = ||X - X a,b,c || 2

[0047]

[0048] In step 1-2-2 of the present invention, the method of vector splicing includes: splicing the sampled feature f sam and X to obtain the query vector f query .

[0049] In step 1-2-3 of the present invention, the method of implicit decoding includes: building a local decoder, and inputting the query vector f query into the local decoder (the local decoder uses the decoding network described in the literature Peng S, Niemeyer M, Mescheder L, et al. Convolutional occupancy networks[C] / / European Conference on ComputerVision. Springer, Cham, 2020:523-540.) to obtain the confidence score of the point category

[0050] The generation of implicit representation semantic labels in step 2 of the present invention includes the following steps:

[0051] Step 2-1, randomly sampling points in the box space: randomly collecting N query points in the box space

[0052] Step 2-2, semantic label generation: for a given voxel label Y ∈ {1, 2,..., K} h×w×d , where K is the number of categories, and the coordinates X of a query point in a three-dimensional space sam , convert Y into one-hot code form Y one-h ∈ R h×w×d×K , calculate the index of the voxel grid to which the point belongs, and set the semantic label y of the point to the one-hot code value under this index value, and at the same time generate the geometric label y g of the point, and the method includes:

[0053]

[0054] where y[1] represents the first element of y;

[0055] Step 2-3: For all the collected points, repeat Steps 2-1 to 2-2 to complete the generation of semantic labels for one training sample.

[0056] Step 2-4: For each training sample in the New York dataset (the New York dataset is the dataset described in the literature Silberman N, Hoiem D, Kohli P, et al. Indoor segmentation and support inference from rgbd images[C] / / European conference on computer vision. Springer, Berlin, Heidelberg, 2012:746-760.), repeat Steps 2-1 to 2-3 to complete the generation of implicit representation semantic labels.

[0057] The training of the implicit representation encoding and decoding network in Step 3 of the present invention includes the following steps:

[0058] Step 3-1: Construct a loss calculation module, including a semantic loss module and a geometric loss module. Among them, the loss calculation method in the semantic loss module is as follows:

[0059]

[0060] where w cl is the weight coefficient of each category, is the confidence score of the cl-th category decoded by the implicit decoding network, and y cl is the one-hot code value of the cl-th category of the semantic label generated in Step 2-4;

[0061] The loss calculation method in the geometric loss module is as follows:

[0062]

[0063] where y g is the geometric label generated in Step 2-4,

[0064] Finally, the total loss term of the loss calculation module is as follows:

[0065] L = L sem + 2 * L geo ;

[0066] Step 3-2: Use the stochastic gradient descent optimizer (the stochastic gradient descent optimizer adopts the one described in the literature Bottou L. Stochastic gradient descent tricks [M] / / Neural networks: Tricks of the trade. Springer, Berlin, Heidelberg, 2012: 421-436), and train for 200 rounds on the New York dataset at a learning rate of 10 -2 to complete the training of the implicit representation encoding and decoding network.

[0067] The generation of the voxel semantic completion result in step 4 of the present invention includes the following steps:

[0068] Step 4-1: Generate spatial query points. According to the specified resolution H u , W u , D u , divide the box-shaped space into a voxel grid of H u ×W u ×D u , calculate the three-dimensional point coordinates of the center of each grid, and generate the final query point set;

[0069] Step 4-2: Generate the scene implicit representation. Input the RGB image and the depth image into the implicit representation encoding network to obtain the implicit representation feature encoding of the scene;

[0070] Step 4-3: Generate point labels. Input the implicit representation feature encoding and the query point set into the implicit representation decoding network to obtain the confidence score of each point, and take the subscript with the maximum confidence score as the class label of this point;

[0071] Step 4-4: Voxel result generation. Set K colors for K categories. For each query point, if the class label of this query point is 1, ignore it. Otherwise, generate a bounding box with a side length of centered at this point, and paint the 6 faces of this box with the color of this category to generate the final voxel completion result.

[0072] Beneficial effects:

[0073] The present invention proposes a scene semantic completion method based on implicit representation. By decoupling the direct connection between the granularity and the network structure, the effect of meeting the variable granularity requirements without adjusting the network and retraining is achieved. The present invention can realize scene completion with arbitrary granularity and can meet the requirements of different application scenarios for the geometric variable granularity of the scene model. Description of the drawings

[0074] The following further specifically describes the present invention in conjunction with the accompanying drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.

[0075] Figure 1 It is a schematic diagram of the processing flow of the present invention.

[0076] Figure 2 It is a schematic diagram of the input RGB image.

[0077] Figure 3 It is a schematic diagram of the input depth image.

[0078] Figure 4 It is a schematic diagram of the voxel completion result with a resolution of 60×36×60.

[0079] Figure 5 It is a schematic diagram of the voxel completion result with a resolution of 240×144×240. Specific Embodiments

[0080] As Figure 1 shown, a three-dimensional scene semantic completion method based on implicit representation disclosed by the present invention is specifically implemented according to the following steps:

[0081] Step 1, building an implicit encoding and decoding network: Build an implicit representation encoding network for generating implicit representations and an implicit representation decoding network for decoding implicit representations.

[0082] Step 2, generating implicit representation semantic labels: For each training sample, randomly sample a certain number of points in three-dimensional space, calculate the category label of each point according to the voxel-level label, and generate implicit representation semantic labels.

[0083] Step 3, training the implicit encoding and decoding network: According to the stochastic gradient descent algorithm, use the implicit representation semantic labels to train the implicit encoding and decoding network.

[0084] Step 4, generating voxel semantic completion results at any resolution: For the input RGB and depth maps, use the implicit representation encoding network to calculate the implicit representation feature encoding of the scene, specify the resolution of the final result by the user, generate a sampling space according to this resolution, uniformly sample points from this sampling space, and use the implicit representation decoding network to calculate the semantic label of each point. Finally, obtain the final voxel completion result through the voxel result generation module.

[0085] The main processes of each step are introduced below

[0086] Among them, Step 1 includes the following steps

[0087] Step 1.1, building the implicit representation encoding network

[0088] Step 1.2, construct the implicit representation decoding network

[0089] Step 1.3, for an RGB image I rgb and a depth image I dep , input them into the implicit representation encoding network to obtain the implicit representation feature encoding f imp , and then input the implicit representation feature encoding and any query point in the box space S centered at the origin of coordinates with length, width, and height D s , W s , H s into the implicit representation decoding network to obtain the class score of this point.

[0090] In actual implementation, D s = W s = H s = 2.

[0091] Step 1.1 includes the following steps:

[0092] Step 1.1.1, two-dimensional feature extraction: construct the RGB image feature extraction network F rgb and the depth image feature extraction network F dep . The RGB image feature extraction network consists of a two-dimensional convolutional layer and two two-dimensional channel decomposition residual networks. The structure of the depth image feature extraction network is the same as that of the RGB image feature extraction network. The difference is that the number of input channels of the two-dimensional convolutional layer in the depth image feature extraction network is 1, while the number of input channels of the two-dimensional convolutional layer in the RGB image feature extraction network is 3. Input the RGB image I rgb into F rgb to obtain the two-dimensional RGB feature Input the depth image I depth into F dep to obtain the two-dimensional depth feature

[0093] The two-dimensional channel decomposition residual network adopts the network described in the literature Li J, Liu Y, Gong D, et al. Rgbd based dimensional decomposition residual network for 3d semantic scene completion[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2019:7693-7702.

[0094] Step 1.1.2, 2D to 3D backprojection: Consider a given feature image f with the shape of H×W×ch, where H is the height of the feature image, W is the width of the feature image, and ch is the number of channels of the feature image. Given the camera intrinsic parameter K and extrinsic parameter P, for a point coordinate x = (u, v) on the feature image, the feature vector of this point on the feature image is f x , calculate the coordinate X of x in the three-dimensional space, X = P -1 K -1 x. If X falls outside the box-shaped space S, then this point is ignored. Otherwise, the feature vector at X is f X and f X = f x . Voxelize the three-dimensional space into a grid V of h×w×d, and the length, width, and height of each grid are and calculate the index (i, j, k) of the grid to which X belongs. For each grid, if no point falls within this grid, then set the feature vector of this grid to a zero vector. If p points fall within this grid, then the calculation method of the feature vector of this grid is as follows:

[0095]

[0096] where V i,j,k,m represents the value of the m-th dimension of the feature vector of the grid with index (i, j, k), f X,m represents the value of the m-th dimension of the feature vector at point X, and Φ X represents the set of all points that fall within this grid.

[0097] Two-dimensional RGB features Obtain RGB voxel features through 2D to 3D backprojection Two-dimensional depth features Obtain depth voxel features through 2D to 3D backprojection

[0098] In actual implementation, the shape of RGB is 480×640×3, the shape of the depth map is 480×640, h = 240, w = 144, d = 240, H = 480, W = 640, ch = 8.

[0099] As Figure 2 shown, it is a schematic diagram of the input RGB image in this embodiment. This figure shows a living room scene, where a part of the wall is blocked by a chair, and there are also self-occlusions on the chair and the table itself;

[0100] As Figure 3 shown, it is a schematic diagram of the input depth image in this embodiment. This figure is Figure 2The depth image corresponding to the RGB image has incomplete geometric information of the living room scene described by the depth map due to occlusion factors.

[0101] Step 1.1.3, multi-scale 3D feature completion, constructing a downsampling network F down and an upsampling network F up , the downsampling network contains a 3D convolutional layer F with a kernel size of 3, a stride of 2, and a padding of 1 d2 , a max pooling layer F mp , and a 3D convolutional layer F with a kernel size of 3, a stride of 1, and a padding of 1 fill , whose function is to reduce the resolution of the voxel features to For a voxel feature f with a shape of 2h′×2w′×2d′×ch′ vox , the calculation process of the downsampling network is as follows:

[0102] f′ vox = F fill (cat(F d2 (f vox ), F mp (f vox )))

[0103] where f′ vox is the calculated voxel feature, with a shape of h′×w′×d′×ch′, and cat(·,·) represents concatenating two voxel features in the channel dimension.

[0104] The upsampling network contains a 3D transposed convolutional layer F with a kernel size of 2, a stride of 2, and a padding of 0 u2 , a trilinear upsampling layer F up , with an upsampling size of 2 times, and a 3D convolutional layer F with a kernel size of 3, a stride of 1, and a padding of 1 fill , whose function is to double the resolution of the voxel features. For a voxel feature f with a shape of h′×w′×d′×ch′ vox , the calculation process of the upsampling network is as follows:

[0105] f′ vox = F fill (cat(F u2 (f vox ), F up (f vox )))

[0106] where f′ vox is the calculated voxel feature, with a shape of 2h′×2w′×2d′×ch′.

[0107] Constructing a feature completion network Fu And calculate the feature pyramid. The feature completion network includes two cascaded downsampling networks and two cascaded upsampling networks, For a voxel feature f with a shape of 4h′×4w′×4d′×ch′ vox , the calculation method of the feature completion network is as follows:

[0108]

[0109]

[0110]

[0111] Among them, the low-resolution voxel feature f low has a shape of h′×w′×d′×ch′, the medium-resolution voxel feature f mid has a shape of 2h′×2w′×2d′×ch′, the high-resolution voxel feature f hig has a shape of 4h′×4w′×4d′×ch′. Upsample f low and f mid to the shape of f high by the method of nearest neighbor upsampling. Finally, by concatenating f low , f mid , f high on the channels, the feature pyramid f pyr is obtained.

[0112] RGB voxel feature After passing through the feature completion network, the RGB feature pyramid is obtained Depth voxel feature After passing through the feature completion network, the depth feature pyramid is obtained

[0113] Step 1.1.4, Fusion: Fusion the RGB feature pyramid and the depth feature pyramid by element-wise addition to obtain the implicit representation feature encoding Among them, the shape of f imp is h×w×d×ch p .

[0114] In actual implementation, ch p = 64.

[0115] Step 1.2 includes the following steps:

[0116] Step 1.2.1, Feature Sampling. For any point X in the box-shaped space S and the implicit representation feature encoding f imp, calculate the grid index (i, j, k) to which the point belongs, and then use the following formula to calculate the three-dimensional coordinate values of the grid vertices with index values (i, j, k), (i, j, k + 1), (i + 1, j, k), (i + 1, j + 1, k), (i + 1, j, k + 1), (i, j + 1, k + 1), (i + 1, j +

[0117] 1, k + 1)

[0118]

[0119] where a ∈ {i, i + 1}, b ∈ {j, j + 1}, c ∈ {k, k + 1}, X a,b,c represents the three-dimensional coordinate values of the grid vertices with index values a, b, c. Take out the feature vector with index value (a, b, c) from the implicit representation feature encoding Then the feature sampling calculation method is:

[0120]

[0121] d a,b,c = ||X - X a,b,c || 2

[0122]

[0123] Step 1.2.2, vector concatenation: Concatenate the sampled feature f sam and X to obtain the query vector f query .

[0124] Step 1.2.3, implicit decoding: Build a local decoder, and input the query vector f query into the local decoder to obtain the confidence score of the category of this point The local decoder adopts the decoding network described in the literature Peng S, Niemeyer M, Mescheder L, et al. Convolutional occupancy networks[C] / / European Conference on ComputerVision. Springer, Cham, 2020: 523 - 540.

[0125] Step 2 includes the following steps:

[0126] Step 2.1, randomly sample points in the box space: Randomly sample N query points in the box space

[0127] Step 2.2, Semantic Label Generation: For a given voxel label Y ∈ {1, 2, …, K} h×w×d , where K is the number of categories, and the coordinates X of a query point in a three-dimensional space sam , convert Y into one-hot code form Y one-h ∈ R h×w×d×K , calculate the index of the voxel grid to which the point belongs, and set the semantic label y of the point to the one-hot code value under this index value, and at the same time generate the geometric label y g .

[0128]

[0129] Among them, y[1] represents the first element of y.

[0130] In actual implementation, K = 12, including the following categories: empty, ceiling, floor, wall, window, chair, bed, sofa, table, TV, furniture, and object. It is better to take N as 10240.

[0131] Step 2.3, For all the collected points, repeat Steps 2-1 to 2-2 to complete the generation of semantic labels for a training sample.

[0132] Step 2.4, For each training sample in the training set, repeat Steps 2-1 to 2-3 to complete the generation of implicit representation semantic labels. The training set uses the dataset described in the literature Silberman N, Hoiem D, Kohli P, et al. Indoor segmentation and support inference from rgbd images[C] / / European conference on computer vision. Springer, Berlin, Heidelberg, 2012:746-760.

[0133] Step 3 includes the following steps:

[0134] Step 3.1, Build a loss calculation module, including a semantic loss module and a geometric loss module. Among them, the loss calculation method in the semantic loss module is as follows:

[0135]

[0136] Among them, w cl is the weight coefficient of each category, used to alleviate the problem of class imbalance. is the confidence score of the cl-th category decoded by the implicit decoding network, and y cl is the one-hot code value of the cl-th category of the semantic label generated in Step 2.4.

[0137] The loss calculation method in the geometric loss module is as follows:

[0138]

[0139] Among them, y g is the geometric label generated in step 2.4,

[0140] Finally, the total loss term of the loss calculation module is as follows:

[0141] L = L sem + 2 * L geo

[0142] Step 3.2, use the stochastic gradient descent optimizer to train for 200 rounds on the training set with a learning rate of 10 -2 to complete the training of the implicit representation encoding and decoding network. The stochastic gradient descent optimizer adopts the optimizer described in the literature Bottou L. Stochastic gradient descent tricks [M] / / Neural networks: Tricks of the trade. Springer, Berlin, Heidelberg, 2012: 421 - 436.

[0143] Step 4 includes the following steps:

[0144] Step 4.1, generate spatial query points. Specify the resolution H u , W u , D u , and divide the box - shaped space into a voxel grid of H u ×W u ×D u . Calculate the three - dimensional point coordinates of the center of each grid to generate the final query point set.

[0145] Step 4.2, generate the scene implicit representation. Input the RGB and depth maps into the implicit representation encoding network to obtain the implicit representation feature encoding of the scene.

[0146] Step 4.3, generate point labels. Input the implicit representation feature encoding and the query point set into the implicit representation decoding network to obtain the confidence score of each point, and take the subscript with the maximum confidence score as the class label of the point.

[0147] Step 4.4, voxel result generation. Set K colors for K classes. For each query point, if the class label of the query point is 1, it is ignored. Otherwise, with this point as the center, generate a cube with a side length of The bounding box is obtained, and the six faces of the box are painted with the color of this category to generate the final voxel completion result.

[0148] As Figure 4 shown, in step 4 above, the resolution is determined to be 60×36×60, and the voxel completion result is obtained. This figure shows the completion result at this resolution, as well as the semantic information predicted by the present invention, such as walls, tables, and chairs. It can be seen that after being completed by the present invention, the occluded parts of the wall, the table, the chair, etc. have better restored the complete geometric information.

[0149] As Figure 5 shown, in step 4 of the above overview, the resolution is determined to be 240×144×240, and the voxel completion result is obtained. This figure shows the completion result at a higher resolution of the same scene. It can be seen that this result Figure 4 shows more geometric details.

[0150] In specific implementation, the present invention also provides a computer storage medium. Among them, the computer storage medium can store a program, and when the program is executed, it can include some or all of the steps in the various embodiments provided by the present invention. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0151] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments of the present invention.

[0152] The present invention provides an idea and method for a three-dimensional scene semantic completion method based on implicit representation. There are many methods and ways to specifically implement this technical solution. The above description is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by the prior art.

Claims

1. A method for semantic completion of 3D scenes based on implicit representation, characterized in that, it includes the following steps: Step 1, construction of implicit encoding and decoding network: Construct an implicit representation encoding network for generating implicit representation and an implicit representation decoding network for decoding implicit representation; Step 2, generation of implicit representation semantic labels: For each training sample, randomly sample a certain number of points in 3D space, and calculate the class label of each point according to the given voxel-level label to generate implicit representation semantic labels; Step 3, training of implicit encoding and decoding network: According to the stochastic gradient descent algorithm, use the implicit representation semantic labels to train the implicit encoding and decoding network; Step 4, generation of voxel semantic completion results: Input RGB images and depth images, use the implicit representation encoding network to calculate the implicit representation feature encoding of the scene, generate a sampling space according to the specified resolution, uniformly sample points from this sampling space, and use the implicit representation decoding network to calculate the semantic label of each point, and finally obtain the final voxel completion result through the voxel result generation module; Among them, the construction of the implicit encoding and decoding network in Step 1 includes the following steps: Step 1-1, build an implicit representation encoding network Step 1-2, build the implicit representation decoding network Construct the implicit representation encoding network described in Step 1-1 including the following steps: Step 1-1-1, two-dimensional feature extraction: construct the RGB image feature extraction network F rgb and the depth image feature extraction network F dep ; Input the RGB image I rgb into F rgb to obtain the two-dimensional RGB feature Input the depth image I depth into F dep to obtain the two-dimensional depth feature Step 1-1-2, 2D to 3D back-projection: Given a box-shaped space S centered at the origin of the coordinate system with dimensions D s , W s , H s and a given feature image f with a shape of H×W×ch, where H is the height of the feature image, W is the width of the feature image, and ch is the number of channels of the feature image. Given the camera intrinsic parameter K and the extrinsic parameter P, for a point coordinate x=(u,v) on the feature image, the feature vector of this point on the feature image is f x , calculate the coordinate X of x in the three-dimensional space X = P -1 K -1 x. If X falls outside the box-shaped space S, ignore this point. Otherwise, the feature vector at X is f X and f X = f x ; Voxelize the three-dimensional space into voxels V of h×w×d, where h, w, and d represent the number of grids in the height, width, and length directions of the voxels respectively; the length, width, and height of each grid are And calculate the indices (i, j, k) of the grid to which X belongs, where i, j, and k represent the index values in the height, width, and length directions respectively; for each grid, if no points fall within the grid, set the feature vector of the grid to a zero vector, and if p points fall within the grid, calculate the feature vector of the grid through max pooling; Two-dimensional RGB features RGB voxel features are obtained through the back-projection from 2D to 3D Two-dimensional depth features Depth voxel features are obtained through the back-projection from 2D to 3D Step 1-1-3, multi-scale three-dimensional feature completion: construct a downsampling network F down and an upsampling network F up ; construct a feature completion network F u and calculate a feature pyramid, the feature completion network includes two cascaded downsampling networks and two cascaded upsampling networks RGB voxel features After passing through the feature completion network, an RGB feature pyramid is obtained Depth voxel features After passing through the feature completion network, a depth feature pyramid is obtained Step 1-1-4, Fusion: Fuse the RGB feature pyramid and the depth feature pyramid by element-wise addition to obtain the implicit representation feature encoding where f imp has a shape of h×w×d×ch p , and ch p is the number of channels of f imp ; Construct the implicit representation decoding network described in Step 1-2 including the following steps: Step 1-2-1, feature sampling; Step 1-2-2, vector concatenation; Step 1-2-3, implicit decoding; The method of feature sampling described in Step 1-2-1 includes: For any point X in the box-shaped space S and the implicit representation feature encoding f imp , calculate the grid index (i, j, k) to which the point belongs, and sample the features according to the value of the index; In step 1-2-2, the method of vector concatenation includes: concatenating the sampled feature f sam and X to obtain the query vector f query ; In step 1-2-3, the method of implicit decoding includes: building a local decoder, and inputting the query vector f query into the local decoder to obtain the confidence score of the category of this point 2. The method for semantic completion of 3D scenes based on implicit representation according to claim 1, characterized in that, the generation of implicit representation semantic labels described in Step 2 includes the following steps: Step 2-1, random sampling of query points in box space: Randomly sample N query points in the box space; Step 2-2, semantic label generation: For a given voxel label Y and the coordinates X of a query point in a three-dimensional space sam , convert Y into one-hot code form Y one-hot , calculate the index of the voxel grid to which the point belongs, and set the semantic label y of the point to the one-hot code value under the value of the index, and at the same time generate the geometric label y of the point g ; Step 2-3, for all sampled points, repeat Steps 2-1 to 2-2 to complete the generation of semantic labels for one training sample; Step 2-4, for each training sample in the New York dataset, repeat Steps 2-1 to 2-3 to complete the generation of implicit representation semantic labels.

3. The method for semantic completion of 3D scenes based on implicit representation according to claim 2, characterized in that, the training of the implicit encoding and decoding network described in Step 3 includes the following steps: Step 3-1, construct a loss calculation module, including a semantic loss module and a geometric loss module; Step 3-2, use the stochastic gradient descent optimizer to complete the training of the implicit encoding and decoding network on the New York dataset.

4. The method for semantic completion of 3D scenes based on implicit representation according to claim 3, characterized in that, the generation of voxel semantic completion results described in Step 4 includes the following steps: Step 4-1, generate spatial query points according to the specified resolution H u , W u , D u , divide the box-shaped space into voxels of H u ×W u ×D u , calculate the three-dimensional point coordinates of the center of each grid, and generate the final query point set; Step 4-2, generate the implicit representation of the scene, and input the RGB image and depth image into the implicit representation encoding network to obtain the implicit representation feature encoding of the scene; Step 4-3, generate point labels, input the implicit representation feature encoding and the query point set into the implicit representation decoding network, obtain the confidence score of each point, and take the subscript with the largest confidence score as the class label of this point; Step 4-4, voxel result generation. Set K colors for K categories. For each query point, if the category label of this query point is 1, it is ignored. Otherwise, with this point as the center, generate a bounding box with a side length of , and paint the 6 faces of this bounding box with the color of this category to generate the final voxel completion result.

Citation Information

Patent Citations

  • Three-dimensional scene perception method and device, electronic equipment, robot and medium

    CN113487664A

  • Semantic scene completion method and system based on point cloud-voxel aggregation network model

    CN113850270A