Indoor semantic scene completion method and system based on bird's-eye view layered interactive perception
By employing a layered interactive perception method based on bird's-eye view, utilizing binocular cameras and layered bird's-eye view representation modeling, and designing a refined complementary interactive network, the problems of semantic consistency and occlusion area completion in indoor semantic scene completion are solved, achieving low-cost and high-precision indoor 3D semantic scene reconstruction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU DIANZI UNIV
- Filing Date
- 2026-01-18
- Publication Date
- 2026-04-10
AI Technical Summary
Existing indoor semantic scene completion technologies suffer from poor semantic consistency when dealing with complex vertical spatial structures, fail to fully integrate semantic and geometric features from different perspectives, and rely on high-cost sensors, making it difficult to effectively complete occluded areas.
A layered interactive perception method based on bird's-eye view is adopted. Multi-view images are acquired through binocular cameras, context-aware features are extracted, and three-dimensional geometric representation volume is generated by combining camera intrinsic parameters. Layered bird's-eye view representation modeling and a refined complementary interactive network are designed to achieve low-cost and high-precision semantic scene completion.
It improves the vertical semantic consistency of indoor scenes, effectively filters noise information and completes occluded areas, reduces hardware costs, and achieves high integrity and high accuracy semantic scene completion.
Smart Images

Figure CN121526928B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of computer vision and three-dimensional scene reconstruction, and proposes an indoor semantic scene completion method and system based on hierarchical interactive perception of bird's eye view. BACKGROUND
[0002] Semantic scene completion aims to simultaneously recover the three-dimensional geometric structure of the scene and the semantic label of each spatial position from limited observation data, and is one of the core technologies for realizing intelligent perception of indoor environment. With the rapid development of indoor robot navigation technology and VR / AR technology, higher requirements are put forward for the completeness, accuracy and real-time performance of indoor scene modeling.
[0003] Existing semantic scene completion technologies mainly fall into three categories: (1) methods based on single-view images and depth maps: relying on semantic segmentation results and depth estimation results of a single 2D image for scene completion, but limited by the limitations of single-view information, it is difficult to handle occluded areas and complex spatial structures, and the completeness of the completion result is poor; (2) methods based on multi-view images: obtaining depth information through multi-view matching, and combining semantic segmentation to achieve completion, but multi-view feature fusion mainly uses simple splicing or shallow interaction, and does not fully utilize the complementarity of different modal features, limiting the completion accuracy; (3) methods based on point clouds and voxels: directly processing three-dimensional point clouds or voxel data, which can capture three-dimensional structures, but are highly dependent on sensors (such as LiDAR) and have high costs, and in the joint completion of semantic information and geometric structure, local details are easily lost and semantic labels are easily confused.
[0004] However, existing technologies still face three major challenges: first, the complex vertical spatial structure of indoor scenes is not well handled, and the functional and interactive relationship between different height layers is not considered, resulting in poor semantic consistency in the vertical direction of the completion result; second, the semantic and geometric features under different perspectives lack effective fusion, making it difficult to filter noise information and complete occluded areas through complementation. SUMMARY
[0005] The present application aims to overcome the shortcomings of existing indoor semantic scene completion technologies and provide an indoor semantic scene completion method and system based on hierarchical interactive perception of bird's eye view, which realizes low-cost, high-precision and high-completeness indoor semantic scene completion through hierarchical bird's eye view representation modeling method and fine complementary interaction network, and provides reliable three-dimensional semantic scene support for indoor intelligent applications.
[0006] In one aspect of the present application, an indoor semantic scene completion method based on hierarchical interactive perception of bird's eye view is provided, including the following steps:
[0007] Step one, through binocular camera, multi-view image input of indoor scene is acquired, context perception features of different levels in left and right images are extracted, and left and right scene context perception features are obtained;
[0008] Step two, the left and right scene context perception features output in step one are combined with camera intrinsic parameters, depth geometry volume is generated by using disparity conversion and denoising, and three-dimensional geometric representation volume is further obtained by three-dimensional convolution preprocessing;
[0009] Step three, for the left scene context perception features output in step one, a hierarchical bird's eye view representation modeling method is designed, and hierarchical bird's eye view Figure Three Grid volume and hierarchical bird's eye view features are output.
[0010] Step four, a fine complementary interaction network is designed, the hierarchical bird's eye view Figure Three Grid volume and three-dimensional geometric representation volume are complementarily interacted in layers, and the interaction results are spliced according to the hierarchical height to obtain a unified scene perception bird's eye view volume.
[0011] Step five, the unified scene perception bird's eye view volume and the three-dimensional geometric representation volume are fused by convolution regularization and channel calibration to generate a joint scene volume representation.
[0012] Step six, the joint scene volume representation and the hierarchical bird's eye view features output in step three are integrated and interacted, and finally a complete semantic scene completion result is output through a three-dimensional convolution network and scene decoding.
[0013] Another aspect of the present application also provides an indoor semantic scene completion system based on bird's eye view hierarchical interaction perception, comprising the following modules:
[0014] The feature extraction module is used for acquiring multi-view images of an indoor scene through a binocular camera, obtaining left and right scene context perception features, and modeling to obtain hierarchical bird's eye view Figure Three Grid volume and hierarchical bird's eye view features.
[0015] The volume representation module is used for obtaining three-dimensional geometric representation volume from left and right scene context perception features combined with camera intrinsic parameters, and obtaining a unified scene perception bird's eye view volume through complementary interaction, and generating a joint scene volume representation.
[0016] The semantic scene completion module is used for integrating and interacting the joint scene volume representation and the hierarchical bird's eye view features, and outputting a complete semantic scene completion result through a three-dimensional convolution network and scene decoding.
[0017] Compared with the prior art, the present application has the following beneficial effects:
[0018] (1) Hierarchical modeling improves semantic consistency: Through adaptive hierarchical modeling and layer-specific feature enhancement, the vertical layered structure of indoor scenes is fully utilized to make the semantic labels of different functional layers more accurate, solving the problem of semantic confusion in the vertical direction of traditional bird's eye view BEV method;
[0019] (2) Fine interaction improves completeness: Through depth confidence weighted intra-layer interaction and function dependent cross-layer interaction, noise information is effectively filtered and occluded areas are completed;
[0020] (3) Low cost: Only rely on binocular camera as input device, without high-priced sensor, reduce the hardware cost. BRIEF DESCRIPTION OF DRAWINGS
[0021] Figure 1 The flowchart of the indoor semantic scene completion method based on bird's eye view hierarchical scene perception and three-dimensional complementary interaction modeling of the application;
[0022] Figure 2 The flowchart of the hierarchical bird's eye view representation modeling method in the application;
[0023] Figure 3 The flowchart of the fine complementary interaction network in the application. DETAILED DESCRIPTION
[0024] The application will be further described in conjunction with the drawings and specific embodiments.
[0025] In one aspect of the application, an indoor semantic scene completion method based on bird's eye view hierarchical interaction perception is provided, Figure 1 The flowchart of the method of the application. As Figure 1 shown, the specific process steps of the method of the application are as follows:
[0026] Step one, using binocular cameras with the same resolution to collect left and right images of indoor scenes, and ensuring the baseline distance of binocular cameras stable during image acquisition to ensure the accuracy of parallax calculation. The left and right images collected are input into a pre-trained feature extraction network, and through multiple convolution layers, the image resolution is gradually down-sampled to extract context perception features.
[0027] Step two, combining the left and right image features output in step one with the pre-calibrated camera intrinsic parameters, generating a three-dimensional geometric representation volume through a stereo volume construction process, and the specific process of constructing the three-dimensional geometric representation volume is as follows:
[0028] First, the matching cost of left and right image features is calculated using a grouping correlation algorithm, which accurately captures pixel-level correspondence by grouping features and calculating intra-group inner product. The matching cost is generated by inter-group correlation calculation:
[0029]
[0030] where, represents the matching cost value at position with disparity and group index . represents the number of input feature channels, i.e., the dimension of each feature vector. represents the number of groups, used to divide the input features into multiple groups for correlation calculation. represents the feature vector of the left image at position with group index . represents the feature vector of the right image at position with group index . represents the inner product operation of two feature vectors.
[0031] Secondly, according to the camera imaging model, the matching cost is converted into the depth value and three-dimensional coordinates in the world coordinate system to obtain the initial depth geometric volume; the formula is:
[0032]
[0033] where, represents the depth value at image coordinates , i.e., the distance from the point to the camera. represents the focal length in the horizontal and vertical directions of the camera. represents the distance between the optical centers of the left and right cameras. represents the disparity value at image coordinates . represents the horizontal pixel coordinate of the intersection (principal point) of the camera optical axis and the image sensor plane in the image coordinate system, represents the vertical pixel coordinate of the principal point in the image coordinate system. respectively represent the horizontal and vertical coordinates of the point in the world coordinate system, which together with the depth determine the three-dimensional position of the point.
[0034] Finally, the depth geometric volume is regularized by using a three-dimensional convolutional network with an encoder-decoder structure, multi-scale feature aggregation is realized by three-dimensional convolutional blocks, geometric features are strengthened and noise is reduced by residual connection, and finally a three-dimensional geometric representation volume is obtained by channel compression through a three-dimensional convolutional block with a kernel size of 1 .
[0035] Step three, a hierarchical bird's eye view representation modeling method is designed for the preprocessed left image features, and a hierarchical bird's eye view Figure Three3D grid volume and layered bird's-eye view features. For example... Figure 2 As shown, layered bird's-eye view Figure Three The specific process of 3D mesh volume and layered bird's-eye view features is as follows:
[0036] First, based on prior knowledge of the height distribution of objects in an indoor scene, a set of height interval boundary values is adaptively determined. For each height level, the adjustable range of the boundary is considered. , The height range boundary values are determined by the adjustable boundary parameter β. Defined in the height range Inside: ,in, Let these be the initial upper and lower boundaries of the k-th layer;
[0037] Secondly, through a differentiable inverse projection operation, the left-image features in the context-aware features are... Mapped to the 3D mesh volume corresponding to each height level k In this operation, a differentiable bilinear sampling kernel K is used. The formula for calculating the differentiable projection operation is:
[0038]
[0039] For differentiable bilinear sampling kernels When the coordinate z falls within the interval of height layer k At that time, The value is set to 1 otherwise, and a Gaussian smoothing function is used to make the sampling kernel differentiable. Based on the camera intrinsic parameters, a mapping relationship between the image coordinate system and the world coordinate system is established, and a validity mask is calculated. When image pixels The corresponding 3D point in the world coordinate system When within the camera's field of view and unobstructed. The value is 1 if the condition is met, and 0 otherwise. Through the above differentiable inverse projection operation, the features of the left image obtained in step one are mapped to the corresponding 3D mesh volume for each height layer. .
[0040] Following the above operations, a set of lightweight 3D convolutional networks is designed for each height layer. Each network group contains multiple 3D convolutional layers and one 3D batch normalization layer, using Leaky ReLU as the activation function. Input a lightweight 3D convolutional network corresponding to the height layer to obtain an enhanced layered bird's-eye view. Figure Three 3D mesh volume .
[0041] Finally, the camera intrinsic parameters are processed by the encoder to obtain the intrinsic parameter encoding vector; the intrinsic parameter encoding vector is extended to the two-dimensional left image feature with matching spatial dimensions, and the two-dimensional left image feature An element-wise multiplication operation is performed to integrate the camera intrinsic information into the image features; finally, a hierarchical bird's eye view feature is obtained through a convolution layer and a normalization layer.
[0042] Step four, for the hierarchical bird's eye view Figure Three volume obtained in step three, a fine-grained complementary interaction network is designed, such as Figure 3 , and the specific implementation process is as follows:
[0043] First, a layer-intra complementary interaction method is designed to flatten the hierarchical bird's eye view Figure Three volume obtained in step three in the spatial and depth dimensions, respectively, and map them through a linear transformation layer to generate the and BEV volume of the three-dimensional volume, respectively. . Among them, represents the attention query vector, represents the key vector, represents the value vector. Then, the confidence is calculated based on the depth probability distribution of , and the specific process is as follows:
[0044]
[0045] where, represents the depth value of the coordinate in the three-dimensional volume, represents the maximum depth dimension, represents the probability value of the depth d at the spatial coordinate . The depth value is converted to probability P through the softmax function, and the maximum probability value in the depth dimension is selected as the confidence of the spatial coordinate, and the formula is:
[0046]
[0047] where is the confidence of the spatial coordinate (x, y). Set the confidence threshold τ to filter low-confidence results and obtain the confidence matrix .
[0048]
[0049] The feature volume after layer-intra interaction is calculated using the above formula , where and denote the softmax function along each row and each column of the input matrix, respectively, where denotes the element-wise product. For each layer, the hierarchical bird's eye view Figure Three dimensional grid volume are interacted by the above method, respectively, to obtain .
[0050] Secondly, a cross-layer functional relationship interaction method is designed. The functional dependency graph matrix is constructed according to the indoor inter-layer functional dependency prior The initial elements of the matrix are set based on the prior knowledge of the inter-layer functional association of the indoor scene. The element in the i-th row and the j-th column of the matrix denotes the "functional dependency strength of the i-th layer to the j-th layer".
[0051] as the query , and other layers as the key , value , cross-attention calculation is performed to obtain the output feature volume of each layer ;
[0052]
[0053] According to the above formula, cross-attention calculation is performed, and is taken as the query , and other layers are taken as the key and value , and the cross-attention calculation is performed in combination with the functional dependency graph matrix to obtain the output feature volume of each layer ; is the dependency weight coefficient; and the of each height layer is spliced in the order of height level to obtain the unified scene perception bird's eye view volume .
[0054] Step five, splicing and in the channel dimension; inputting a three-dimensional residual convolutional network for regularization; using a squeeze and excitation module for channel calibration; dividing the calibrated volume into 4 groups, respectively, using 3D dilated convolution with different dilation rates to capture multi-scale features; splicing the multi-scale features in the channel, and aggregating to obtain the joint scene volume representation through a convolution layer and a group normalization layer.
[0055] Step six, the layered bird's eye view features obtained in step three are up-sampled by transposed convolution, an outer product operation is performed with the joint scene volume representation, and the processing result is input into a 3D-UNet network with pre-trained weights for regularization. Finally, the result is processed by a decoding head: semantic labels and occupancy states are output through three-dimensional convolution to obtain a complete semantic scene completion result.
[0056] The pre-trained weights come from the general feature processing capability obtained by training a large-scale data set, which can significantly improve the adaptability and robustness of the model in different scenes. By introducing pre-trained weights, the model can better assign appropriate weights to the features in the model, thereby providing high-quality regularization effect. This process not only improves the efficiency of the algorithm, but also to some extent reduces the bias caused by insufficient data or special scenes.
[0057] In another aspect of the present application, an indoor semantic scene completion system based on layered interactive perception of bird's eye view is also provided, comprising the following modules:
[0058] The feature extraction module is used to obtain multi-view images of the indoor scene through a binocular camera, obtain left and right scene context perception features, and model and obtain layered bird's eye view Figure Three grid volume and layered bird's eye view features.
[0059] The volume representation module is used to obtain a three-dimensional geometric representation volume from the left and right scene context perception features combined with the camera intrinsic parameters, and obtain a unified scene perception bird's eye view volume through complementary interaction, to generate a joint scene volume representation.
[0060] The semantic scene completion module is used to integrate and interact the joint scene volume representation and the layered bird's eye view features, and output a complete semantic scene completion result through a three-dimensional convolution network and scene decoding.
[0061] The feature extraction module includes a perception feature extraction unit and a layered bird's eye view feature extraction unit.
[0062] The perception feature extraction unit is used to obtain multi-view images of the indoor scene through a binocular camera, extract context perception features of different levels in the left and right images, and obtain left and right scene context perception features.
[0063] The layered bird's eye view feature extraction unit is used to design layered bird's eye view representation modeling for the left scene context perception features, and output layered bird's eye view Figure Three grid volume and layered bird's eye view features.
[0064] The volume representation module includes a three-dimensional geometric representation volume unit and a unified scene perception bird's eye view unit.
[0065] The three-dimensional geometric representation volume unit is used for perceiving features of left and right scene contexts, combining with camera intrinsic parameters, using parallax conversion and denoising to generate a depth geometry volume, and obtaining a three-dimensional geometric representation volume through three-dimensional convolution preprocessing.
[0066] The unified scene perception bird's eye view volume unit is used for designing a complementary interaction network, combining a hierarchical bird's eye view Figure Three The hierarchical bird's eye view volume is obtained by performing complementary interaction between the hierarchical grid volume and the three-dimensional geometric representation volume in layers, and splicing interaction results according to layer levels.
[0067] The embodiments of the application are described in detail above with reference to the drawings, but the application is not limited to the described embodiments. For those skilled in the art, various changes, modifications, replacements and variations of the embodiments can be made without departing from the principles and spirits of the application, and still fall within the protection of the application.
[0068] Embodiments:
[0069] Dataset Introduction: The NYU-V2 dataset (full name: NYU Depth V2 dataset) is a 3D scene understanding dataset for real indoor scenes proposed by Silberman et al. in 2012. The NYU-CAD dataset is improved based on the NYU-V2 dataset by Firman et al. in 2016, and its core goal is to solve the misalignment problem between depth images and 3D labels in the NYU-V2 dataset, so as to provide higher quality training and testing data for semantic scene completion algorithms. NYU-CAD does not directly use the original sensor-acquired depth images of NYU-V2, but synthesizes depth images by projecting 3D semantic labels onto the image plane. This synthesis method greatly reduces the misalignment between depth information and 3D labels (voxel occupancy, semantic labels), improves the consistency and reliability of the data, and helps to more accurately evaluate the real performance of algorithms in scene completion and semantic prediction tasks.
[0070] In other aspects of data composition, NYU-CAD maintains compatibility with NYU-V2: its RGB images and 2D semantic segmentation ground truths directly follow the original data of the NYU-V2 dataset, only replacing the depth image part. This design not only retains the rich diversity of indoor scenes of NYU-V2, but also optimizes the quality of depth data to provide a more reasonable test benchmark for the performance improvement of semantic scene completion algorithms.
[0071] Experimental introduction: In terms of data division, in order to maintain consistency with mainstream research, the experimental settings of Song et al. (2017), Dourado et al. (2022) and other scholars were referred to, and 1449 samples were divided into 795 training samples and 654 test samples. The model was trained for 40 epochs using the AdamW optimizer on a single NVIDIA GeForce RTX 4090 using the Pytorch framework, with a learning rate of 1 × 10e-4 and a batch size of 8. To use memory more effectively, the 3D-UNet input was set to 128 × 128 × 16.
[0072] Evaluation index introduction: For semantic scene completion (SSC), the intersection over union (IoU) between the true value GT and the predicted cube is measured, excluding voxels outside the perspective or space. Calculate the IoU intersection ratio score of each category, and then take the average of all categories to get the average intersection ratio (mIoU) score. The calculation formula is as follows:
[0073]
[0074] Wherein, wherein, TP (True Positive) represents the number of voxels that are "predicted as occupied / target category and true value as occupied / target category"; FP (False Positive) represents the number of voxels that are "predicted as occupied / target category but true value as non-occupied / non-target category"; FN (False Negative) represents the number of voxels that are "predicted as non-occupied / non-target category but true value as occupied / target category".
[0075] Related method introduction: SSCNet is a semantic scene completion method proposed by Princeton University in 2017 at the top conference of computer vision CVPR. Its core task is to predict the complete 3D voxel occupancy of the scene and the semantic label of each voxel from a single view depth image, covering the visible and occluded areas within the camera frustum. The significant features of this method include: an end-to-end 3D convolutional network architecture is adopted to fuse the scene completion and semantic labeling tasks, taking advantage of the coupling relationship between the two to improve performance, which is better than traditional methods that handle the two tasks separately; a context module based on 3D dilated convolution is designed to expand the receptive field without losing resolution, effectively capturing 3D context information across objects and solving the problem of sparse 3D voxel data; based on the truncated signed distance function (TSDF), the voxel encoding is optimized to eliminate the dependence on the perspective and adjust the gradient distribution, providing better geometric feature signals; multi-scale feature responses are fused to adapt to the identification needs of objects of different sizes. In terms of performance, SSCNet performs well on the NYU-CAD real dataset, with a significant improvement in semantic scene completion average intersection over union (mIoU) over traditional methods. The related dataset and code have been open-sourced, providing complete 3D geometric and semantic information support for scene understanding tasks such as robot navigation and object grasping.
[0076] MonoScene is a semantic scene completion (SSC) method proposed by the French National Institute for Research in Computer Science and Automation in 2022 at the top conference of computer vision CVPR. Its core innovation is that it can simultaneously infer the dense 3D voxel geometry and semantic labels of indoor and outdoor scenes from a single RGB image, and also predict the reasonable scene content outside the camera field of view. The main features of this method include: a 2D and 3D UNet series architecture is adopted, with an innovative FLoSP module to connect the two networks, inspired by optical principles to project multi-scale 2D features along the line of sight to 3D space, achieving efficient flow and decoupling of 2D-3D information; a 3D context relationship prior (3D CRP) layer is introduced to learn four types of semantic relationships between voxels and capture global context information, improving spatiotemporal semantic consistency; a new loss function is designed, including a scene-class affinity loss to optimize global class affinity and a frustum proportion loss to calibrate local frustum class distribution, effectively alleviating the ambiguity problem of occlusion; without the need for depth maps, point clouds, and other geometric inputs, it is highly versatile and suitable for indoor and outdoor scenes. The related code and pre-trained models have been open-sourced, providing an efficient solution for scene understanding applications that lack depth sensors.
[0077] BRGScene is a 3D semantic scene completion method proposed by Shanghai Jiaotong University, Ningbo Digital Twin Research Institute, Tokyo Institute of Technology, PhiGent Robotics and other units. The related achievements were published in 2024. The core innovation is to integrate stereo vision geometry and overhead view representation. High-precision scene geometry completion and semantic labeling can be achieved through only binocular RGB images. It is especially good at predicting small objects at a distance. The related code has been open-sourced, providing a low-cost and reliable scene understanding solution for autonomous driving, robot navigation and other fields.
[0078] FFNet is a method for semantic scene completion (SSC) published in 2022. Its core is to optimize the use of RGB-D data to improve the accuracy of 3D geometry occupancy and object class estimation. The code has been open-sourced, providing a more efficient feature fusion solution for scene understanding based on RGB-D data.
[0079] SG-SSC is an SSC method proposed by Harbin Institute of Technology, Nanyang Technological University, Xiamen University and other units. It was published in 2025. The core is to solve the problem of unreliable depth data and complex scene semantic confusion. Based on single-view RGB-D images, it realizes high-precision 3D geometry occupancy and semantic label prediction. The related code has been open-sourced.
[0080] Experimental results:
[0081] Table 1: Comparison of IoU of the present invention and other methods on NYU-CAD dataset
[0082]
[0083] Ceil, Sofa, Wall, Chair and Table in Table 1 correspond to the key object elements in the five types of indoor scenes: ceiling, sofa, wall, chair and table. The mIoU of the present invention on NYU-CAD reaches 39.62, leading the suboptimal FFNet by 4.6%. The IoU of the Sofa class in the ground key target reaches 54.92, which is 15.2% higher than the MonoScene method. The IoU of the Chair class reaches 36.69, which is 16.7% higher than the SSCNet. This reflects the superior performance of the present invention in semantic completion in indoor scenes. The experimental results are shown in Table 1. As can be seen from the table, the overall effect of the present invention is better than the latest technical means.
Claims
1. An indoor semantic scene completion method based on bird's eye view layered interactive perception, characterized in that, The method comprises the following steps: Step 1: obtaining multi-view images of an indoor scene, and obtaining left and right scene context perception features; Step 2: using disparity conversion and denoising to generate a depth geometric volume from the left and right scene context perception features combined with camera intrinsic parameters, and obtaining a three-dimensional geometric representation volume through three-dimensional convolution preprocessing; Step 3: designing a hierarchical bird's eye view representation modeling for the left scene context perception features, and outputting a hierarchical bird's eye view three-dimensional grid volume and hierarchical bird's eye view features, the specific implementation process being: First, based on prior knowledge of the height distribution of objects in an indoor scene, a set of height interval boundary values B={ }; For each height level, the adjustable range of the boundary is combined. , The height range boundary values are determined by the adjustable boundary parameter β. Defined in the height range Inside: ,in Let these be the initial upper and lower boundaries of the k-th layer; Second, the two-dimensional left image features in the context-aware features are mapped to the corresponding three-dimensional grid volume of each height layer k by a differentiable back-projection operation The operation is implemented by a differentiable bilinear sampling kernel K. The operation is implemented by a differentiable bilinear sampling kernel K. Design a set of lightweight three-dimensional convolutional networks for each height layer Each set of networks contains multiple three-dimensional convolutional layers and a three-dimensional batch normalization layer Input the lightweight three-dimensional convolutional network corresponding to the height layer to obtain an enhanced layered bird's eye view three-dimensional grid volume Finally, the camera intrinsic parameters are processed by the encoder to obtain an intrinsic parameter encoding vector; the intrinsic parameter encoding vector is extended to the same spatial dimension as the two-dimensional left image feature matching, and the two-dimensional left image feature An element-by-element multiplication operation is performed to integrate the camera intrinsic parameter information into the image feature; finally, a hierarchical bird's eye view feature is obtained through a convolution layer and a normalization layer. Step 4: designing a complementary interaction network to interact the hierarchical bird's eye view three-dimensional grid volume and the three-dimensional geometric representation volume within layers, and splicing the interaction results according to layer heights to obtain a unified scene perception bird's eye view volume, the specific implementation process being: First, design complementary interactions within the layers to transform the layered bird's-eye view 3D mesh volume. The three-dimensional geometric representation volume is flattened in the spatial and depth dimensions, respectively, and then the solid volume is generated through linear transformation layer mapping. and BEV volume ,in, Represents the attention query vector. Represents the key vector. Represents a value vector; Then based on the depth probability distribution of The confidence is calculated, the confidence threshold τ is set, and the confidence matrix is filtered ; Will After transposing, calculate the softmax function value for each column and compare it with... Perform element-wise multiplication, and then combine with the confidence matrix. Perform element-wise multiplication to obtain the intermediate value F1; After calculating the softmax function value of each row, perform element-wise multiplication with the intermediate value F1 to complete the interaction and obtain the feature volume after intra-layer interaction. For each layer of the layered bird's-eye view, the 3D mesh volume All interactions were performed to obtain ; Secondly, the cross-layer function relationship interaction is designed, and a function dependency graph matrix is constructed according to the indoor inter-layer function dependency priori The initial elements of the matrix are set based on the prior knowledge of the indoor scene inter-layer function association, and the element in the i th row and the j th column of the matrix represents the function dependency strength of the i th layer to the j th layer. Will As a query , with other layers As a key , value , combined with dependent weight coefficient And the functional dependency graph matrix Cross attention calculation is carried out to obtain the output feature volume of each layer ; stitching the height layers in height level order , to obtain a unified scene-aware bird's-eye view volume ; Step 5: performing convolution regularization and channel calibration fusion on the unified scene perception bird's eye view volume and the three-dimensional geometric representation volume to generate a joint scene volume representation; Step 6: performing integrated interaction processing on the joint scene volume representation and the hierarchical bird's eye view features, and outputting a complete semantic scene completion result through a three-dimensional convolution network and scene decoding.
2. The bird's eye view based hierarchical interaction-aware indoor semantic scene completion method according to claim 1, characterized in that, The specific implementation process of obtaining the context perception features of different levels in the left and right images is: obtaining multi-view images of an indoor scene through a binocular camera, inputting the left and right images collected by the binocular camera into a pre-trained feature extraction network, gradually downsampling the image resolution through multiple convolution layers, and extracting left and right scene context perception features.
3. The indoor semantic scene completion method based on bird's eye view layered interaction perception according to claim 2, characterized in that, The specific implementation process of step 2 is: First, a grouping correlation algorithm is used to calculate the matching cost of the left and right image features, which captures the pixel-level correspondence by grouping the features and calculating the inner product within the group; Second, according to the camera imaging model, the matching cost is converted into depth values and three-dimensional coordinates in the world coordinate system to obtain an initial depth geometric volume. Finally, the three-dimensional convolution network with the encoder-decoder structure is used to regularize the deep geometric volume, realize multi-scale feature aggregation, pass through the residual connection, and then pass through the three-dimensional convolution block with a convolution kernel of 1 to compress the channel to obtain the final three-dimensional geometric representation volume .
4. The indoor semantic scene completion method based on bird's eye view layered interaction perception according to claim 3, characterized in that, The step five is specifically implemented as follows: and In the channel dimension splicing, the input three-dimensional residual convolutional network is used for regularization, and then the squeeze excitation module is used for channel calibration. The calibrated volume is averagely divided into h groups, 3D hollow convolution with different dilation rates is used to capture multi-scale features, the multi-scale features are spliced in the channel, and the joint scene volume representation is obtained through the convolution layer and group normalization layer aggregation.
5. The bird's eye view based hierarchical interaction-aware indoor semantic scene completion method according to claim 4, characterized in that, The specific implementation process of step 6 is: upsampling the hierarchical bird's eye view features through transpose convolution, performing an outer product operation with the joint scene volume representation, inputting the processing result into a 3D-UNet network with pre-trained weights for regularization, and finally processing the result through a decoding head: outputting semantic labels and occupancy states through three-dimensional convolution to obtain a complete semantic scene completion result.
6. An indoor semantic scene completion system based on aerial view layered interactive perception, for implementing the indoor semantic scene completion method of any one of claims 1 to 5, characterized in that, The method comprises the following modules: A feature extraction module is used to obtain multi-view images of an indoor scene through a binocular camera, obtain left and right scene context perception features, and model to obtain a hierarchical bird's eye view three-dimensional grid volume and hierarchical bird's eye view features; A volume representation module is used to obtain a three-dimensional geometric representation volume from the left and right scene context perception features combined with camera intrinsic parameters, and obtain a unified scene perception bird's eye view volume through complementary interaction to generate a joint scene volume representation; A semantic scene completion module is used to perform integrated interaction processing on the joint scene volume representation and the hierarchical bird's eye view features, and output a complete semantic scene completion result through a three-dimensional convolution network and scene decoding.
7. The bird's eye view based hierarchical interaction-aware indoor semantic scene completion system according to claim 6, characterized in that, The feature extraction module comprises a perception feature extraction unit and a hierarchical bird's eye view feature extraction unit; The perception feature extraction unit is used to obtain multi-view images of an indoor scene through a binocular camera, extract context perception features of different levels in the left and right images, and obtain left and right scene context perception features. The hierarchical bird's eye view feature extraction unit is configured to design hierarchical bird's eye view representation modeling for the left scene context-aware feature, and output a hierarchical bird's eye view three-dimensional grid volume and a hierarchical bird's eye view feature.
8. The bird's eye view based hierarchical interaction-aware indoor semantic scene completion system according to claim 7, characterized in that, The volume representation module comprises a three-dimensional geometric representation volume unit and a unified scene-aware bird's eye view volume unit. The three-dimensional geometric representation volume unit is configured to, for the left and right scene context-aware features, combine camera intrinsic parameters, use parallax conversion and denoising to generate a depth geometric volume, and obtain a three-dimensional geometric representation volume through three-dimensional convolution preprocessing. The unified scene-aware bird's eye view volume unit is configured to design a complementary interaction network, perform complementary interaction between the hierarchical bird's eye view three-dimensional grid volume and the three-dimensional geometric representation volume in layers, and splice the interaction results according to layer heights to obtain a unified scene-aware bird's eye view volume.
Citation Information
Patent Citations
Semantic scene completion method based on feature representation decomposition and aerial view fusion
CN116630975A
Methods for generating at least one ground truth from a bird's-eye view
DE102022214330A1