Semantic scene completion method and device, electronic equipment and readable storage medium

By using the FoundationStereo model for depth estimation and semantic feature extraction, the problem of restoring the three-dimensional spatial structure of visual images in semantic scene completion is solved, achieving higher accuracy in semantic scene completion.

CN120876892APending Publication Date: 2025-10-31SHENZHEN SWEET POTATO ROBOT CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510930419.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately recover 3D spatial structures when performing semantic scene completion based on visual images, especially in occluded and similar texture areas. Furthermore, insufficient acquisition and utilization of depth information leads to low accuracy in semantic scene completion tasks.

Method used

The FoundationStereo model is used for depth estimation. Multi-scale image feature maps, disparity probability distribution maps, and disparity maps are obtained through the binocular stereo depth estimation model. The disparity dimension is converted to the depth dimension. Semantic features are extracted by combining the depth map and the multi-scale feature map to determine the voxel occupancy data and voxel semantic data of the semantic data.

Benefits of technology

It improves the accuracy of semantic scene completion, accurately restores the three-dimensional spatial structure, and enhances the consistency of understanding the semantic and spatial geometric information of objects in driving scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876892A_ABST
    Figure CN120876892A_ABST
Patent Text Reader

Abstract

The invention discloses a semantic scene completion method and device, electronic equipment and a readable storage medium. The method comprises the following steps: acquiring a binocular image; performing depth estimation processing on the binocular image to obtain a multi-scale image feature map, a disparity probability distribution map and a disparity map of the monocular image; performing conversion processing from a parallax dimension to a depth dimension on the parallax image of the monocular image to obtain a depth image of the monocular image; processing the multi-scale feature map and the depth map of the monocular image to obtain a first semantic feature map of the monocular image; performing depth feature extraction processing on the parallax probability distribution diagram of the monocular image to obtain a depth probability distribution diagram of the monocular image; and determining voxel occupation data and voxel semantic data of the semantic occupation data based on the depth map of the monocular image, the first semantic feature map and the depth probability distribution map so as to perform semantic scene completion. According to the invention, semantic scene completion can be carried out based on visual image data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of intelligent driving technology, and in particular to a semantic scene completion method, device, electronic device, and readable storage medium. Background Technology

[0002] Semantic scene completion is a core task of autonomous driving technology. It aims to jointly reason and reconstruct the geometric and semantic information of a driving scene based on data collected from various sensors (such as visual image data and 3D sparse data). On one hand, the semantic scene completion task needs to accurately reconstruct the geometric information such as the shape, position, and size of each object in the driving scene to construct a complete and accurate 3D geometric model of the driving scene. On the other hand, the semantic scene completion task also needs to determine the semantic label corresponding to each object in the driving scene. Through semantic scene completion, vehicles can understand their surroundings, facilitating path planning and decision-making.

[0003] Semantic occupancy data, as key data for semantic scene completion tasks, is typically represented in the form of 3D voxel mesh data. Semantic scene completion tasks require determining the occupancy state and semantic label of each voxel in the 3D voxel mesh data to perform semantic scene completion. Therefore, there is an urgent need for a method capable of performing semantic scene completion based on visual image data. Summary of the Invention

[0004] To address the aforementioned technical issues, this disclosure provides a semantic scene completion method, apparatus, electronic device, and readable storage medium for performing semantic scene completion based on visual image data.

[0005] A first aspect of this disclosure provides a semantic scene completion method, including:

[0006] Acquire binocular images;

[0007] The binocular image is subjected to depth estimation processing to obtain a multi-scale image feature map, disparity probability distribution map, and disparity map of the monocular image; the monocular image is either the left or right eye image in the binocular image.

[0008] The disparity map of the monocular image is converted from the disparity dimension to the depth dimension to obtain the depth map of the monocular image.

[0009] The multi-scale feature map and depth map of the monocular image are processed to obtain the first semantic feature map of the monocular image;

[0010] Depth feature extraction is performed on the disparity probability distribution map of the monocular image to obtain the depth probability distribution map of the monocular image;

[0011] Based on the depth map, first semantic feature map, and depth probability distribution map of the monocular image, the voxel occupancy data and voxel semantic data of the semantic occupancy data are determined to complete the semantic scene.

[0012] A second aspect of this disclosure provides a semantic scene completion device, comprising:

[0013] A binocular image acquisition module is used to acquire binocular images;

[0014] A binocular stereo depth estimation module is used to perform depth estimation processing on the binocular image to obtain a multi-scale image feature map, a disparity probability distribution map, and a disparity map of the monocular image; the monocular image is either the left or right eye image in the binocular image.

[0015] The binocular stereo depth estimation module is also used to perform disparity dimension to depth dimension conversion processing on the disparity map of the monocular image to obtain the depth map of the monocular image.

[0016] The semantic feature extraction module is used to process the multi-scale feature map and depth map of the monocular image to obtain the first semantic feature map of the monocular image;

[0017] The depth feature extraction module is used to perform depth feature extraction processing on the disparity probability distribution map of the monocular image to obtain the depth probability distribution map of the monocular image.

[0018] The semantic scene completion module is used to determine the voxel occupancy data and voxel semantic data of the semantic occupancy data based on the depth map, the first semantic feature map and the depth probability distribution map of the monocular image, so as to perform semantic scene completion.

[0019] A third aspect of this disclosure provides a computer-readable storage medium storing a computer program for executing the semantic scene completion method provided in the first aspect.

[0020] A fourth aspect of this disclosure provides an electronic device, the electronic device comprising: a processor; a memory for storing executable instructions of the processor; the processor being configured to read the executable instructions from the memory and execute the instructions to implement the semantic scene completion method provided in the first aspect above.

[0021] A fifth aspect of this disclosure provides a computer program product that, when an instruction processor in the computer program product is executed, performs the semantic scene completion method provided in the first aspect of this disclosure.

[0022] This disclosure provides a semantic scene completion method, apparatus, electronic device, and readable storage medium. The electronic device acquires a binocular image, performs depth estimation processing on the binocular image, and obtains a multi-scale image feature map, a disparity probability distribution map, and a disparity map for the monocular image. Then, the electronic device performs a disparity-to-depth dimension conversion processing on the disparity map of the monocular image to obtain a depth map of the monocular image. Next, the electronic device processes the multi-scale feature map and depth map of the monocular image to obtain a first semantic feature map of the monocular image, and performs depth feature extraction processing on the disparity probability distribution map of the monocular image to obtain a depth probability distribution map of the monocular image. Finally, based on the depth map, the first semantic feature map, and the depth probability distribution map of the monocular image, the electronic device determines the voxel occupancy data and voxel semantic data of the semantic occupancy data for semantic scene completion. Therefore, in this disclosure, the electronic device can obtain a more accurate depth map based on depth estimation processing of the binocular image, and further combine the depth map and the multi-scale image feature map to perform semantic feature extraction processing based on geometric priors to obtain richer and more accurate semantic features, thereby improving the accuracy of subsequent semantic scene completion. Attached Figure Description

[0023] Figure 1 This is a flowchart illustrating a semantic scene completion method provided in an exemplary embodiment of this disclosure.

[0024] Figure 2 This is a schematic diagram of a semantic scene completion network provided in an exemplary embodiment of this disclosure.

[0025] Figure 3 This is a flowchart illustrating a semantic scene completion method provided in another exemplary embodiment of this disclosure.

[0026] Figure 4 This is a flowchart illustrating a semantic scene completion method provided in another exemplary embodiment of this disclosure.

[0027] Figure 5 This is a flowchart illustrating a semantic scene completion method provided in another exemplary embodiment of this disclosure.

[0028] Figure 6 This is a flowchart illustrating a semantic scene completion method provided in another exemplary embodiment of this disclosure.

[0029] Figure 7 This is a schematic diagram of the structure of a semantic scene completion device provided in an exemplary embodiment of this disclosure.

[0030] Figure 8 This is a structural diagram of an electronic device provided in an exemplary embodiment of this disclosure. Detailed Implementation

[0031] To explain this disclosure, exemplary embodiments of the disclosure will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the disclosure, and not all of them. It should be understood that the disclosure is not limited to exemplary embodiments.

[0032] It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps set forth in these embodiments do not limit the scope of this disclosure.

[0033] Application Overview

[0034] The intelligent driving described in this disclosure encompasses multiple fields, including autonomous driving, driver assistance systems, and robotic systems. Autonomous driving technology aims to enable intelligent vehicles to drive fully autonomously in various complex road conditions without human intervention; it is also known as driverless driving and represents an advanced form of intelligent driving. Driver assistance systems, through a series of sensors and algorithms, provide drivers with real-time road condition information, warnings, and support for some driving operations, such as automatic parking and adaptive cruise control, aiming to improve driving safety and convenience. Robotic systems further extend intelligent driving technology to service robots, industrial robots, and other fields, enabling robots to autonomously navigate, avoid obstacles, and complete specific tasks in complex environments, such as logistics delivery and warehouse management, demonstrating the broad application potential of intelligent driving technology in different scenarios.

[0035] Semantic scene completion is a core task of autonomous driving technology. It aims to jointly reason and reconstruct the geometric and semantic information of a driving scene based on data collected from various sensors (such as visual image data and 3D sparse data). On one hand, the semantic scene completion task needs to accurately reconstruct the geometric information such as the shape, position, and size of each object in the driving scene to construct a complete and accurate 3D geometric model of the driving scene. On the other hand, the semantic scene completion task also needs to determine the semantic label corresponding to each object in the driving scene. Through semantic scene completion, vehicles can understand their surroundings, facilitating path planning and decision-making.

[0036] Semantic occupancy data, a key component of semantic scene completion tasks, is typically represented as a 3D voxel mesh. Semantic scene completion requires determining the occupancy state and semantic label of each voxel in the 3D voxel mesh to generate semantic occupancy data. Traditional semantic scene completion methods largely rely on 3D sparse data (such as point cloud data and depth maps) to generate semantic occupancy data. However, the high cost of sensors for acquiring 3D sparse data limits the application scope of semantic occupancy data generation based on 3D sparse data. Image sensors for acquiring visual images, on the other hand, are inexpensive. Therefore, semantic scene completion based on visual images has gradually become a research hotspot. However, semantic scene completion based on visual images still faces many challenges. On the one hand, recovering the 3D spatial structure from a 2D image is difficult. For example, when dealing with occluded regions and regions with similar textures, it is difficult to accurately determine the depth and spatial geometry of objects. On the other hand, traditional semantic scene completion techniques have shortcomings in acquiring and utilizing depth information, leading to inaccurate understanding of the 3D spatial structure and reducing the accuracy of the semantic scene completion task. Therefore, how to perform semantic scene completion based on visual images is an urgent problem to be solved.

[0037] To address the aforementioned issues, this disclosure provides a target scene prediction method. An electronic device acquires binocular images and performs depth estimation processing on them to obtain a multi-scale image feature map, a disparity probability distribution map, and a disparity map for the monocular image. Then, the electronic device performs a disparity-to-depth dimension conversion on the disparity map of the monocular image to obtain a depth map. Next, the electronic device processes the multi-scale feature map and depth map of the monocular image to obtain a first semantic feature map, and performs depth feature extraction processing on the disparity probability distribution map of the monocular image to obtain a depth probability distribution map. Finally, based on the depth map, the first semantic feature map, and the depth probability distribution map of the monocular image, the electronic device determines the voxel occupancy data and voxel semantic data of the semantic occupancy data, obtaining semantic occupancy data. Therefore, in this disclosure, the electronic device can obtain a more accurate depth map by performing depth estimation processing on binocular images. Furthermore, by combining the depth map and the multi-scale image feature map with geometric prior-based semantic feature extraction processing, richer and more accurate semantic features are obtained, thereby improving the accuracy of subsequent semantic scene completion.

[0038] Exemplary methods

[0039] Figure 1 This is a flowchart illustrating a semantic scene completion method provided in an exemplary embodiment of this disclosure. This embodiment can be applied to electronic devices, such as... Figure 1 As shown, it includes the following steps:

[0040] Step 101: Obtain stereo images.

[0041] For example, an electronic device can acquire binocular images of a driving scene using a binocular image sensor. The binocular images can include a left-eye image and a right-eye image. The image dimensions of the left-eye and right-eye images are (H, W, C), where H represents the image height, W represents the image width, and C represents the number of image channels. In the case where the binocular images are RGB (Red, Green, Blue) images, both the left-eye image and the right-eye image have 3 channels.

[0042] Step 102: Perform depth estimation processing on the binocular images to obtain multi-scale image feature maps, disparity probability distribution maps, and disparity maps of the monocular images. The monocular image is either the left or right eye image from the binocular images.

[0043] For example, after acquiring a binocular image, the electronic device can input the binocular image into a binocular stereo depth estimation model. Correspondingly, the binocular stereo depth estimation model can perform depth estimation processing on the binocular image and output a disparity map of the monocular image. Simultaneously, during the depth estimation processing of the binocular image by the binocular stereo depth estimation model, the electronic device can also acquire a multi-scale image feature map and a disparity probability distribution map of the monocular image. Here, the monocular image can be either the left or right eye image in the binocular image. This embodiment uses the left eye image as an example for description; other cases are similar and will not be described in detail here.

[0044] It should be noted that the stereo depth estimation model can be the FoundationStereo model or other types of stereo depth estimation models, and this disclosure does not limit the specific model. This disclosure uses the FoundationStereo model as an example for illustration; other cases are similar and will not be described in detail here.

[0045] The FoundationStereo model comprises an STA (Side-Tuning Adapter) module, a hybrid cost volume construction module, a hybrid cost volume filtering module, a normalization module, and a Gated Recurrent Unit (GRU). The STA module includes a parallel configuration of the DepthAnythingV2 model and a CNN (Convolutional Neural Network) model. Furthermore, the DepthAnythingV2 model uses the DINOv2 model, based on the ViT (VisionTransformer) structure, as its backbone network.

[0046] In some embodiments, the FoundationStereo model performs depth estimation on stereo images as follows:

[0047] Step 1: The left-eye image is processed using the DepthAnythingV2 model in the STA module to extract depth features, obtaining its depth feature map. Simultaneously, the left-eye image is processed using a multi-layer CNN model in the STA module to extract multi-scale image features, obtaining its multi-scale image feature map. Then, the depth feature map and multi-scale image feature map of the left-eye image are concatenated along the channel dimension to obtain the concatenated feature map of the left-eye image.

[0048] Step two involves using the DepthAnythingV2 model in the STA module to extract depth features from the right-eye image, obtaining its depth feature map. Simultaneously, a multi-layer CNN model in the STA module is used to extract image features from the right-eye image, obtaining its multi-scale image feature map. Then, the depth feature map and multi-scale image feature map of the right-eye image are concatenated along the channel dimension to obtain the concatenated feature map of the right-eye image.

[0049] It should be noted that the DepthAnythingV2 model and the multi-layer CNN model together constitute the STA structure in the FoundationStereo model. The DepthAnythingV2 model uses the DINOv2 model based on the ViT structure as its backbone network. The DINOv2 model performs image feature extraction processing on the left-eye image to obtain a multi-scale image feature map of the left-eye image. Furthermore, electronic devices can use the multi-scale image feature map of the left-eye image output by the DINOv2 model as the multi-scale image feature map of the aforementioned monocular image. Compared to traditional CNN models that rely on local convolution operations and are prone to losing global feature information of the image, the DINOv2 model, based on the ViT structure, can establish global dependencies of image features, making it superior in 3D semantic scene completion tasks.

[0050] Step 3: Using the Hybrid Cost Volume (HCV) construction module, group-wise correlation is performed on the stitched feature maps of the left and right images to obtain the correlation feature map between them. Then, the stitched feature maps of the left and right images, along with the correlation feature map, are concatenated along the channel dimension to obtain the Hybrid Cost Volume Feature (HCV). Where C represents the mixed cost body V C The number of channels, D dispV represents the mixed cost body C The parallax, H / 4 represents the mixed cost volume V C The height, W / 4 represents the hybrid cost body V C Width, D disp =max disp / 4, max disp This represents the maximum disparity between the left and right eye images. Compared to the image size of a monocular image, the mixing cost volume V... C The height and width are both 1 / 4 of the monocular image.

[0051] Step four: The mixed cost volume V is filtered using the mixed cost volume filtering module. C Filtering is performed to obtain the filtered mixed cost body. Where C represents the filtered hybrid cost volume V′ C The number of channels, D disp V′ represents the filtered mixed cost volume. C The disparity, H / 4 represents the filtered mixed cost volume V′. C The height, W / 4, represents the filtered mixed cost volume V′. C Width, D disp =max disp / 4, max disp This represents the maximum disparity between the left and right eye images. Filtering can be performed using AHCF (Attention Hybrid Cost Filtering). AHCF filtering uses a 3D convolutional model and a self-attention (self-transformer) model to filter the hybrid cost volume V. C Filtering is performed. The 3D convolutional model and the self-attention model are parallel structures.

[0052] Step 5: The filtered mixed cost volume V′ is processed using a normalization module. C Normalization is performed along the disparity dimension to obtain the disparity probability distribution map of the left eye image. Among them, D disp Disparity probability distribution plot (disp) volume The disparity, H / 4 represents the disparity probability distribution map disp volume The height, W / 4 represents the disparity probability distribution map disp volume Width, D disp =max disp / 4, max dispThis represents the maximum disparity between the left and right eye images. Normalization can be performed using the softmax function, or other normalization functions; this embodiment is not limited to any particular normalization method. Furthermore, the electronic device can use the disparity probability distribution map of the left eye image as the disparity probability distribution map of the aforementioned monocular image. Then, along the disparity dimension, the disparity probability distribution map of the left eye image is processed... volume The disparity probabilities in the images are weighted and averaged to obtain the initial disparity map (disp) of the left eye image. init ∈R C ×H / 4×W / 4 Where C represents the initial disparity map disp init The number of channels, C is 1, H / 4 represents the initial disparity map disp init The height, W / 4 represents the initial disparity map disp init The width.

[0053] Step 6: The initial disparity map (disp) of the left-eye image is processed through a gated loop unit. unit Iterative optimization is performed to obtain an iteratively optimized initial disparity map. Then, the iteratively optimized initial disparity map is upsampled to obtain the target disparity map disp∈R of the left eye image. C×H×W Where C represents the number of channels in the target disparity map disp, where C is 1; H represents the height of the target disparity map disp; and W represents the width of the target disparity map disp. After upsampling, the height and width of the target disparity map in the left-eye image are consistent with the height and width of the left-eye image. Furthermore, the electronic device can use the target disparity map of the left-eye image as the disparity map of the aforementioned monocular image.

[0054] It should be noted that using the FoundationStereo model for semantic scene completion has two advantages. First, it reduces the difficulty of recovering 3D spatial structure from 2D images. In particular, the FoundationStereo model can accurately determine the depth and spatial geometry of objects when dealing with occluded regions and similar texture regions in 2D images. Second, it can compensate for the shortcomings of traditional semantic scene completion tasks in acquiring and utilizing depth information, improving the accuracy of the semantic scene completion task's understanding of 3D spatial structure and thus enhancing the overall precision of the semantic scene completion task.

[0055] Step 103: Perform a disparity dimension to depth dimension conversion on the disparity map of the monocular image to obtain the depth map of the monocular image.

[0056] For example, after obtaining the disparity map of a monocular image, the electronic device can further acquire the parameters of the binocular image sensor that acquired the binocular image. Then, based on the parameters of the binocular image sensor, the electronic device can perform a disparity-to-depth dimension conversion on the disparity map of the monocular image to obtain the depth map of the monocular image. The process of performing this disparity-to-depth dimension conversion on the disparity map of the monocular image based on the parameters of the binocular image sensor to obtain the depth map will be described in detail later and will not be repeated here.

[0057] Step 104: Process the multi-scale feature map and depth map of the monocular image to obtain the first semantic feature map of the monocular image.

[0058] For example, an electronic device can perform feature fusion processing on the multi-scale image feature maps of a monocular image to obtain a second semantic feature map of the monocular image. Then, the electronic device can generate a geometric prior matrix for the monocular image based on the depth map of the monocular image. This geometric prior matrix can characterize the spatial geometric relationships between objects in the driving scene in the front-back, horizontal, and vertical directions, thereby improving the semantic scene completion task's understanding of these spatial geometric relationships. Subsequently, the electronic device can perform feature fusion processing on the geometric prior matrix and the second semantic feature map to obtain a first semantic feature map of the monocular image. This first semantic feature map contains the semantic information and spatial geometric information of objects in the driving scene, and the semantic information and spatial geometric information of the same object are correlated, further improving the consistency of the semantic scene completion task's understanding of the semantic information and spatial geometric information of objects in the driving scene, and avoiding confusion between the semantic information and spatial geometric information of different objects.

[0059] Step 105: Perform depth feature extraction processing on the disparity probability distribution map of the monocular image to obtain the depth probability distribution map of the monocular image.

[0060] For example, the electronic device performs depth feature extraction processing on the disparity probability distribution map of the monocular image, converting the disparity probability distribution map into a depth probability distribution map, thus obtaining the depth probability distribution map of the monocular image. This provides a data foundation for subsequent semantic scene completion tasks to further generate semantic occupancy data based on the depth probability distribution map. The process of the electronic device performing depth feature extraction processing on the disparity probability distribution map of the monocular image to obtain the depth probability distribution map will be described in detail later and will not be repeated here.

[0061] Step 106: Based on the depth map, first semantic feature map and depth probability distribution map of the monocular image, determine the voxel occupancy data and voxel semantic data of the semantic occupancy data in order to complete the semantic scene.

[0062] For example, an electronic device can determine the voxel occupancy data and voxel semantic data of semantic occupancy data based on the depth map, the first semantic feature map, and the depth probability distribution map of a monocular image. Then, the electronic device can perform semantic scene completion based on the voxel occupancy data and voxel semantic data. The process of determining the voxel occupancy data and voxel semantic data of semantic occupancy data based on the depth map, the first semantic feature map, and the depth probability distribution map of the monocular image for semantic scene completion will be described in detail later and will not be repeated here.

[0063] In this embodiment, the electronic device acquires binocular images, performs depth estimation processing on the binocular images, and obtains a multi-scale image feature map, a disparity probability distribution map, and a disparity map for the monocular image. Then, the electronic device performs a disparity-to-depth dimension conversion processing on the disparity map of the monocular image to obtain a depth map. Next, the electronic device processes the multi-scale feature map and depth map of the monocular image to obtain a first semantic feature map, and performs depth feature extraction processing on the disparity probability distribution map of the monocular image to obtain a depth probability distribution map. Finally, based on the depth map, the first semantic feature map, and the depth probability distribution map of the monocular image, the electronic device determines the voxel occupancy data and voxel semantic data of the semantic occupancy data for semantic scene completion. Therefore, in this disclosure, the electronic device can obtain a more accurate depth map based on depth estimation processing of the binocular images, and further combine the depth map and multi-scale image feature map to perform semantic feature extraction processing based on geometric priors to obtain richer and more accurate semantic features, thereby improving the accuracy of subsequent semantic scene completion.

[0064] Figure 2 This is a schematic diagram of a semantic scene completion network provided in an exemplary embodiment of this disclosure. For example... Figure 2As shown, the semantic scene completion network 200 includes: a binocular stereo depth estimation model 210, a depth feature extraction model 220, a semantic feature extraction model based on geometric priors 230, and a semantic scene completion model 240. Specifically, the binocular stereo depth estimation model 210 receives binocular images acquired by a binocular image sensor, performs depth estimation processing on the binocular images, and obtains a multi-scale image feature map, a disparity probability distribution map, and a disparity map for the monocular image. The binocular stereo depth estimation model 210 also performs disparity-to-depth dimension conversion processing on the disparity map of the monocular image to obtain a depth map of the monocular image. The depth feature extraction model 220 performs depth feature extraction processing on the disparity probability distribution map of the monocular image to obtain a depth probability distribution map of the monocular image. The semantic feature extraction model 230 based on geometric priors processes the multi-scale feature map and depth map of the monocular image to obtain a semantic feature map (i.e., a first semantic feature map) of the monocular image. The semantic scene completion model 240 is used to process the depth map, first semantic feature map and depth probability distribution map of monocular images to obtain voxel occupancy data and voxel semantic data of semantic occupancy data, and to perform semantic scene completion based on voxel occupancy data and voxel semantic data.

[0065] Steps 102 and 103 are executed by the binocular stereo depth estimation model 210, step 104 is executed by the semantic feature extraction model 230 based on geometric priors, step 105 is executed by the depth feature extraction model 220, and step 106 is executed by the semantic scene completion model 240.

[0066] For example, the stereo depth estimation model 210 in this disclosure may include the FoundationStereo model provided in the above embodiments, the depth feature extraction model 220 may be implemented based on a multilayer perceptron (MLP), the semantic feature extraction model 230 based on geometric prior may be implemented through a semantic feature extraction network (context net), and the semantic scene completion model 240 may be implemented based on a feature dimensionality enhancement module (view transformation), a 3D local encoder, and a decoding head. The specific implementation process will be described in detail below.

[0067] like Figure 3 As shown above, in the above Figure 1 Based on the illustrated embodiment, step 104 may include the following steps:

[0068] Step 1041: Perform feature fusion processing on the multi-scale image feature map of the monocular image to obtain the second semantic feature map of the monocular image.

[0069] For example, to improve the understanding of semantic information in driving scenarios in semantic scene completion tasks, after acquiring multi-scale image feature maps of a monocular image, the electronic device can further perform feature fusion processing on the multi-scale image feature maps based on an FPN (Feature Pyramid Network) to obtain a second semantic feature map of the monocular image. The multi-scale image feature maps contain varying degrees of richness in image and semantic feature information. For instance, small-scale image feature maps have higher spatial resolution and contain richer image feature information, while large-scale image feature maps have lower spatial resolution but contain richer semantic feature information. By fusing the image feature maps of different scales using an FPN network, the electronic device obtains a semantic feature map that retains the rich image feature information from the small-scale image feature maps while incorporating the rich semantic feature information from the large-scale image feature maps. The image size of the second semantic feature map of the monocular image is (C, H / 8, W / 8), where H / 8 represents the height of the second semantic feature map, W / 8 represents the width of the second semantic feature map, and C represents the number of channels in the second semantic feature map. Compared to the image size of a monocular image, the height and width of the second semantic feature map are both 1 / 8 of those of a monocular image.

[0070] Step 1042: Generate the geometric prior matrix of the monocular image based on the depth map of the monocular image.

[0071] For example, to improve the understanding of spatial geometric information in driving scenes for semantic scene completion tasks, electronic devices can determine the spatial geometric relationships between objects in the driving scene in the depth and planar directions (including horizontal and vertical directions) based on the depth map of a monocular image, generating a geometric prior matrix for the monocular image. This geometric prior matrix can characterize the spatial geometric relationships between objects in the driving scene in the front-back, horizontal, and vertical directions. The process of generating the geometric prior matrix for the monocular image based on the depth map will be described in detail later and will not be repeated here.

[0072] Step 1043: Perform feature fusion processing on the geometric prior matrix and the second semantic feature map to obtain the first semantic feature map of the monocular image.

[0073] For example, after obtaining the geometric prior matrix and the second semantic feature map, the electronic device can further perform feature fusion processing on the geometric prior matrix and the second semantic feature map to obtain the first semantic feature map of the monocular image. The height and width of the first semantic feature map are consistent with those of the second semantic feature map. The first semantic feature map contains semantic information and spatial geometric information of objects in the driving scene, and the semantic information and spatial geometric information of the same object are correlated.

[0074] In this embodiment, the electronic device performs feature fusion processing on the multi-scale image feature maps of a monocular image to obtain a second semantic feature map of the monocular image. Then, based on the depth map of the monocular image, the electronic device generates a geometric prior matrix for the monocular image. Subsequently, the electronic device performs feature fusion processing on the geometric prior matrix and the second semantic feature map to obtain a first semantic feature map of the monocular image. This improves the semantic scene completion task's understanding of the spatial geometric relationships between objects in the driving scene in the front-back, horizontal, and vertical directions, and further enhances the consistency of the semantic scene completion task's understanding of the semantic information and spatial geometric information of objects in the driving scene, avoiding confusion between the semantic information and spatial geometric information of different objects.

[0075] In the above Figure 3 Based on the illustrated embodiment, step 1042 may include the following steps:

[0076] Step 1: According to the preset depth block division strategy, the depth map of the monocular image is divided into depth blocks to obtain multiple depth blocks.

[0077] For example, to improve the understanding of spatial geometric information in a driving scene during semantic scene completion tasks, after acquiring the depth map of a monocular image, the electronic device can perform depth block partitioning processing on the depth map of the monocular image according to a preset depth block partitioning strategy to obtain multiple depth blocks. For instance, the electronic device can divide a depth map of image size H×W into (H / 8)×(W / 8) depth blocks according to the preset depth block partitioning strategy.

[0078] Step 2: For any depth block among multiple depth blocks, determine the depth distance and spatial distance between the depth block and other depth blocks, and obtain the depth distance matrix and spatial distance matrix.

[0079] For example, after dividing the depth map into multiple depth blocks, the electronic device can calculate the depth distance and spatial distance between any depth block and other depth blocks. These other depth blocks can be any of the multiple depth blocks. The electronic device can define any two depth blocks as a depth block pair; the two depth blocks in a pair can be the same depth block or different depth blocks. Then, after obtaining the depth distances between the two depth blocks in all depth block pairs, the electronic device can generate a depth distance matrix. Each element in the depth distance matrix represents the depth distance between the two depth blocks in a depth block pair. The depth distance matrix reflects the differences in depth direction between different depth blocks and can represent the spatial geometric relationships between objects in the forward and backward directions in a driving scene. Similarly, after obtaining the spatial distances between the two depth blocks in all depth block pairs, the electronic device can generate a spatial distance matrix. Each element in the spatial distance matrix represents the spatial distance between the two depth blocks in a depth block pair. The spatial distance matrix reflects the differences in planar dimensions between different depth blocks and can represent the spatial geometric relationships between objects in the horizontal and vertical directions in a driving scene.

[0080] For example, an electronic device divides a depth map of image size H×W into (H / 8)×(W / 8) depth blocks. Correspondingly, the dimensions of the depth distance matrix and the spatial distance matrix are both [(H / 8×W / 8), (H / 8×W / 8)].

[0081] In some embodiments, the electronic device determines the depth distance between two depth blocks (referred to as the first depth block and the second depth block for easy distinction) contained in any depth block pair as follows:

[0082] Step A: The electronic device can determine the depth distance between the first pixel and the second pixel as the absolute difference between the depth value of the first pixel in the first depth block and the depth value of the second pixel in the second depth block. Here, the first pixel is any pixel in the first depth block, and the second pixel is any pixel in the second depth block. The formula corresponding to the processing in Step A above is: D uj,mn =|z ij -z mn |. Among them, z ij z represents the depth value of the first pixel at coordinates (i, j) in the first depth block. mn This represents the depth value of the second pixel at coordinates (m, n) in the second depth block. ij,mn This represents the depth distance between the first pixel and the second pixel.

[0083] Step B: The electronic device can determine the average depth distance based on the depth distance between all first pixels and second pixels, and determine the average depth distance as the depth distance between the first depth block and the second depth block.

[0084] In some embodiments, the electronic device determines the spatial distance between two depth blocks (referred to as the first depth block and the second depth block for easy distinction) contained in any depth block pair as follows:

[0085] Step A: The electronic device can determine the horizontal distance between the first and second pixels by the absolute difference between the horizontal coordinates of the first pixel in the first depth block and the horizontal coordinates of the second pixel in the second depth block. Similarly, the electronic device can also determine the vertical distance between the first and second pixels by the absolute difference between the vertical coordinates of the first and second pixels. Then, the electronic device can determine the spatial distance between the first and second pixels by the sum of the horizontal and vertical distances between the first and second pixels. Here, the first pixel is any pixel in the first depth block, and the second pixel is any pixel in the second depth block. The formula corresponding to the processing in Step A is: S ij,mn = |im|+|jn|. Where i and j represent the x and y coordinates of the first pixel at coordinates (i, j) in the first depth block, respectively; m and n represent the x and y coordinates of the second pixel at coordinates (m, n) in the second depth block, respectively; S ij,mn This represents the spatial distance between the first pixel and the second pixel.

[0086] Step B: The electronic device can determine the average spatial distance based on the spatial distance between all first pixels and second pixels, and define the average spatial distance as the spatial distance between the first depth block and the second depth block.

[0087] Step 3: Based on the weight matrix, perform a weighted summation operation on the depth distance matrix and the spatial distance matrix to obtain the geometric prior matrix of the monocular image.

[0088] For example, the weight matrix includes a first weight matrix corresponding to the depth distance matrix and a second weight matrix corresponding to the spatial distance matrix. The first and second weight matrices can be learnable weight matrices that are continuously updated during model training using optimization algorithms (such as gradient descent). After the electronic device determines the depth distance matrix and the spatial distance matrix, it can further perform a weighted summation operation on the depth distance matrix and the spatial distance matrix based on the first and second weight matrices to obtain the geometric prior matrix of the monocular image. The geometric prior matrix includes both the differences in the depth direction between different depth blocks reflected by the depth distance matrix and the differences in the plane between different depth blocks reflected by the spatial distance matrix, and can characterize the spatial geometric relationships between objects in the driving scene in the front-back, horizontal, and vertical directions.

[0089] In this embodiment, the electronic device performs depth block segmentation on the depth map of a monocular image to obtain multiple depth blocks. Then, the electronic device determines the depth distance and spatial distance between any depth block and other depth blocks, obtaining a depth distance matrix and a spatial distance matrix. Subsequently, based on a weight matrix, the electronic device performs a weighted summation operation on the depth distance matrix and the spatial distance matrix to obtain the geometric prior matrix of the monocular image. Thus, based on the geometric prior matrix, the semantic scene completion task can improve its understanding of the spatial geometric relationships between objects in the driving scene in the front-back, horizontal, and vertical directions.

[0090] In the above Figure 3 Based on the illustrated embodiment, step 1043 may include the following steps:

[0091] Based on geometric self-attention, feature fusion processing is performed on the geometric prior matrix and the second semantic feature map to obtain the first semantic feature map of the monocular image.

[0092] For example, after obtaining the geometric prior matrix and the second semantic feature map of a monocular image, the electronic device can further perform feature fusion processing on the geometric prior matrix and the second semantic feature map based on a geometry self-attention mechanism to obtain the first semantic feature map of the monocular image. The image size of the first semantic feature map of the monocular image is (C, H / 8, W / 8). Here, H / 8 represents the height of the first semantic feature map, W / 8 represents the width of the first semantic feature map, and C represents the number of channels in the first semantic feature map. The height and width of the first semantic feature map are consistent with those of the second semantic feature map. The first semantic feature map contains semantic information and spatial geometric information of objects in the driving scene, and the semantic information and spatial geometric information of the same object are associated.

[0093] In this embodiment of the disclosure, the electronic device performs feature fusion processing on the geometric prior matrix and the second semantic feature map based on geometric self-attention to obtain a first semantic feature map of the monocular image. Thus, based on the first semantic feature map, the consistency of the semantic scene completion task's understanding of the semantic information and spatial geometric information of objects in the driving scene can be improved, avoiding confusion between the semantic information and spatial geometric information of different objects.

[0094] like Figure 4 As shown above, in the above Figure 1 Based on the illustrated embodiment, step 105 may include the following steps:

[0095] Step 1051: Using a multilayer perceptron, the disparity probability distribution map of the monocular image is mapped from the disparity dimension to the depth dimension to obtain the first depth feature map of the monocular image.

[0096] For example, after an electronic device obtains the disparity probability distribution map of a monocular image, it can use a multilayer perceptron (MLP) to non-linearly transform the disparity features of each pixel in the disparity probability distribution map, thereby mapping each pixel from the disparity dimension to the depth dimension and obtaining the first depth feature map of the monocular image. Compared to the traditional disparity-to-depth dimension mapping based on the parameters of a binocular image sensor, using a data-driven learning-based multilayer perceptron for this mapping improves the cross-device and cross-scene generalization ability of the semantic scene completion task. The image size of the first depth feature map is (D, H / 4, W / 4), where H / 4 represents the height of the first depth feature map, W / 4 represents the width of the first depth feature map, and D represents the number of depth channels in the first depth feature map.

[0097] Step 1052: Perform multi-level feature fusion processing on the first depth feature map of the monocular image to obtain the second depth feature map of the monocular image.

[0098] For example, after obtaining the first depth feature map of a monocular image, the electronic device can perform multi-level feature fusion processing on the first depth feature map of the monocular image to obtain the second depth feature map of the monocular image. Specifically, the electronic device can perform multi-level feature fusion processing on the first depth feature map of the monocular image using a stereo depth volume encoder. The image size of the second depth feature map is (D, H / 4, W / 4), where H / 4 represents the height of the second depth feature map, W / 4 represents the width of the second depth feature map, and D represents the number of depth channels in the second depth feature map.

[0099] Step 1053: In the depth dimension, the second depth feature map of the monocular image is normalized to obtain the depth probability distribution map of the monocular image.

[0100] For example, after obtaining the second depth feature map of a monocular image, the electronic device can further normalize the second depth feature map of the monocular image in the depth dimension to obtain a depth probability distribution map of the monocular image. The image size of the depth probability distribution map of the monocular image is (D, H / 4, W / 4). Here, H / 4 represents the height of the depth probability distribution map, W / 4 represents the width of the depth probability distribution map, and D represents the number of depth channels in the depth probability distribution map. The normalization process can be performed using the softmax function, or other normalization functions; this embodiment of the disclosure is not limited to any particular method.

[0101] It should be noted that the image size of the first semantic feature map of the monocular image is (C, H / 8, W / 8). To ensure that the image sizes of the first semantic feature map and the depth probability distribution map of the monocular image are consistent, and to facilitate subsequent semantic scene completion tasks, the electronic device further performs a 2x downsampling process on the depth probability distribution map, resulting in a downsampled depth probability distribution map. The image size of the downsampled depth probability distribution map is (D, H / 8, W / 8).

[0102] In this embodiment, the electronic device uses a multilayer perceptron to map the disparity probability distribution map of a monocular image from the disparity dimension to the depth dimension, obtaining a first depth feature map of the monocular image. Then, the electronic device performs multi-level feature fusion processing on the first depth feature map of the monocular image to obtain a second depth feature map. Subsequently, the electronic device normalizes the second depth feature map of the monocular image in the depth dimension to obtain a depth probability distribution map of the monocular image. In this way, the electronic device converts the disparity probability distribution map into a depth probability distribution map, providing a data foundation for subsequent semantic scene completion tasks to generate semantic occupancy data based on the depth probability distribution map.

[0103] like Figure 5 As shown above, in the above Figure 1 Based on the illustrated embodiment, step 103 may include the following steps:

[0104] Step 1031: Obtain the parameters of the image sensor that acquires binocular images.

[0105] For example, the electronic device acquires parameters of a binocular image sensor that captures binocular images. These parameters include baseline distance and focal length. The baseline distance of the binocular image sensor is the horizontal distance between the optical centers of the left and right eye image sensors. The greater the baseline distance of the binocular image sensor, the greater the parallax of the target object in the left and right eye images.

[0106] Step 1032: For each pixel in the disparity map of the monocular image, determine the depth value of the pixel based on the parameters of the image sensor and the disparity value of the pixel.

[0107] For example, after acquiring the baseline distance and focal length of the binocular image sensor, the electronic device can further determine the depth value of each pixel in the disparity map of the monocular image by dividing the product of the baseline distance and focal length by the disparity value of that pixel. Here, the disparity value represents the horizontal displacement of that pixel in the left and right eye images. The formula corresponding to the above processing is: depth = baseline × f x / disparity. Where depth represents the depth value of a pixel in the depth map, baseline represents the baseline distance of the stereo image sensors, and f... x This represents the focal length of the binocular image sensor, and disparity represents the disparity value of a pixel in the disparity map.

[0108] Step 1033: Generate a depth map of the monocular image based on the depth values ​​of the pixels.

[0109] For example, after the electronic device determines the depth value of each pixel, for each pixel, the electronic device can further map the depth value of the pixel to the depth value of the corresponding pixel in the depth map of the monocular image based on the coordinates of each pixel, thereby generating the depth map of the monocular image.

[0110] In this embodiment, for the disparity map of the monocular image output by the binocular stereo depth estimation model, the electronic device performs a disparity dimension to depth dimension conversion on the disparity map of the monocular image based on the parameters of the binocular image sensor to obtain the depth map of the monocular image. In this way, the electronic device can obtain accurate spatial depth information in the driving scene through the depth map of the monocular image, and also provides accurate spatial depth information for the semantic scene completion task, thereby improving the accuracy of subsequent semantic occupancy data generation.

[0111] like Figure 6 As shown above, in the above Figure 1 Based on the illustrated embodiment, step 106 may include the following steps:

[0112] Step 1061: Process the first semantic feature map, depth probability distribution map and depth map of the monocular image to obtain the voxel occupancy data of the semantic occupancy data.

[0113] For example, in step one, the electronic device can use a context-aware query generator to perform voxel query processing on the first semantic feature map and the depth probability distribution map to obtain a three-dimensional voxel query V. Q ∈R C×H×W×D Where H represents a three-dimensional voxel query V Q The length of can also represent a three-dimensional voxel query V. Q The number of voxel grids along the length direction, W represents the 3D voxel query V. Q The width can also represent a three-dimensional voxel query V. Q The number of voxel grids in the width direction, where D represents the 3D voxel query V. Q The height can also represent the three-dimensional voxel query V. Q The number of voxel grids in the height direction, C represents the 3D voxel query V. Q The number of channels in the mid-voxel mesh. Step two, the electronic device queries V based on the three-dimensional voxel. Q Depth information-based query proposal processing is performed using depth maps and intrinsic and extrinsic parameters of stereo image sensors to obtain binary 3D voxel mesh data M∈{1,0}. H×W×D In this context, 1 indicates that the voxel grid is occupied, and 0 indicates that the voxel grid is not occupied. Voxel grids in an occupied state are called occupied voxel grids (also known as visible voxel grids), and voxel grids in a non-occupied state are called non-occupied voxel grids (also known as invisible voxel grids). Since the binary 3D voxel grid data M represents the occupancy state of each voxel grid in the semantic occupancy data, the electronic device can determine the above binary 3D voxel grid data M as the voxel occupancy data of the semantic occupancy data.

[0114] Step 1062: Process the first semantic feature map, the depth probability distribution map, the voxel occupancy data and the preset semantic category data to obtain the voxel semantic data of the semantic occupancy data.

[0115] For example, in step one, the electronic device can query V from voxels based on binary three-dimensional voxel mesh data M (i.e., voxel occupancy data). QIn the first step, the electronic device filters the occupied voxel grid to obtain the query proposal Q. Here, the query proposal Q represents the initial spatial position of the occupied voxel grid in 3D space. In the second step, the electronic device can perform an outer product operation on the first semantic feature map and the depth probability distribution map to obtain a high-dimensional feature map F. In the third step, the electronic device can perform feature fusion processing on the high-dimensional feature map F and the query proposal Q based on 3D deformable cross-attention to obtain the feature-fused query proposal. Among them, the query proposal after feature fusion It includes semantic and spatial geometric information (including depth and spatial information) from the first semantic feature map and the depth probability distribution map. Step four: The electronic device fuses the features into a query proposal. Voxel query V Q The non-occupied voxel meshes in the data are merged to obtain sparse 3D voxel mesh data F. 3D ∈R C ×H×W×D Step five: The electronic device can apply deformable self-attention to the sparse three-dimensional voxel mesh data F. 3D Diffusion processing is performed to obtain dense three-dimensional voxel mesh data. Step six: Electronic devices can use 3D convolutional networks (such as ResNet networks) to process dense 3D voxel mesh data. Feature extraction is performed to obtain multi-scale three-dimensional voxel mesh data. Step seven: The electronic device can use a three-dimensional feature pyramid network to perform feature fusion processing on the multi-scale three-dimensional voxel mesh data to obtain high-dimensional three-dimensional voxel mesh data. In this process, the electronic device uses a three-dimensional feature pyramid network to perform feature fusion processing on multi-scale three-dimensional voxel mesh data to obtain high-dimensional three-dimensional voxel mesh data. The features of 3D voxel mesh data at various scales are integrated, containing more comprehensive driving scene information. Step eight: The electronic device can, based on a preset semantic category N... class For high-dimensional three-dimensional voxel mesh data Voxel semantic prediction processing is performed to obtain three-dimensional voxel semantic feature data. Among them, semantic categories N class This includes one air category and multiple non-air categories. Non-air categories may include motor vehicle categories, non-motor vehicle categories, pedestrian categories, building categories, road categories, etc., which are not limited in this embodiment. Step nine: The electronic device can use a three-dimensional convolutional network to normalize the three-dimensional voxel semantic feature data along the semantic category dimension to obtain three-dimensional voxel semantic probability distribution data. Among them, the three-dimensional voxel semantic probability distribution data P semThe probability of each voxel grid belonging to each semantic category is represented. Normalization can be performed using the softmax function or other normalization functions; this embodiment is not limited to any particular function. Step ten: For each voxel grid, the electronic device can determine the semantic category corresponding to that voxel grid as the semantic category with the highest probability, thereby obtaining the voxel semantic data of the semantic occupancy data.

[0116] Step 1063: Generate semantic occupancy data based on voxel occupancy data and voxel semantic data.

[0117] For example, after obtaining voxel occupancy data and voxel semantic data, the electronic device can further perform semantic scene completion based on the voxel occupancy data and voxel semantic data.

[0118] In this embodiment, the electronic device processes the first semantic feature map, depth probability distribution map, and depth map of a monocular image to obtain voxel occupancy data of semantic occupancy data. Then, the electronic device processes the first semantic feature map, depth probability distribution map, voxel occupancy data, and preset semantic category data to obtain voxel semantic data of semantic occupancy data. Afterward, the electronic device performs semantic scene completion based on the voxel occupancy data and voxel semantic data. Thus, the electronic device can determine the voxel occupancy data and voxel semantic data of semantic occupancy data based on the depth map, first semantic feature map, and depth probability distribution map of a monocular image to perform semantic scene completion.

[0119] Exemplary device

[0120] Figure 7 This is a schematic diagram of the structure of a semantic scene completion device provided in an exemplary embodiment of the present disclosure. The semantic scene completion device 700 includes: a binocular image acquisition module 710, a binocular stereo depth estimation module 720, a semantic feature extraction module 730, a depth feature extraction module 740, and a semantic scene completion module 750.

[0121] A binocular image acquisition module 710 is used to acquire binocular images;

[0122] The binocular stereo depth estimation module 720 is used to perform depth estimation processing on the binocular image to obtain a multi-scale image feature map, a disparity probability distribution map, and a disparity map of the monocular image; the monocular image is either the left or right eye image in the binocular image.

[0123] The binocular stereo depth estimation module 720 is also used to perform disparity dimension to depth dimension conversion processing on the disparity map of the monocular image to obtain the depth map of the monocular image.

[0124] The semantic feature extraction module 730 is used to process the multi-scale feature map and depth map of the monocular image to obtain the first semantic feature map of the monocular image;

[0125] The depth feature extraction module 740 is used to perform depth feature extraction processing on the disparity probability distribution map of the monocular image to obtain the depth probability distribution map of the monocular image.

[0126] The semantic scene completion module 750 is used to determine the voxel occupancy data and voxel semantic data of the semantic occupancy data based on the depth map, the first semantic feature map and the depth probability distribution map of the monocular image, so as to perform semantic scene completion.

[0127] In some embodiments, the semantic feature extraction module 730 includes:

[0128] The first feature fusion unit is used to perform feature fusion processing on the multi-scale image feature map of the monocular image to obtain the second semantic feature map of the monocular image.

[0129] A geometric prior matrix generation unit is used to generate a geometric prior matrix of the monocular image based on the depth map of the monocular image.

[0130] The second feature fusion unit is used to perform feature fusion processing on the geometric prior matrix and the second semantic feature map to obtain the first semantic feature map of the monocular image.

[0131] In some embodiments, the geometric prior matrix generation unit is specifically used for:

[0132] According to the preset depth block partitioning strategy, the depth map of the monocular image is divided into depth blocks to obtain multiple depth blocks;

[0133] For any one of the plurality of depth blocks, determine the depth distance and spatial distance between the depth block and other depth blocks, and obtain a depth distance matrix and a spatial distance matrix;

[0134] Based on the weight matrix, the depth distance matrix and the spatial distance matrix are weighted and summed to obtain the geometric prior matrix of the monocular image.

[0135] In some embodiments, the second feature fusion unit is specifically used for:

[0136] Based on geometric self-attention, feature fusion processing is performed on the geometric prior matrix and the second semantic feature map to obtain the first semantic feature map of the monocular image.

[0137] In some embodiments, the depth feature extraction module 740 includes:

[0138] The dimension mapping unit is used to perform disparity dimension to depth dimension mapping processing on the disparity probability distribution map of the monocular image through a multilayer perceptron to obtain the first depth feature map of the monocular image.

[0139] A multi-layer feature fusion unit is used to perform multi-layer feature fusion processing on the first depth feature map of the monocular image to obtain the second depth feature map of the monocular image.

[0140] The normalization unit is used to normalize the second depth feature map of the monocular image in the depth dimension to obtain the depth probability distribution map of the monocular image.

[0141] In some embodiments, the binocular stereo depth estimation module 720 includes:

[0142] The parameter acquisition unit is used to acquire the parameters of the image sensor that acquires the binocular images;

[0143] A depth value determination unit is used to determine the depth value of each pixel in the disparity map of the monocular image based on the parameters of the image sensor and the disparity value of the pixel.

[0144] A depth map generation unit is used to generate a depth map of the monocular image based on the depth values ​​of the pixels.

[0145] In some embodiments, the semantic scene completion module 750 includes:

[0146] The voxel occupancy data determination unit is used to process the first semantic feature map, depth probability distribution map and depth map of the monocular image to obtain voxel occupancy data of semantic occupancy data.

[0147] A voxel semantic data determination unit is used to process the first semantic feature map, the depth probability distribution map, the voxel occupancy data and the preset semantic category data to obtain the voxel semantic data of the semantic occupancy data.

[0148] The semantic scene completion unit is used to complete the semantic scene based on the voxel occupancy data and the voxel semantic data.

[0149] The beneficial technical effects corresponding to the exemplary embodiments of this device can be found in the corresponding beneficial technical effects of the exemplary method section above, and will not be repeated here.

[0150] Exemplary electronic devices

[0151] Figure 8 A structural diagram of an electronic device provided in an embodiment of this disclosure includes at least one processor 11 and a memory 12.

[0152] The processor 11 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 10 to perform desired functions.

[0153] The memory 12 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 11 may execute one or more computer program instructions to implement the semantic scene completion methods and / or other desired functions of the various embodiments of this disclosure described above.

[0154] In one example, the electronic device 10 may also include an input device 13 and an output device 14, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).

[0155] The input device 13 may also include, for example, a keyboard, a mouse, etc.

[0156] The output device 14 can output various information to the outside, including, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0157] Of course, for the sake of simplicity, Figure 8 Only some of the components of the electronic device 10 relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device 10 may include any other suitable components depending on the specific application.

[0158] Exemplary computer program products and computer-readable storage media

[0159] In addition to the methods and devices described above, embodiments of this disclosure may also provide a computer program product, including computer program instructions, which, when executed by a processor, cause the processor to perform the steps in the semantic scene completion methods of the various embodiments of this disclosure described in the "Exemplary Methods" section above.

[0160] Computer program products can be written in any combination of one or more programming languages ​​to perform the operations of embodiments of this disclosure. These programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0161] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions that, when executed by a processor, cause the processor to perform the steps in the semantic scene completion methods of the various embodiments of this disclosure described in the "Exemplary Methods" section above.

[0162] Computer-readable storage media may take the form of any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may include, but is not limited to, systems, apparatuses, or devices that are electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0163] The basic principles of this disclosure have been described above with reference to specific embodiments. However, the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.

[0164] Various modifications and variations can be made to this disclosure without departing from its spirit and scope. Therefore, this disclosure is also intended to include such modifications and variations if they fall within the scope of the claims of this disclosure and their equivalents.

Claims

1. A semantic scene completion method, comprising: Acquire binocular images; The binocular images are subjected to depth estimation processing to obtain multi-scale image feature maps, disparity probability distribution maps, and disparity maps of the monocular images; The monocular image is either the left or right eye image in the binocular image; The disparity map of the monocular image is converted from the disparity dimension to the depth dimension to obtain the depth map of the monocular image. The multi-scale feature map and depth map of the monocular image are processed to obtain the first semantic feature map of the monocular image; Depth feature extraction is performed on the disparity probability distribution map of the monocular image to obtain the depth probability distribution map of the monocular image; Based on the depth map, first semantic feature map, and depth probability distribution map of the monocular image, the voxel occupancy data and voxel semantic data of the semantic occupancy data are determined in order to complete the semantic scene.

2. The method according to claim 1, wherein, The process of processing the multi-scale feature map and depth map of the monocular image to obtain the first semantic feature map of the monocular image includes: The multi-scale image feature map of the monocular image is subjected to feature fusion processing to obtain the second semantic feature map of the monocular image; Based on the depth map of the monocular image, generate the geometric prior matrix of the monocular image; The geometric prior matrix and the second semantic feature map are subjected to feature fusion processing to obtain the first semantic feature map of the monocular image.

3. The method according to claim 2, wherein, The step of generating the geometric prior matrix of the monocular image based on the depth map of the monocular image includes: According to a preset depth block partitioning strategy, the depth map of the monocular image is partitioned into multiple depth blocks. For any one of the plurality of depth blocks, determine the depth distance and spatial distance between the depth block and other depth blocks, and obtain a depth distance matrix and a spatial distance matrix; Based on the weight matrix, the depth distance matrix and the spatial distance matrix are weighted and summed to obtain the geometric prior matrix of the monocular image.

4. The method according to claim 2, wherein, The step of performing feature fusion processing on the geometric prior matrix and the second semantic feature map to obtain the first semantic feature map of the monocular image includes: Based on geometric self-attention, feature fusion processing is performed on the geometric prior matrix and the second semantic feature map to obtain the first semantic feature map of the monocular image.

5. The method according to claim 1, wherein, The step of performing depth feature extraction processing on the disparity probability distribution map of the monocular image to obtain the depth probability distribution map of the monocular image includes: By using a multilayer perceptron, the disparity probability distribution map of the monocular image is mapped from the disparity dimension to the depth dimension to obtain the first depth feature map of the monocular image. A multi-level feature fusion process is performed on the first depth feature map of the monocular image to obtain the second depth feature map of the monocular image. In the depth dimension, the second depth feature map of the monocular image is normalized to obtain the depth probability distribution map of the monocular image.

6. The method according to claim 1, wherein, The step of converting the disparity map of the monocular image from the disparity dimension to the depth dimension to obtain the depth map of the monocular image includes: Obtain the parameters of the image sensor that acquires the binocular images; For each pixel in the disparity map of the monocular image, the depth value of the pixel is determined based on the parameters of the image sensor and the disparity value of the pixel. A depth map of the monocular image is generated based on the depth values ​​of the pixels.

7. The method according to claim 1, wherein, The process of determining voxel occupancy data and voxel semantic data based on the depth map, first semantic feature map, and depth probability distribution map of the monocular image for semantic scene completion includes: The first semantic feature map, depth probability distribution map and depth map of the monocular image are processed to obtain the voxel occupancy data of the semantic occupancy data; The first semantic feature map, the depth probability distribution map, the voxel occupancy data, and the preset semantic category data are processed to obtain the voxel semantic data of the semantic occupancy data; Based on the voxel occupancy data and the voxel semantic data, semantic scene completion is performed.

8. A semantic scene completion device, comprising: A binocular image acquisition module is used to acquire binocular images; A binocular stereo depth estimation module is used to perform depth estimation processing on the binocular image to obtain a multi-scale image feature map, a disparity probability distribution map, and a disparity map of the monocular image; the monocular image is either the left or right eye image in the binocular image. The binocular stereo depth estimation module is also used to perform disparity dimension to depth dimension conversion processing on the disparity map of the monocular image to obtain the depth map of the monocular image. The semantic feature extraction module is used to process the multi-scale feature map and depth map of the monocular image to obtain the first semantic feature map of the monocular image; The depth feature extraction module is used to perform depth feature extraction processing on the disparity probability distribution map of the monocular image to obtain the depth probability distribution map of the monocular image. The semantic scene completion module is used to determine the voxel occupancy data and voxel semantic data of the semantic occupancy data based on the depth map, the first semantic feature map and the depth probability distribution map of the monocular image, so as to perform semantic scene completion.

9. A computer-readable storage medium storing a computer program for performing the semantic scene completion method according to any one of claims 1-7.

10. An electronic device, the electronic device comprising: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the semantic scene completion method according to any one of claims 1-7.

Citation Information

Cited By

  • Point cloud processing method and electronic equipment

    CN121437822A

  • Point cloud processing method and electronic device

    CN121437822B