3DGS and neural SDF semantic level hybrid reconstruction method for occlusion perception

Through the hybrid reconstruction method of 3DGS and neural SDF, the problems of uneven point cloud distribution and occlusion in complex scenes are solved, and high-quality 3D reconstruction and intelligent completion of occluded areas are achieved. It is suitable for virtual reality, augmented reality, architectural design, cultural heritage protection and autonomous driving.

CN120655823APending Publication Date: 2025-09-16TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL +1
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510718928.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing 3DGS-based reconstruction methods have problems such as uneven point cloud distribution, blurred boundaries, and broken geometric structures in complex scenes. In particular, it is difficult to accurately restore the edges of objects and fine structure areas, and it is difficult to deal with occlusion problems between multiple objects.

Method used

An occlusion-aware 3DGS and neural SDF semantic-level hybrid reconstruction method is adopted. Semantic maps, depth maps and normal maps are generated through image preprocessing. 3DGS and neural SDF are combined for spatial alignment, and the Marching Cubes algorithm is used for 3D object modeling. Semantic-level region densification and hierarchical depth map restoration are introduced, and the MLP network is optimized to improve the reconstruction quality.

Benefits of technology

It significantly improves the structural coherence and detail expression capabilities of the scene, enhances the sensitivity and accuracy of key structures, and effectively restores the geometric structure of occluded objects. The generated three-dimensional scene data is suitable for downstream applications such as semantic editing, virtual-reality fusion, and interactive operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655823A_ABST
    Figure CN120655823A_ABST
Patent Text Reader

Abstract

The invention discloses an occlusion perception 3DGS and neural SDF semantic level hybrid reconstruction method, which comprises the following steps: image preprocessing: initially obtaining an image video frame through a camera, starting from image preprocessing, processing an input image, including semantic segmentation, depth estimation, normal estimation and internal and external parameter calibration, a corresponding semantic graph, a depth graph, a normal graph and camera parameters are generated; aligning the 3DGS with the neural SDF, after preprocessing is completed, performing spatial alignment on geometric information generated based on the 3DGS and a neural symbol distance function, and performing restoration in combination with regional densification and a hierarchical depth map; and finishing the final three-dimensional object modeling and rapid reconstruction by adopting a Marking Cubes algorithm. According to the occlusion perception 3DGS and neural SDF semantic level hybrid reconstruction method provided by the invention, each object in a scene can be quickly reconstructed with high quality, and applications such as downstream editing tasks, interaction and the like are supported at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of visual processing technology, and in particular to an occlusion-aware 3DGS and neural SDF semantic-level hybrid reconstruction method. Background Art

[0002] NeRF is a 3D scene representation method based on neural networks. It renders high-quality images from any perspective by learning a continuous 3D radiation field. The core idea of ​​NeRF is to use an MLP to encode the radiation field in 3D space. The input is the 3D coordinate position (x, y, z) and the viewing direction (θ, φ), and the output is the color c and volume density σ at that position. The color value of each pixel is calculated by integrating along each ray. NeuS and VolSDF are two 3D surface reconstruction technologies based on NeRF. They introduce SDF as an implicit expression of the three-dimensional surface and propose a volume rendering method based on SDF, but each adopts different strategies to optimize the reconstruction quality and efficiency.

[0003] 3DGS represents a scene as a series of Gaussian spheres with color and opacity attributes. Each Gaussian sphere is described by position, color, opacity, and a covariance matrix. By combining these Gaussian spheres, complex three-dimensional scenes can be rendered efficiently. The 3DGS rendering process uses projected 2D Gaussian spheres to calculate the color of each pixel.

[0004] However, although the current reconstruction method based on 3DGS has advantages in rendering efficiency, it has problems such as uneven point cloud distribution, blurred boundaries, and broken geometric structures in complex scenes. In particular, it is difficult to accurately restore the edges of objects and fine structure areas. In areas with weak texture or unclear color changes, existing methods often find it difficult to provide sufficient depth clues, resulting in low reconstruction density and blurred shapes in the corresponding areas. In addition, traditional three-dimensional reconstruction methods have difficulty dealing with the problem of mutual occlusion between multiple objects. Summary of the Invention

[0005] The purpose of the present invention is to provide an occlusion-aware 3DGS and neural SDF semantic-level hybrid reconstruction method that can quickly and high-quality reconstruct each object in the scene while supporting downstream editing tasks, interactive applications, etc.

[0006] The technical solution adopted by the occlusion-aware 3DGS and neural SDF semantic-level hybrid reconstruction method disclosed in the present invention is:

[0007] An occlusion-aware 3DGS and neural SDF semantic-level hybrid reconstruction method includes the following steps:

[0008] Image preprocessing: Initially, the camera acquires the image video frame. Starting with image preprocessing, the input image processing steps include semantic segmentation, depth estimation, normal estimation, and internal and external parameter calibration, thereby generating the corresponding semantic map, depth map, normal map, and camera parameters.

[0009] 3DGS is aligned with the neural SDF. After preprocessing, the geometric information generated based on 3DGS is spatially aligned with the neural signed distance function, and combined with regional densification and hierarchical depth map restoration;

[0010] Finally, the Marching Cubes algorithm is used to convert the semantic SDF into a triangular mesh model, or the depth map is converted into SDF and then Marching Cubes is used to complete the final three-dimensional object modeling and rapid reconstruction.

[0011] As a preferred solution, the image preprocessing steps are as follows:

[0012] Obtain the corresponding instance semantic segmentation mask S={S1,S2,...,S N};

[0013] The corresponding depth map D = {D1, D2, ..., D N};

[0014] The corresponding normal map N={N1,N2,...,N N};

[0015] Colmap calculates the camera external parameters to obtain the pose P = {P1, P2, ..., P N};

[0016] The intrinsic parameter of the camera calibration is K, where N is the number of images.

[0017] As a preferred solution, the steps of aligning the 3DGS with the neural SDF are as follows:

[0018] Initial geometry construction of 3DGS: In the initial stage, 3DGS can quickly approximate the geometric structure of the scene by optimizing parameters such as the Gaussian center position, scale, and transparency;

[0019] 3DGS guides MLP sampling: The depth value of the Gaussian point cloud is used as a guiding signal to guide MLP sampling, ensuring that the sampling points can cover the key geometric areas. If the density of the Gaussian point cloud is high enough, the Gaussian points can be directly used as the input of the MLP without additional sampling.

[0020] Neural SDF field construction: After sampling is completed, the MLP network learns the SDF value and semantic label of the scene based on the information of the sampling points;

[0021] Rasterization rendering normal map: After the SDF field is constructed, the normal map can be quickly rendered using 3DGS rasterization technology to align it more accurately. Since the SDF itself is a continuous field, the normal vector can be calculated by the gradient of the SDF, as shown in the following formula:

[0022]

[0023] in, Represents the gradient of the neural SDF function;

[0024] Object-level mesh extraction: This can be done in two ways: MLP predicts SDF and semantic labels, or semantic-level 3DGS renders the depth map into SDF and then uses Marching cubes to obtain the mesh.

[0025] As a preferred solution, the MLP network is as follows:

[0026] Input layer: The input three-dimensional coordinates are first mapped to a high-dimensional space through the position encoding of the formula:

[0027] γ(x)=[sin(2 0 πx),cos(2 0 πx),...,sin(2 L-1 πx),cos(2 L-1 πx)]

[0028] L is the frequency level of the positional encoding;

[0029] Hidden layer: The hidden layer uses a variable-depth MLP structure, with 4, 6, or 8 layers set according to the complexity of the scene and objects. Each layer contains 64, 128, or 256 neurons. All hidden layers use ReLU as the activation function;

[0030] Output layer: The output layer contains two output nodes, corresponding to the SDF value and semantic label of the point respectively;

[0031] The stable Gaussian points are locked to avoid meaningless adjustments or splits in subsequent training, while the training focus is shifted to the neural SDF to further optimize the implicit surface shape.

[0032] As a preferred solution, the regional densification: using semantic segmentation information, giving higher weights to key areas, and accumulating the gradients of all Gaussian points in the same semantic area and taking their absolute values ​​to reduce the impact of gradient offset. In addition, the key semantic areas are weighted in combination with semantic weights to ensure that the key areas can obtain higher point cloud density while avoiding generating too many point clouds in non-key areas.

[0033]

[0034] Among them, L represents the loss function, x i represents the spatial coordinates of the i-th Gaussian point, S k represents the kth semantic region, w k represents the semantic weight of the region, w k The setting can be adjusted according to the semantic category of the scene, giving higher weights to weak texture areas and lower weights to dense areas such as foreground.

[0035] As a preferred solution, the steps of repairing the layered depth map are as follows:

[0036] Determine the occlusion relationship between objects: For each semantic block, first find its nearest neighbor and combine the depth information of each pixel to compare the relative depths between objects. Suppose the depth range of object A is If the depth of object A is the minimum Greater than the maximum depth of object B Right now It is believed that object B may be located behind object A and there is an occlusion area;

[0037] Removing occlusions between different layers: Advanced models such as Lama are used to remove the effects of occlusions between different layers, allowing each layer to be clearly presented. At the same time, models such as Metric3D and SAM are applied to generate depth maps and semantic maps for each layer.

[0038] Depth map alignment and fusion: To effectively fill in missing parts in the depth map, depth information is first supplemented using depth maps from multiple perspectives. Based on the camera's intrinsic and extrinsic parameters, the depth maps from different perspectives are geometrically aligned. The aligned depth values ​​are then fused with depth maps from different layers.

[0039]

[0040] Among them, K ref represents the internal parameter matrix under the reference perspective, T ref represents the external parameter matrix under the reference perspective, K i represents the internal parameter matrix under the i-th perspective, T ref Denotes the external parameter matrix under the i-th perspective, D i ′ represents the depth map of the i-th level, w i represents the confidence of the i-th level, and l represents the occlusion level;

[0041] Post-processing: The fused depth map contains noise, and image filtering method is used for smoothing to further remove the noise.

[0042] The beneficial effects of the occlusion-aware 3DGS and neural SDF semantic-level hybrid reconstruction method disclosed in the present invention are as follows: by aligning and fusing the 3DGS point cloud representation with the neural implicit surface modeling, it combines the rendering efficiency of explicit representation with the geometric accuracy of implicit representation, significantly improving the structural coherence and detail expression ability of the overall scene, introducing a semantic-level regional densification mechanism, and guiding it through semantic segmentation, enhancing the Gaussian point density and gradient response in areas of high semantic value, significantly improving the model's sensitivity and accuracy to key structures, and through an occlusion-aware multi-level deep restoration mechanism, combining advanced generative models at the image and depth levels to complete the occluded areas with semantic consistency, effectively restoring the geometric structure of the occluded objects. The generated semantically enhanced three-dimensional scene data not only has high geometric and semantic accuracy, but is also naturally adapted to downstream application scenarios such as semantic editing, virtual-reality fusion, and interactive operations. Compared with traditional reconstruction solutions. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 This is a flowchart of an occlusion-aware 3DGS and neural SDF semantic-level hybrid reconstruction method of the present invention. DETAILED DESCRIPTION

[0044] The present invention will be further described and explained below in conjunction with specific embodiments and accompanying drawings:

[0045] Please refer to Figure 1 , an occlusion-aware 3DGS and neural SDF semantic-level hybrid reconstruction method, including the following steps:

[0046] Image preprocessing: Initially, image and video frames are obtained through cameras such as monocular, binocular, and RGBD. Starting from image preprocessing, the input image processing steps include semantic segmentation, depth estimation, normal estimation, and internal and external parameter calibration, thereby generating the corresponding semantic map, depth map, normal map, and camera parameters;

[0047] 3DGS is aligned with the neural SDF. After preprocessing, the geometric information generated based on 3DGS is spatially aligned with the neural signed distance function, and combined with regional densification and hierarchical depth map restoration;

[0048] Finally, the Marching Cubes algorithm is used to convert the semantic SDF into a triangular mesh model, or the depth map is converted into SDF and then Marching Cubes is used to complete the final three-dimensional object modeling and rapid reconstruction.

[0049] The image preprocessing steps are as follows:

[0050] Obtain the corresponding instance semantic segmentation mask S={S1,S2,...,S N};

[0051] The corresponding depth map D = {D1, D2, ..., D N};

[0052] The corresponding normal map N={N1,N2,...,N N};

[0053] Colmap calculates the camera external parameters to obtain the pose P = {P1, P2, ..., P N};

[0054] The intrinsic parameter of the camera calibration is K, where N is the number of images.

[0055] The steps for aligning the 3DGS with the neural SDF are as follows:

[0056] Initial geometry construction of 3DGS: In the initial stage, 3DGS can quickly approximate the geometric structure of the scene by optimizing parameters such as the Gaussian center position, scale, and transparency;

[0057] 3DGS guides MLP sampling: The depth value of the Gaussian point cloud is used as a guiding signal to guide MLP sampling, ensuring that the sampling points can cover the key geometric areas. If the density of the Gaussian point cloud is high enough, the Gaussian points can be directly used as the input of the MLP without additional sampling.

[0058] Neural SDF field construction: After sampling is completed, the MLP network learns the SDF value and semantic label of the scene based on the information of the sampling points;

[0059] Rasterization rendering normal map: After completing the construction of the SDF field, the normal map can be quickly rendered through 3DGS rasterization technology to make it more accurately aligned. Since the SDF itself is a continuous field, the normal vector can be obtained by calculating the gradient of the SDF, as shown in the following formula:

[0060]

[0061] in, Represents the gradient of the neural SDF function;

[0062] Object-level mesh extraction: This can be done in two ways: MLP predicts SDF and semantic labels, or semantic-level 3DGS renders the depth map into SDF and then uses Marching cubes to obtain the mesh.

[0063] The MLP network is as follows:

[0064] Input layer: The input three-dimensional coordinates are first mapped to a high-dimensional space through the position encoding of the formula:

[0065] γ(x)=[sin(20 πx),cos(2 0 πx),...,sin(2 L-1 πx),cos(2 L-1 πx)]

[0066] L is the frequency level of the positional encoding;

[0067] Hidden layer: The hidden layer uses a variable-depth MLP structure, with 4, 6, or 8 layers set according to the complexity of the scene and objects. Each layer contains 64, 128, or 256 neurons. All hidden layers use ReLU as the activation function;

[0068] Output layer: The output layer contains two output nodes, corresponding to the SDF value and semantic label of the point respectively;

[0069] The stable Gaussian points are locked to avoid meaningless adjustments or splits in subsequent training, while the training focus is shifted to the neural SDF to further optimize the implicit surface shape.

[0070] Regional densification: Using semantic segmentation information, higher weights are assigned to key areas (such as object edges and texture-rich areas). The gradients of all Gaussian points in the same semantic area are accumulated and their absolute values ​​are taken to reduce the effect of gradient offset. In addition, the key semantic areas are weighted based on semantic weights to ensure that the key areas can obtain higher point cloud density while avoiding generating too many point clouds in non-key areas.

[0071]

[0072] Among them, L represents the loss function, x i represents the spatial coordinates of the i-th Gaussian point, S k represents the kth semantic region, w k represents the semantic weight of the region, w k The setting can be adjusted according to the semantic category of the scene, giving higher weights to weak texture areas and lower weights to dense areas such as foreground.

[0073] Occlusion is a common problem in the process of generating meshes from depth maps. Especially in indoor scenes, when multiple objects occlude each other, a single-view depth map often cannot fully capture the geometric information of the objects.

[0074] The steps for repairing the layered depth map are as follows:

[0075] Determine the occlusion relationship between objects: For each semantic block, first find its nearest neighbor and combine the depth information of each pixel to compare the relative depths between objects. Suppose the depth range of object A is If the depth of object A is the minimum Greater than the maximum depth of object B Right now It is believed that object B may be located behind object A and there is an occlusion area;

[0076] Removing occlusions between different layers: Advanced models such as Lama are used to remove the effects of occlusions between different layers, allowing each layer to be clearly presented. At the same time, models such as Metric3D and SAM are applied to generate depth maps and semantic maps for each layer.

[0077] Depth map alignment and fusion: To effectively fill in missing parts in the depth map, depth information is first supplemented using depth maps from multiple perspectives. Based on the camera's intrinsic and extrinsic parameters, the depth maps from different perspectives are geometrically aligned. The aligned depth values ​​are then fused with depth maps from different layers.

[0078]

[0079] Among them, K ref represents the internal parameter matrix under the reference perspective, T ref represents the external parameter matrix under the reference perspective, K i represents the internal parameter matrix under the i-th perspective, T ref Denotes the external parameter matrix under the i-th perspective, D i ′ represents the depth map of the i-th level, w i represents the confidence of the i-th level, and l represents the occlusion level;

[0080] Post-processing: The fused depth map contains noise, and image filtering method is used for smoothing to further remove the noise.

[0081] (1) Achieve high-integrity and high-precision 3D scene reconstruction:

[0082] By aligning and fusing 3DGS point cloud representation with Neural Implicit Surface Modeling (Neural SDF), this paper combines the rendering efficiency of explicit representation with the geometric accuracy of implicit representation, significantly improving the structural coherence and detail expression of the overall scene. This is particularly effective at object boundaries and in complex structural areas, effectively avoiding geometric fragmentation and blurring.

[0083] (2) Improve the modeling quality and expression capabilities of key semantic areas:

[0084] This paper introduces a semantic-level region densification mechanism. Guided by semantic segmentation, it enhances Gaussian point density and gradient response in areas of high semantic value (such as object edges and texture fuzzy areas), significantly improving the model's sensitivity and accuracy to key structures. At the same time, it automatically suppresses point density in non-critical areas, improving modeling efficiency and avoiding resource waste.

[0085] (3) Realize intelligent reasoning and reconstruction of occluded areas:

[0086] This paper uses an occlusion-aware multi-level depth restoration mechanism and combines advanced generative models (LaMa+Stable Diffusion) to complete semantically consistent occluded areas at the image and depth levels, effectively restoring the geometric structure of the occluded objects and achieving more complete and realistic object-level reconstruction results.

[0087] (4) Stronger downstream scalability and interactive adaptability:

[0088] The semantically enhanced 3D scene data generated by this method not only possesses a high degree of geometric and semantic accuracy, but is also naturally suited for downstream applications such as semantic editing, virtual-reality fusion, and interactive operations. Compared to traditional reconstruction solutions, this method demonstrates greater practical value and expansion potential in terms of real-time rendering performance, editing convenience, and semantic controllability.

[0089] These improvements have enabled the technology to demonstrate remarkable results in a variety of practical applications, including virtual reality, augmented reality, architectural design, cultural heritage preservation, and autonomous driving. For example, in the AR and VR fields, by providing higher-quality, more fluid virtual environments, it attracts more users and increases platform subscription fees and advertising revenue. While the direct economic benefits of cultural heritage preservation are limited, digital displays can attract more tourists and indirectly increase tourism revenue.

[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the scope of protection of the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the essence and scope of the technical solutions of the present invention.

Claims

1. An occlusion-aware 3DGS and neural SDF semantic-level hybrid reconstruction method, characterized by: The following steps are involved: Image preprocessing: Initially, the camera acquires the image video frame. Starting with image preprocessing, the input image processing steps include semantic segmentation, depth estimation, normal estimation, and internal and external parameter calibration, thereby generating the corresponding semantic map, depth map, normal map, and camera parameters. 3DGS is aligned with the neural SDF. After preprocessing, the geometric information generated based on 3DGS is spatially aligned with the neural signed distance function, and combined with regional densification and hierarchical depth map restoration; Finally, the Marching Cubes algorithm is used to convert the semantic SDF into a triangular mesh model, or the depth map is converted into SDF and then Marching Cubes is used to complete the final three-dimensional object modeling and rapid reconstruction.

2. The occlusion-aware 3DGS and neural SDF semantic-level hybrid reconstruction method according to claim 1, characterized in that: The image preprocessing steps are as follows: Obtain the corresponding instance semantic segmentation mask S={S1,S2,...,S N }; The corresponding depth map D = {D1, D2, ..., D N }; The corresponding normal map N={N1,N2,...,N N }; Colmap calculates the camera external parameters to obtain the pose P = {P1, P2, ..., P N }; The intrinsic parameter of the camera calibration is K, where N is the number of images.

3. The occlusion-aware 3DGS and neural SDF semantic-level hybrid reconstruction method according to claim 1, characterized in that: The steps for aligning the 3DGS with the neural SDF are as follows: Initial geometry construction of 3DGS: In the initial stage, 3DGS can quickly approximate the geometric structure of the scene by optimizing parameters such as the Gaussian center position, scale, and transparency; 3DGS guides MLP sampling: The depth value of the Gaussian point cloud is used as a guiding signal to guide MLP sampling, ensuring that the sampling points can cover the key geometric areas. If the density of the Gaussian point cloud is high enough, the Gaussian points can be directly used as the input of the MLP without additional sampling. Neural SDF field construction: After sampling is completed, the MLP network learns the SDF value and semantic label of the scene based on the information of the sampling points; Rasterization rendering normal map: After the SDF field is constructed, the normal map can be quickly rendered using 3DGS rasterization technology to align it more accurately. Since the SDF itself is a continuous field, the normal vector can be calculated by the gradient of the SDF, as shown in the following formula: in, Represents the gradient of the neural SDF function; Object-level mesh extraction: This can be done in two ways: MLP predicts SDF and semantic labels, or semantic-level 3DGS renders the depth map into SDF and then uses Marching cubes to obtain the mesh.

4. The occlusion-aware 3DGS and neural SDF semantic-level hybrid reconstruction method according to claim 3, characterized in that: The MLP network is as follows: Input layer: The input three-dimensional coordinates are first mapped to a high-dimensional space through the position encoding of the formula: γ(x)=[sin(2 0 πx),cos(2 0 πx),...,sin(2 L-1 πx),cos(2 L-1 πx)] L is the frequency level of the positional encoding; Hidden layer: The hidden layer uses a variable-depth MLP structure, with 4, 6, or 8 layers set according to the complexity of the scene and objects. Each layer contains 64, 128, or 256 neurons. All hidden layers use ReLU as the activation function; Output layer: The output layer contains two output nodes, corresponding to the SDF value and semantic label of the point respectively; The stable Gaussian points are locked to avoid meaningless adjustments or splits in subsequent training, while the training focus is shifted to the neural SDF to further optimize the implicit surface shape.

5. The occlusion-aware 3DGS and neural SDF semantic-level hybrid reconstruction method according to claim 1, characterized in that: The regional densification uses semantic segmentation information to assign higher weights to key areas, and accumulates the gradients of all Gaussian points in the same semantic area and takes their absolute values ​​to reduce the impact of gradient offset. In addition, the key semantic areas are weighted in combination with semantic weights to ensure that the key areas can obtain higher point cloud density while avoiding generating too many point clouds in non-key areas. Among them, L represents the loss function, x i represents the spatial coordinates of the i-th Gaussian point, S k represents the kth semantic region, w k represents the semantic weight of the region, w k The setting can be adjusted according to the semantic category of the scene, giving higher weights to weak texture areas and lower weights to dense areas such as foreground.

6. The occlusion-aware 3DGS and neural SDF semantic-level hybrid reconstruction method according to claim 1, characterized in that: The steps of repairing the hierarchical depth map are as follows: Determine the occlusion relationship between objects: For each semantic block, first find its nearest neighbor and combine the depth information of each pixel to compare the relative depths between objects. Suppose the depth range of object A is If the depth of object A is the minimum Greater than the maximum depth of object B Right now It is believed that object B may be located behind object A and there is an occlusion area; Removing occlusions between different layers: Advanced models such as Lama are used to remove the effects of occlusions between different layers, allowing each layer to be clearly presented. At the same time, models such as Metric3D and SAM are applied to generate depth maps and semantic maps for each layer. Depth map alignment and fusion: To effectively fill in missing parts in the depth map, depth information is first supplemented using depth maps from multiple perspectives. Based on the camera's intrinsic and extrinsic parameters, the depth maps from different perspectives are geometrically aligned. The aligned depth values ​​are then fused with depth maps from different layers. Among them, K ref represents the internal parameter matrix under the reference perspective, T ref represents the external parameter matrix under the reference perspective, K i represents the internal parameter matrix under the i-th perspective, T ref Denotes the external parameter matrix under the i-th perspective, D i ′ represents the depth map of the i-th level, w i represents the confidence of the i-th level, and l represents the occlusion level; Post-processing: The fused depth map contains noise, and image filtering method is used for smoothing to further remove the noise.

Citation Information

Cited By

  • Repair and 3D reconstruction method and system based on polluted face picture

    CN120852248A

  • Block reconstruction method and device of 3DGS, electronic equipment and storage medium

    CN121725162A

  • A patch reconstruction method and device for 3D GS, electronic equipment and storage medium

    CN121725162B

  • Monocular visual relocalization method and system based on WiFi fingerprint assistance and three-dimensional Gaussian scene representation

    CN122714563A