Two-stage surface reconstruction method, device, and storage medium based on SDF volume rendering

By using two-stage training method and adaptive scale factors in SDF body rendering, the problem that SDF-based body rendering is difficult to capture complex geometric details in scenes with high uncertainty is solved, and high-fidelity and fine surface reconstruction are achieved.

CN119107396BActive Publication Date: 2025-06-24SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411017572.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-29
Publication Date
2025-06-24
Estimated Expiration
2044-07-29

AI Technical Summary

Technical Problem

Existing SDF-based bulk rendering methods are difficult to capture complex geometric details in scenarios with high uncertainty, resulting in visible wrong surfaces in surface reconstruction.

Method used

Using a two-stage surface reconstruction method based on SDF, the pre-trained geometric network and color network are used to estimate the stochastic step numerical gradient and the true analysis gradient, combined with Eikonal loss, color loss and deviation correction loss, training under the adaptive scale factor is carried out.

Benefits of technology

Reduces false surfaces in surface reconstruction, improves flexibility and effectiveness of SDF representation, optimizes geometric representation deviations, and achieves high fidelity and fine surface reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119107396B_ABST
    Figure CN119107396B_ABST
Patent Text Reader

Abstract

The present invention relates to a two-stage surface reconstruction method, device, and storage medium based on SDF volume rendering. First, the conversion from SDF to volume density is improved by discarding the global scale factor and using a local adaptive scale factor. Second, a novel loss function is implemented, aiming to align the maximum probability distance in volume rendering with the zero level set, thereby improving the deviation problem of geometric representation. In addition, the above improvements are incorporated into a two-stage optimization framework to address the over-regularization imposed by geometric constraints. In the rough optimization stage, the operation of the SDF field is similar to that of the volume density field and is minimally affected by topological transformations. In the refinement stage, a surface with increased smoothness is implemented. The method also provides a stochastic step numerical gradient estimation technique to maintain the natural zero level set in the rough stage. Through the above design, the method can achieve high-fidelity surface reconstruction suitable for large-scale and intricate geometries.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of volume rendering, and in particular, to a two-stage surface reconstruction method, device, and storage medium based on SDF volume rendering. Background Art

[0002] The importance of volume rendering technology based on the Signed Distance Function (SDF) for surface reconstruction is self-evident. However, in practical applications, especially for scenes with large uncertainties, the SDF-based volume rendering method often fails to capture fine geometric structures, resulting in visible surface defects.

[0003] The signed distance function provides a new method for 3D surface reconstruction. However, integrating SDF into the volume density function poses great challenges. For scenes with high uncertainties, complex geometric details are often not captured or visible false surfaces are created. Although volume rendering based on volume density can reconstruct the surface with precise positioning, the surface has a certain roughness. Although SDF-based volume rendering can generate smoother surfaces, it has problems in correctly positioning parts of the surface. This difference highlights the limitations of SDF-based surface reconstruction methods and also emphasizes the need to improve the modeling ability.

[0004] In summary, there is currently a lack of a surface reconstruction method to further improve the performance and reliability of neural network surface reconstruction. Summary of the Invention

[0005] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide a two-stage surface reconstruction method, device, and storage medium based on SDF volume rendering to solve or partially solve the problem of visible false surfaces in surface reconstruction.

[0006] The purpose of the present invention can be achieved by the following technical solutions:

[0007] In one aspect of the present invention, a two-stage surface reconstruction method based on SDF volume rendering is provided. Based on the acquired 2D image, using a pre-trained geometry network and color network, surface reconstruction is achieved through volume rendering. The training process of the geometry network and color network includes the following steps:

[0008] Step S1, based on the acquired sample images and corresponding camera parameters, calculate the ray information along the view direction of the camera position. Take the point coordinates on the ray as the inputs of the geometry network and color network respectively, and obtain the predicted distance sign function information and color information respectively.

[0009] Step S2, on the premise of taking uncertainty into account, obtain the estimated gradient of the distance signed function information through random step size numerical gradient estimation. Combine the Eikonal loss of the estimated gradient, the color loss calculated based on the color information, and the deviation correction loss, and achieve the training of the first stage under the premise of an adaptive scale factor and a low lower limit of the scale factor;

[0010] Step S3, calculate the true analytical gradient of the distance signed function information. Combine the Eikonal loss of the true analytical gradient, the color loss calculated based on the color information, and the smoothness constraint loss, and achieve the training of the second stage under the premise of an adaptive scale factor and a high lower limit of the scale factor.

[0011] As a preferred technical solution, step S1 is implemented by the following formula:

[0012]

[0013]

[0014] where, and are the geometric network and the color network respectively, { is the camera position emits a ray along the view direction , is the point on the ray, and and are the predicted distance signed function information, scale factor, and geometric feature respectively, is the preset conversion function from the distance signed function to the volume density, is the volume density converted from the distance signed function, is the predicted color information.

[0015] As a preferred technical solution, the conversion function is the rendering distance function.

[0016] As a preferred technical solution, in step S2, the loss function is:

[0017]

[0018] where, is the loss function for the training of the first stage, is the color loss, is the Eikonal loss of the estimated gradient of the distance signed function information, is the deviation correction loss, and are the loss weights.

[0019] As a preferred technical solution, in the step S2, the estimated gradient of the distance signed function information component is:

[0020]

[0021] where is the point coordinate on the ray in the view direction emitted from the camera position, is the predicted distance signed function information, and , is a preset value.

[0022] As a preferred technical solution, the deviation correction loss is:

[0023]

[0024] where is the point coordinate on the ray in the view direction emitted from the camera position, is the predicted distance signed function information, is the deviation correction factor, , is the volume density transformed from the distance signed function, is the ray set in each training batch.

[0025] As a preferred technical solution, in the step S3, the loss function is:

[0026]

[0027] where is the loss function for the second-stage training, is the color loss, is the true analytical gradient of the signed function information, is the smoothness constraint loss, and are the loss weights.

[0028] As a preferred technical solution, the smoothness constraint loss is:

[0029]

[0030] where is the ray in the view direction emitted from the camera position, is the point normal vector at, For the distance signed function information the true analytical gradient of , is a random unit vector is the smooth coefficient, and is the number of rays within one batch and is the number of sampled points for each ray.

[0031] Another aspect of the present invention provides an electronic device, including: one or more processors and a memory, where the memory stores one or more programs, and the one or more programs include instructions for executing the foregoing two-stage surface reconstruction method based on SDF volume rendering.

[0032] Another aspect of the present invention provides a computer-readable storage medium, including one or more programs for execution by one or more processors of an electronic device, and the one or more programs include instructions for executing the foregoing two-stage surface reconstruction method based on SDF volume rendering.

[0033] Compared with the prior art, the present invention has at least one of the following beneficial effects:

[0034] (1) Reducing the false surfaces generated by surface reconstruction: Aiming at the problem that the deviation of volume rendering and geometric over-regularization often lead to visible false surfaces in surface reconstruction, the present invention trains the geometric network and the color network in a coarse-fine two-stage training manner. In the first stage, the estimated gradient of the distance signed function information is obtained by using random step numerical gradient estimation, and in the second stage, the loss is calculated using the true analytical gradient. By introducing random step numerical gradient estimation, the natural zero level set of the rough stage is maintained while avoiding geometric over-regularization.

[0035] (2) High flexibility: Aiming at the current problem of forcibly characterizing the volume density influence ability for points with the same SDF value, the present invention improves the conversion from the distance signed function (SDF) to the volume density by transitioning from a global scale factor to a local adaptive scale factor, allowing adaptive volume density values, thereby improving the flexibility and effectiveness of the SDF representation.

[0036] (3) Optimizing the deviation of geometric representation: The present invention provides a method for calculating the deviation correction loss, aiming to align the maximum probability distance in volume rendering with the zero level set, thereby improving the deviation problem of geometric representation. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 is a flowchart of the training process of the geometric network and the color network in the embodiment;

[0038] Figure 2Comparison analysis diagram of the SDF-based and density-based volume rendering methods in the embodiment;

[0039] Figure 3 Visualization schematic diagram of the volume density deviation in the embodiment;

[0040] Figure 4 Heat map of the normal variance predicted using a random step size in the embodiment;

[0041] Figure 5 Comparison of visual results on the Tanks and Temples dataset in the embodiment;

[0042] Figure 6 Comparison of visual results on the Advance subset of the Tanks and Temples dataset in the embodiment;

[0043] Figure 7 Comparison of visual results on the ScanNet++ dataset in the embodiment

[0044] Figure 8 Schematic diagram for comparing the visual effects and metrics of the ablation experiment in the embodiment;

[0045] Figure 9 Analysis schematic diagram of the numerical gradient estimation of the random step size in the embodiment;

[0046] Figure 10 Analysis schematic diagram of the explicit bias correction in the embodiment. Detailed implementation manners

[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0048] The following are the relevant definitions involved in this embodiment:

[0049] A neural radiance field (NeRF) is a method for recovering a 3D scene from a set of 2D images. Its input is 2D images and the camera parameters of the corresponding images. By sampling the 2D image pixels and projecting rays, 3D coordinate points are sampled in the corresponding rays as the input of a neural network. These inputs are mapped to the corresponding colors and volume densities. Finally, through volume rendering technology, the 3D attributes in Euclidean space are projected onto the 2D imaging plane for rendering loss calculation with the input 2D images.

[0050] NeRF introduces a neural network to represent the neural radiance field for novel view synthesis. The core idea of this method is to learn a continuous three-dimensional scene representation through a neural network, which can describe the color and transparency at any given position and view angle in a scene. In this way, images of any new view can be generated through the neural network, thus realizing novel view synthesis. To optimize these scenes, NeRF adopts the method of differentiable volume rendering. Specifically, it calculates the color of each pixel by integrating along the ray path of each pixel, and then optimizes the parameters of the neural network through backpropagation. In this way, the global information of the entire scene can be taken into account during the optimization process, thus generating accurate and realistic novel view images.

[0051] The signed distance function (SDF) is defined as the orthogonal distance from a point to the boundary of a certain set, and its sign depends on whether the point is inside the boundary. This function takes positive values at points inside the boundary. As the point approaches the boundary, its value decreases, and when it reaches the boundary, the value of the signed distance function is zero; outside the boundary, the function value is negative.

[0052] Example 1

[0053] Existing methods still face challenges in dealing with complex geometric details and solving noise problems. Incorporating SDF into the volume density function is not easy. In practical applications, especially for scenes with large uncertainties, the volume rendering technique based on the directed distance function often fails to capture fine geometric structures, resulting in visible surface defects. The specific reasons include certain deviations in the conversion from SDF to volume density, and implicitly limiting the expressive power of the volume density. Excessive geometric regularization of SDF is also one of the reasons.

[0054] This is illustrated by comparing two methods for reconstructing the same scene: Instant-NGP (Müller et al. 2022) uses volume rendering based on volume density, while Neuralangelo (Li et al. 2023) uses volume rendering based on SDF. Both methods use a similar multi-resolution hash table representation. The mesh is generated for Instant-NGP through TSDF-fusion (Barian et al. 1996). As Figure 2As shown, where (a) is Neuralangelo, (b) is Instant-NGP, (c) is the method of this embodiment, and (d) is the GT grid. Although Instant-NGP can reconstruct the surface with precise positioning, the surface has a certain roughness. While Neuralangelo can generate a smoother surface, it has problems in correctly positioning parts of the surface. This difference highlights the limitations of the SDF-based surface reconstruction method and also emphasizes the need to improve the modeling ability.

[0055] To address the above problems, this embodiment provides a two-stage surface reconstruction method based on SDF volume rendering to further improve the performance and reliability of neural network surface reconstruction. First, the conversion from SDF to volume density is improved by discarding the global scale factor and using a local adaptive scale factor. Different from previous methods, this method allows adaptive volume density values, thus improving the flexibility and effectiveness of the SDF representation, rather than forcing the same volume density for points with the same SDF value. Second, a new loss function is implemented, aiming to align the maximum probability distance in volume rendering with the zero level set, thereby improving the deviation problem of geometric representation. In addition, the above innovations are incorporated into a two-stage optimization framework to address the over-regularization imposed by geometric constraints. Initially, a rough optimization stage is adopted, in which the SDF field is operated similarly to the volume density field, with the least influence from topological transformation. Subsequently, a refinement stage is carried out to achieve a surface with increased smoothness. This method also provides a random step numerical gradient estimation technique to maintain the natural zero level set in the rough stage. Through the above design, this method can achieve high-fidelity surface reconstruction suitable for large-scale and intricate geometries.

[0056] Specifically, based on the acquired 2D images, this method uses a pre-trained geometry network and color network to achieve surface reconstruction through volume rendering. Among them, see Figure 1 , the training process of the geometry network and the color network includes the following steps:

[0057] Step S1: Based on the acquired sample images and corresponding camera parameters, calculate the ray information of the camera position along the view direction, and use the point coordinates on the rays as the inputs of the geometry network and the color network respectively to obtain the predicted distance sign function information and color information.

[0058] Step S2: On the premise of considering uncertainty, obtain the estimated gradient of the distance sign function information through random step numerical gradient estimation. Combine the Eikonal loss of the estimated gradient, the color loss calculated based on the color information, and the deviation correction loss, and achieve the training of the first stage on the premise of an adaptive scale factor and a low lower limit of the scale factor.

[0059] Step S3: Calculate the true analytical gradient of the distance signed function information. Combine the Eikonal loss of the true analytical gradient, the color loss calculated based on color information, and the smoothness constraint loss to achieve the training of the second stage under the premise of an adaptive scale factor and a high lower limit of the scale factor.

[0060] When performing surface reconstruction using the geometry network and the color network, from the geometry network the predicted SDF and the geometric features at point . Then, through a predefined function and the global scale factor transform the SDF into volume density , thereby modeling the 3D scene as a volume density field to achieve surface reconstruction.

[0061] In this method, the training processes of the geometry network and the color network have the following characteristics:

[0062] (1) Different from the way of using a global scale factor, this method enables the same SDF level set to exhibit different volume densities. This allows the predicted density to reach any non - negative value, thus enhancing its expressive ability.

[0063] (2) Introduce stochastic step - size numerical gradient estimation to solve the over - regularization problem in the optimization process of the SDF - based method.

[0064] (3) Design an explicit bias correction aimed at alleviating bias and convergence problems.

[0065] (4) Design a two - stage optimization framework using a coarse - to - fine method, which allows optimizing rough surfaces while avoiding the above problems. Finally, through a refinement stage, it generates a surface with rich details and high fidelity.

[0066] The above characteristics will be elaborated separately in the following text.

[0067] (1) Improvement of the transformation from SDF to volume density.

[0068] Some existing SDF - based volume rendering methods combine volume rendering with the SDF representation. Different from density - based methods, geometric shapes are represented by the zero - level set of the SDF. From the geometry network the predicted SDF and the geometric features at point .

[0069] Then, through a predefined function and the global scale factor Convert the SDF into volume density . For example, VolSDF defines the density as the scaled cumulative distribution function of the negative SDF:

[0070]

[0071] In the above formula, is defined as the scale factor, and it can be seen that the global scale factor is used.

[0072] It should be noted that there are many conversion functions from SDF to volume density, which can be uniformly expressed as . In this embodiment, the above formula is mainly used as an example for illustration. Without conflict, other conversion functions can be used as alternatives.

[0073] Volume rendering methods based on SDF usually adopt a predefined function and the global scale factor to convert the SDF value into volume density, as described above. These methods often result in points with the same SDF value having uniform volume density values. This global scale factor mechanism limits the representation ability of the density field derived from the SDF field. Intuitively, the previously best-performing novel view synthesis methods can generate arbitrary non-negative volume density values within . On the contrary, combining the SDF representation with the global scale factor for surface reconstruction can only result in density values within .

[0074] To solve the problem of the representation limitation of converting SDF to density, a strategy similar to the adaptive scale factor is used instead of using the global scale factor for the conversion from SDF to density. This strategy uses a non-linear mapping to obtain a unique scale related to a given point, and different coordinates can obtain different scale factors through the network to achieve "adaptive". Given the camera position and the view direction , the ray emitted from along the direction is represented as . A set of points are sampled on this ray. The SDF of the point is obtained from the geometric network , geometric features , as well as the corresponding scale factor Note that the scale factor in this embodiment is obtained through a neural network with coordinate input, and different scale factors are used for different coordinates. Additionally, a color network is utilized to predict color , whose input is geometric features and the viewing direction . More precisely, its definition is as follows:

[0075] (5-1)

[0076]

[0077] where is a predefined conversion function from SDF to volume density, is the volume density converted from SDF.

[0078] The rendering color of the ray can be calculated as:

[0079]

[0080] With this design, the volume density is no longer the same within the same SDF level set. This method ensures that the volume density within the same SDF level set is no longer uniformly the same. Instead, they can vary, and by mapping the input coordinates to a continuous representation of their corresponding scale factors, any non-negative value can be achieved, enabling the volume density field derived from the SDF field to have a stronger representation ability. This design greatly enhances the flexibility and accuracy of our volume density modeling, making the reconstruction more realistic and detailed.

[0081] (2) Improvement for volume density deviation.

[0082] The deviation problem is a key issue that often needs to be solved in SDF-based volume rendering. As Figure 3 shown are the volume density deviations in (a) the ideal case and (b) the deviation case. Therefore, it is necessary to align the geometric representation in the volume rendering framework with the representation of the implicit surface. For the volume rendering framework, due to the ambiguity of the geometric representation of the volume density, that is, the geometric surface can be expressed as any volume density level set, and the most intuitive way to represent the geometric shape is through the rendering distance :

[0083] (5-2)

[0084] Preferably, the position where the rendering weight achieves its maximum value - that is, the position where the ray arrives and collides with the highest probability, or in other words, the position that contributes the most to the color - can be considered as the geometric representation within the volume rendering framework:

[0085] (5 - 3)

[0086] where is the maximum probability distance. For an implicit surface, the zero-level set provides a direct geometric representation. Ideally, whether convergence is achieved or not, the geometric representations of volume rendering (i.e., the rendering distance and the maximum probability distance) and the geometric representation of the implicit surface (i.e., the zero-level set) should be aligned, as shown in Figure 3 (a).

[0087] However, in the actual optimization process, conflicts as shown in Figure 3 (b) may occur, resulting in misalignment between the two representations. However, it should be noted that due to the deviation of the existing volume density, the rendering distance is significantly affected, especially in the early stage of convergence, as shown in Figure 3 (b).

[0088] Meanwhile, if geometric constraints are imposed on the SDF field in the case shown in Figure 3 (b), this deviation will be enhanced, resulting in the model being unable to converge correctly. Geometric constraints such as the Eikonal loss may prefer the situation of the SDF shown in Figure 3 (b), that is, the SDF at this time has minimized the Eikonal loss, but the weights at this time have also minimized the photometric loss , and the model has reached a local optimum at this time, and it is difficult for the local scale factor to converge further. If forced to converge, that is, reducing the local scale factor at this time, although it will align the maximum probability distance and the rendering distance with the zero-level set, it may result in an incorrect zero-level set. Therefore, the deviation problem is crucial for the SDF-based volume rendering method.

[0089] To address the aforementioned problems, this embodiment provides an explicit deviation correction method, in which the maximum probability distance is aligned with the zero-level set. Specifically, a deviation correction loss is defined as follows:

[0090] (5 - 4)

[0091] where is a deviation correction factor. The loss function is designed to penalize the positive part of , which encourages the SDF to take negative values after the maximum probability distance. This mechanism effectively aligns the maximum probability distance with the zero-level set by ensuring that the SDF value in the maximum probability distance region is below zero. In the experiment, it is approximated by directly using the sampling points with the largest number, although this introduces a certain degree of deviation, it does not affect the overall effectiveness. Preferably, to address the aforementioned implementation deviation, As a mask, only for the light rays greater than 0 are corrected for explicit deviation. This method can effectively prevent misalignment caused by estimating the maximum probability distance.

[0092] (3) Improvement for geometric over-regularization.

[0093] Some existing methods often produce incorrect surfaces due to over-regularizing the geometry, as shown in Fig. 2(a). However, the method based on volumetric density is not constrained by topological changes, as Figure 2 (b) shows. Based on the above results, this method considers first freely optimizing the SDF field and then refining it to a smooth surface through geometric regularization.

[0094] To this end, a novel two-stage optimization method is proposed. This method allows the optimization process to initially mimic the behavior based on volumetric density in the first stage and then refine it to a smooth surface in the second stage.

[0095] For the first stage, the goal is to solve the problem of over-regularization. One solution is to eliminate or reduce any geometric constraints and avoid conditioning the color on the predicted normal, but this method usually results in unnatural zero level sets.

[0096] To address the above problems, this embodiment provides a simple but effective method that can retain the natural level sets of large-scale structures while allowing the formation of complex structures without being hindered by geometric regularization. In the method, geometric regularization is selected to be applied to the estimated gradient , rather than directly applied to the gradient , and uncertainty is introduced through a specific design. Specifically, the component of the estimated gradient is , where and , and the calculation method of the gradient is implemented by finite difference. The difference is that the gradient estimation step size for each iteration is uniformly sampled from the range of.

[0097] Visualize the variance of the normal predicted by the random step size, as Figure 4 shown in (a) the reference image, (b) the variance of the normal. It can be observed that at a larger scale, the normal predicted using the random step size shows a smaller variance. However, for finer structures, such as Figure 4Details such as chair legs and table brackets shown have a large variance. The random step size gradient estimation provided in this embodiment can be regarded as introducing a certain degree of uncertainty in the application of geometric regularization. This means that the gradients of the estimated large-scale features are relatively stable and vary little in different sampling steps, while the gradients of complex structures show significant fluctuations in different sampling steps. This increased uncertainty relaxes the strict topological constraints on complex fine structures while still ensuring a natural and consistent zero-level set on larger-scale structures.

[0098] Specifically, some existing methods usually directly apply geometric constraints to the geometry of the gradient of the SDF corresponding to each sampling point and uniformly process all regions, that is, global geometric constraints. However, for regions with complex topological structures, global geometric constraints may encounter difficulties in optimization. To solve this problem, the above method of random step size numerical gradient estimation is introduced. This method allows us to apply geometric constraints to a special set that contains the gradients of the SDFs of the sampling points we have slightly processed. It adds a certain degree of uncertainty to the regions of the sampling points that belong to the fine structures, which makes the geometric constraints no longer act globally and uniformly, but can act adaptively on large-scale structures and relax the geometric constraints on small-scale structures. In the experiment, it is found that this method can effectively accelerate the convergence rate and relieve the topological structure limitations brought by geometric constraints. Therefore, this method enables better handling of complex topological structures.

[0099] (4) Overall optimization framework.

[0100] Based on the above improvements, a general framework is needed to integrate them. Therefore, a two-stage optimization framework, a rough-to-fine optimization method, is provided, which integrates the above design elements and finally realizes high-fidelity, fine surface reconstruction. The core idea of the two-stage optimization framework starts with a rough optimization stage, the goal of which is to establish a rough SDF and avoid the harmful effects brought by over-regularization or bias, which may lead to incorrect surface representations. Subsequently, in the second stage, the optimization process combines additional smoothing constraints and manual convergence strategies. These adjustments are carefully calibrated to produce a finer and smoother SDF, thus enhancing the details and fidelity in the final surface reconstruction.

[0101] In the rough optimization stage, the goal is to reconstruct the approximate rough structure of the 3D content. Random step size numerical gradient estimation and explicit bias correction are used to solve the initial reconstruction of the rough structure of the 3D content. In addition, the SDF-to-density conversion of VolSDF and the local scale modeling described in formula (5-1) are used to facilitate the formation of this main formula. Therefore, the training loss in this stage is formulated as:

[0102] (5 - 5)

[0103] Among them, is the color loss, which quantifies the difference between the rendered color and the true color. and is the loss weight. represents the set of rays in each training batch. is the ray 's true color. In addition, the local scale factor in Equation (5 - 1) can, to some extent, prevent the model from converging to the surface. Therefore, manually adjust the lower bound of the scale to prevent the ambiguity caused by too small a scale.

[0104] In the refinement stage, the estimated gradient is no longer used because the basic 3D content has been preliminarily restored and the problem of over - regularization is no longer a concern. Similarly, the explicit bias correction is removed because the significant surface errors caused by the bias have been preliminarily solved in the first stage. In summary, this stage combines the standard Eikonal loss and the smoothness constraint loss to achieve training. Among them, the smoothness constraint loss is:

[0105] (5 - 6)

[0106] Among them and is a random unit vector. is a smooth coefficient. In addition, when the network approaches convergence in this stage, the SDF - to - density conversion method of TUVR is adopted, ensuring the minimum deviation and retaining the details of the detailed objects. The loss function used in the second stage is defined as follows:

[0107] (5 - 7)

[0108] Among them , is the loss weight. Similar to the manual adjustment in the first stage, the goal in this stage is to promote the convergence from volume rendering to surface rendering, thus fully aligning the implicit surface with volume rendering. To achieve this, the lower bound of the scale factor is exponentially increased to a larger value .

[0109] To verify the effectiveness of this method, experimental evaluations were conducted on two benchmark datasets: Tanks and Temples (Knapitsch et al. 2017) and ScanNet++ (Yeshwanth et al. 2023). The Tanks and Temples dataset is characterized by its large-scale and diverse real-world scenes. In the experiment, six scenes from the training subset were used, which are the same as those used by Neuralangelo to maintain comparability. In addition, the validation was extended to four large indoor scenes in the Advance subset to further evaluate the robustness of the method. For the ScanNet++ dataset, which is characterized by its high-quality indoor scenes and supplemented with DSLR-quality images, eight scenes were selected for our analysis. Meshes were extracted by marching cubes, and the F1 score was calculated for evaluation.

[0110] (1) Results of the Tanks and Temples dataset.

[0111] As shown in Table 1, in terms of the average F1 score for evaluating the model performance, this method performed excellently, outperforming the previous state-of-the-art methods. This result fully demonstrates the superiority and effectiveness of this method. The explicit bias correction technique of this method has shown remarkable effects in practical applications.

[0112] Figure 5 For the visual result comparison of (a) NeuS (b) Neuralangelo (c) this embodiment (d) GT point cloud on the Tanks and Temples dataset, as Figure 5 shown at the top, after being processed by our bias correction technique, the roof of the barn maintained the integrity of its structure without collapsing. This is in sharp contrast to other methods, as other methods usually cannot prevent such collapses. This result fully demonstrates the effectiveness and reliability of the explicit bias correction technique of this method. In addition, the two-stage optimization method has also shown remarkable effects in practical applications. Through this method, the problem of excessive geometric regularization was effectively alleviated, as Figure 5 shown at the bottom. This enables this method to maintain good geometric performance while maintaining accuracy. This method has shown superiority in many aspects. It not only achieved better results than Neuralangelo in detailed surface representation but also significantly reduced the model parameters, reaching 1 / 8 of Neuralangelo's. This result fully demonstrates the efficiency of this method.

[0113] Table 1 Comparison of numerical results on the training set of the Tanks and Temples dataset

[0114]

[0115] The performance of this method on the Advance subset of Tanks and Temples is significantly better than previous work, as shown in Table 2. This result fully demonstrates the advantages and effectiveness of this method. In the Figure 6 comparison, the enhanced accuracy and integrity shown by this method in reconstructing large-scale surfaces are presented, where (a) reference image, (b) MonoSDF, (c) Neuralangelo, (d) this embodiment. This result once again verifies the superiority of this method in dealing with large-scale surface reconstruction. This method can also capture finer-grained details, which is particularly evident in the comparison with the Neuralangelo method. This shows that this method has higher accuracy and better performance in dealing with detail information. These excellent performances benefit from the explicit bias correction adopted in the first stage and the TUVR modeling adopted in the second stage. In the first stage, the fine structure is restored to the zero level set through explicit bias correction, and then in the second stage, a sufficiently small bias is maintained by adopting TUVR modeling. This method can refine the high-quality surface at the initial stage, thus maintaining high-quality reconstruction results in subsequent stages. This method shows superior performance in dealing with various different scenarios and tasks and surpasses previous work in many aspects.

[0116] Table 2 Comparison of numerical results on the Advance subset of the Tanks and Temples dataset

[0117]

[0118] (2) Results on the ScanNet++ dataset.

[0119] The quantified results are shown in Table 3. From this table, it can be seen that in most scenarios, this method surpasses other compared methods. In terms of the F1 score, the results of this method are comparable to those of methods with prior knowledge, indicating that this method performs well in terms of accuracy and precision. More specifically, as Figure 7As shown, where (a) MonoSDF, (b) Neuralangelo, (c) this embodiment, (d) GT mesh. When this method processes the reconstruction of large flat surfaces such as walls, it has a more accurate surface reconstruction result compared to the Neuralangelo method. This indicates that this method has more advantages in processing the reconstruction of large flat surfaces and can reconstruct the surface of an object more accurately. In addition, this method retains more fine structural details than the MonoSDF method. This shows that when this method processes the reconstruction of complex and fine structures, it can retain more detailed information, making the reconstruction result more realistic and accurate. Generally speaking, this method demonstrates superior performance in processing the reconstruction of various different scenarios. Whether it is in processing the reconstruction of large flat surfaces or in processing the reconstruction of complex and fine structures, this method can provide high-quality and accurate reconstruction results.

[0120] Table 3 Comparison of numerical results on the ScanNet++ dataset

[0121]

[0122] (3) Ablation experiment analysis.

[0123] To verify the effectiveness of the proposed technique, an ablation study was conducted on the Meetingroom scene of the Tanks and Temples dataset.

[0124] As Figure 8 As shown in (a), (f) is the GT point cloud. If the global scale is applied for the conversion from SDF to volume density, it will lead to inaccurate surface reconstruction. This is mainly because it is assumed that the volume density within the same level set is uniform. This assumption causes the surface with delicate texture to converge to the wrong position. Therefore, when dealing with surfaces with complex textures, this method can avoid this problem. Figure 8 (b) shows that without the random step size numerical gradient estimation, it will hinder the model from forming arbitrary topologies, resulting in incorrect surface reconstruction. This indicates that in this method, the random step size numerical gradient estimation plays a crucial role in forming arbitrary topologies and ensuring surface accuracy. In Figure 8 In (c), the incorrect ground collapse is due to the deviation of the volume density. If the explicit deviation correction of this method is not applied, this deviation problem will lead to obvious surface errors. This once again proves the importance of the explicit deviation correction method in ensuring surface accuracy. Figure 8 (d) shows the result of optimizing in a single stage. Under the random step size numerical gradient estimation, the preference of the Eikonal loss for smoothness is somewhat compromised, resulting in a rougher surface. This shows that the two-stage optimization strategy has advantages in ensuring surface smoothness and accuracy. Finally, Figure 8(e) shows our complete model. It can be seen that while retaining most of the details, this method achieves a smooth surface with high fidelity. This result fully demonstrates that this method can provide high-quality and accurate reconstruction results when dealing with various different scenarios and tasks.

[0125] (4) Analysis of the numerical gradient estimation with random step sizes.

[0126] Further experiments show the impact of the numerical gradient estimation with random step sizes of this method on the optimization process, as Figure 9 shown, where (a) the depth map generated at 7500 steps in the first stage, (b) the reference image, and (c) the final mesh. The generated depth map indicates that Model A is severely constrained by geometric regularization, making it difficult to change its topology during the optimization process. Model B adopts a progressive numerical gradient estimation technique, and even after 7500 steps, it does not generate an accurate depth map. However, by using the numerical gradient estimation with random step sizes of this embodiment, an approximately accurate scene geometry is achieved within only 7500 steps.

[0127] Through further experiments, as Figure 9 shown, the important impact of the numerical gradient estimation with random step sizes on the optimization process is demonstrated. In these experiments, the generated depth maps provide us with an intuitive understanding of the model performance. It is observed that Model A is severely constrained by geometric regularization, which makes it difficult for the topology of the model to change during the optimization process. This severe constraint limits the flexibility of the model and makes it difficult to handle the reconstruction of complex scenes. It is observed that Model B adopts a progressive numerical gradient estimation technique. However, even after 7500 steps of optimization, Model B does not generate an accurate depth map. This indicates that there may be problems with the progressive numerical gradient estimation technique in terms of optimization efficiency and effectiveness. However, when using the numerical gradient estimation with random step sizes of this embodiment, the situation changes significantly. Within only 7500 steps, this embodiment (Model C) achieves an approximately accurate scene geometry. This result fully demonstrates the advantages of the numerical gradient estimation with random step sizes in terms of optimization efficiency and effectiveness.

[0128] At this stage, it can be seen that the depth map of Model C is similar to that of Instant-NGP, but the depth map of this embodiment shows a more natural and smooth depth contour. This indicates that, similar to Instant-NGP, the method can freely change the topology shape during the optimization process while also maintaining a natural zero-level set surface, which greatly enhances the accuracy and reliability of the model.

[0129] In addition, it can be observed that the final mesh of Model B still exhibits inaccurate floor collapse. This is because during the optimization process, Model B was unable to effectively handle the problem of large-scale planes. However, the mesh of Model C, through random step-size numerical gradient estimation, maintains a correct and smooth floor. This is mainly because the floor is a large-scale plane, which enables randomness to continuously apply the Eikonal constraint in large-scale regions, thus generating natural zero-level sets in these vast regions.

[0130] However, for the method using progressive steps, once the step size is reduced to a certain extent, it no longer imposes the Eikonal constraint on large-scale surfaces, which leads to unnatural zero-level sets. This shows that the random step-size numerical gradient estimation in this embodiment has better performance and effect compared to the progressive-step method when dealing with large-scale surfaces.

[0131] (5) Analysis of explicit bias correction.

[0132] To confirm the versatility and effectiveness of the explicit bias correction method in this embodiment, further experiments were conducted to verify its potential as a plug-and-play bias correction method independent of the complete model of this method. This correction technique was implemented in the coarse optimization stage of this method and tested on various renderers, including NeuS, VolSDF, and TUVR. These experiments aimed to evaluate the adaptability and effectiveness of bias correction in various SDF-to-volume density modeling.

[0133] To confirm the versatility and effectiveness of the explicit bias correction method, a series of further experiments were conducted, aiming to verify its potential as a plug-and-play bias correction method that can be used independently of the complete model of this method.

[0134] First, this explicit bias correction technique was implemented in the coarse optimization stage. It was tested on various renderers, including NeuS, VolSDF, and TUVR. These renderers are currently commonly used and highly representative. By testing on these renderers, the performance and effect of the explicit bias correction method can be evaluated more comprehensively and accurately. The purpose of these experiments was to evaluate the adaptability and effectiveness of the explicit bias correction method in various SDF-to-volume density modeling. Through these experiments, it was confirmed that the explicit bias correction method can not only be used as part of the complete model but also as an independent, plug-and-play bias correction method, showing good adaptability and effect for various different models and tasks.

[0135] As shown in Figure 10, where (a) does not use display deviation correction and (b) uses display deviation correction, an obvious deviation problem can be observed in the ceiling of the room. This problem is manifested as the collapse of the ceiling shape. Even when using the TUVR modeling technique that is mathematically proven to be unbiased in the continuous form, this problem still exists. This phenomenon reveals that in practical applications, even a theoretically unbiased modeling technique may lead to deviations in actual results due to various factors, such as model complexity, data noise, etc. However, when applying the explicit deviation correction method provided by this method, the situation has changed significantly. It can be seen that the problem of ceiling collapse has been significantly improved. This result fully proves the effectiveness of the explicit deviation correction method. Through this method, not only can the deviation of the model be corrected, but also the accuracy and stability of the model can be improved. The explicit deviation correction method provides an effective solution to the deviation problem in modeling. By applying this method, while ensuring the accuracy of the model, the stability and reliability of the model can also be maintained.

[0136] In summary, for the problem that the existing methods do not consider geometric over-regularization, this method can make the optimization process of the SDF-based volume rendering method very similar to that of the volume density-based volume rendering method, enabling our method to freely change the topological structure, thus having a faster optimization process.

[0137] Embodiment 2

[0138] This embodiment provides an electronic device, including: one or more processors and a memory. The memory stores one or more programs, and the one or more programs include instructions for executing the two-stage surface reconstruction method based on SDF volume rendering as described in Embodiment 1.

[0139] Embodiment 3

[0140] This embodiment provides a computer-readable storage medium, including one or more programs for execution by one or more processors of an electronic device. The one or more programs include instructions for executing the two-stage surface reconstruction method based on SDF volume rendering as described in Embodiment 1.

[0141] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A two-stage surface reconstruction method based on SDF volume rendering, characterized in that: Based on the acquired 2D image, surface reconstruction is achieved through volume rendering using a pre-trained geometry network and a color network, wherein the training process of the geometry network and the color network includes the following steps: Step S1, based on the acquired sample image and the corresponding camera parameters, calculate the ray information of the camera position along the view direction, use the point coordinates on the ray as the input of the geometric network and the color network, and obtain the predicted distance sign function information and color information respectively; Step S2, taking uncertainty into account, obtaining an estimated gradient of the distance sign function information by random step numerical gradient estimation, combining the Eikonal loss of the estimated gradient, the color loss calculated based on the color information, and the bias correction loss, and implementing the first stage of training under the premise of an adaptive scale factor and a low lower limit of the scale factor; Step S3, calculating the true analysis gradient of the distance sign function information, combining the Eikonal loss of the true analysis gradient, the color loss calculated based on the color information and the smoothness constraint loss, and realizing the second stage of training under the premise of an adaptive scale factor and a high lower limit of the scale factor.

2. A two-stage surface reconstruction method based on SDF volume rendering according to claim 1, characterized in that: The step S1 is implemented by the following formula: , , in, , They are the geometric network and the color network, is the camera position Emitted along the view direction The coordinates of the point on the ray, , , They are respectively the predicted distance sign function information, scale factor, and geometric features. is the conversion function from the preset distance sign function to volume density, is the volume density converted from the distance sign function, is the predicted color information.

3. A two-stage surface reconstruction method based on SDF volume rendering according to claim 2, characterized in that: The conversion function is a rendering distance function.

4. The two-stage surface reconstruction method based on SDF volume rendering according to claim 1, characterized in that: In step S2, the loss function is: , in, is the loss function of the first stage of training, For color loss, Eikonal loss for estimating gradients for distance sign function information, is the bias correction loss, and is the loss weight.

5. The two-stage surface reconstruction method based on SDF volume rendering according to claim 1, characterized in that: In the step S2, the estimated gradient of the distance sign function information is Quantity for: , in, is the coordinate of the point on the ray emitted from the camera position along the view direction, is the predicted distance sign function information, and , is the default value.

6. The two-stage surface reconstruction method based on SDF volume rendering according to claim 1, characterized in that: The deviation correction loss for: , in, is the coordinate of the point on the ray emitted from the camera position along the view direction, is the predicted distance sign function information, is the deviation correction factor, , is the volume density converted from the distance sign function, For each training batch Rays collection.

7. The two-stage surface reconstruction method based on SDF volume rendering according to claim 1, characterized in that: In step S3, the loss function is: , in, is the loss function of the second stage of training, For color loss, is the true analytical gradient of the sign-free function information, is the smoothness constraint loss, and is the loss weight.

8. The two-stage surface reconstruction method based on SDF volume rendering according to claim 1, characterized in that: The smoothness constraint loss for: , in, is the coordinate of the point on the ray emitted from the camera position along the view direction, For point The normal vector, is the distance sign function information The true analytical gradient of , is a random unit vector, is the smoothness coefficient, For a batch The number of internal rays, The number of sampling points for each ray.

9. An electronic device, characterized in that: include: One or more processors and a memory, wherein the memory stores one or more programs, wherein the one or more programs include instructions for executing the two-stage surface reconstruction method based on SDF volume rendering as claimed in any one of claims 1-8.

10. A computer-readable storage medium, characterized in that: The method comprises one or more programs for execution by one or more processors of an electronic device, wherein the one or more programs comprise instructions for executing the two-stage surface reconstruction method based on SDF volume rendering as claimed in any one of claims 1 to 8.