Nerve radiation field two-stage three-dimensional reconstruction method based on sign distance function

By dividing the three-dimensional reconstruction process into two stages, using symbol distance function and multi-resolution hash encoding neural radiation field method, the problems of large computing resources and rough surface in the prior art are solved, and a more efficient and accurate three-dimensional reconstruction effect is achieved.

CN120070800APending Publication Date: 2025-05-30BEIJING UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510143845.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing three-dimensional reconstruction methods consume a lot of computing resources when dealing with complex scenarios, require a lot of prior knowledge and manual intervention, and the generated surface is rough, making it difficult to accurately describe the boundaries of complex objects shapes.

Method used

The two-stage three-dimensional reconstruction method of neural radiation field based on symbol distance function is used, and the reconstruction process is divided into two stages, respectively, the scene color and volume density are processed, training is accelerated using implicit representation of symbol distance function and multi-resolution hash encoding, and the reconstruction accuracy is improved through multi-view geometric constraints.

Benefits of technology

The training efficiency and reconstruction effect of three-dimensional reconstruction are improved, the generated surface is more accurate, the shape boundaries of complex objects can be better described, and the rendering efficiency and operational convenience are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070800A_ABST
    Figure CN120070800A_ABST
Patent Text Reader

Abstract

The invention discloses a two-stage three-dimensional reconstruction method for a neural radiation field based on a symbolic distance function. A more accurate and real object surface grid is generated from a two-dimensional image. The method comprises the following steps of: dividing a reconstruction process into two stages, and respectively modeling the color and the volume density of an object by using two different networks; in the first stage, a zero-order set of a symbol distance field is used for representing the surface of an object, and rough grids are preliminarily extracted; in the second stage, loss is optimized continuously, and the vertex position and the surface density are adjusted step by step to achieve refinement of the grid surface; multi-resolution hash coding is utilized to accelerate training; three-dimensional position and view vector information are added into each layer of a multi-layer perceptron, supervision network training is displayed through multi-view geometric constraints, and reconstruction quality is improved. According to the method, the three-dimensional reconstruction process is divided into two stages, a good view synthesis effect is kept by using implicit representation and geometric constraint, and the precision and speed of model training and reconstruction are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of 3D reconstruction, and more specifically, to a two-stage 3D reconstruction method for neural radiance fields based on the signed distance function. Background Art

[0002] 3D reconstruction generally refers to the recovery of a 3D model of an object or scene from an image sequence. Currently, the application fields of 3D reconstruction mainly include virtual reality, augmented reality, intelligent driving, medical imaging, and cultural heritage protection, etc.

[0003] Traditional 3D reconstruction methods mainly rely on the principles of multi-view geometry, and complete the conversion from 2D images to 3D models by estimating camera poses and extracting and matching image features. These methods mainly include voxel-based, surface-based, point-cloud-based, and depth-map-based reconstruction techniques. By analyzing information such as color, brightness, and texture between different perspectives, these methods use photometric consistency and geometric consistency to estimate the depth information of the scene, and then construct a 3D model of the target scene. Usually, they require high computing resources, a large amount of prior knowledge, and manual intervention, and have limited effects when dealing with complex scenes.

[0004] 3D reconstruction algorithms based on deep learning utilize the characteristics of neural networks to convert the texture information of images into high-level semantic information, avoiding the complexity and errors of feature point extraction and matching in traditional methods, and being able to reconstruct complex scenes more accurately. They have become the mainstream methods for current 3D reconstruction tasks. Among 3D reconstruction methods based on deep learning, Neural Radiance Fields (NeRF for short) is an advanced 3D scene reconstruction and rendering technology. By using a multi-layer perceptron (MLP) and a volume rendering process to represent and generate complex 3D scenes, it can generate high-quality images from different perspectives. NeRF has achieved good results in reconstructing and rendering large-scale scenes with realistic details from a set of images. However, its representation and rendering processes rely on implicit functions and specialized ray marching algorithms, resulting in complex operations and slow rendering speeds. Instant-NGP can accurately represent scenes with a small and efficient MLP by using multi-resolution hash positional encoding as additional learned features, greatly accelerating the training and inference speeds. Based on NeRF, NeRF2Mesh further explores how to convert implicit volume representations into explicit surface representations to improve rendering efficiency and operational convenience. NeRF2Mesh extracts the object surface mesh from NeRF through an adaptive surface refinement algorithm and performs fine processing to generate a mesh with high-quality texture. However, since NeRF2Mesh uses a volume density-based method to represent the scene, the generated surface is relatively rough and cannot accurately describe the boundaries of complex object shapes.

[0005] Following NeRF, subsequent research combined implicit surfaces with differentiable volume rendering, represented the implicit surface as a Signed Distance Function (SDF), and used the zero-level SDF set to describe the geometry, thus achieving high-quality reconstruction on a single object. The signed distance function is a function used to describe the distance from any point in space to the nearest object surface. The SDF not only returns the distance value but also carries sign information to distinguish whether the point is inside, outside, or on the surface of the object. Specifically, points with a positive SDF value are outside the object, and the distance value represents the shortest distance from the point to the object surface. Points with a negative SDF value are inside the object, and points with an SDF value of 0 are exactly on the surface of the object. NeuS combines the advantages of surface-based rendering methods and volume-based rendering methods, constrains the scene space to a signed distance function, and applies volume rendering to train this robust representation. However, due to the relatively complex MLP structure of NeuS, the query calculation for each sampling point consumes a large amount of resources, and the scene geometry and color are not separated, resulting in low training efficiency and a long optimization process, especially when dealing with complex scenes with a large number of input images.

[0006] The present invention discloses a two-stage three-dimensional reconstruction model of a neural radiance field based on a signed distance function, aiming to generate a more accurate and realistic surface mesh from images. The novelty and creativity of the present invention are reflected in: ① dividing the three-dimensional reconstruction process into two stages, separating the scene color and volume density, and using two networks for training. Compared with the reconstruction method using a single network, the training efficiency is higher and the reconstruction effect is better; ② using an implicit representation based on the signed distance function to characterize the scene, obtaining better geometric effects, using the object surface represented by the zero-level set of the signed distance function as the basis for optimization, and adopting a multi-resolution hash encoding structure to accelerate training; ③ explicitly integrating the three-dimensional position and view vector inputs into each layer of the multi-layer perceptron (MLP), and using multi-view geometric constraints to explicitly supervise the training of the SDF network, improving the accuracy and geometric quality of the reconstruction. Summary of the Invention

[0007] In view of this, an embodiment of the present invention provides a two-stage three-dimensional reconstruction method of a neural radiance field based on a signed distance function to achieve the reconstruction from multi-view images to the three-dimensional mesh of an object.

[0008] To achieve the above object, the embodiment of the present invention provides the following solutions:

[0009] Step 1. Initialize the network parameters and load the scene images for reconstructing the three-dimensional mesh;

[0010] Construct a 3D reconstruction model that includes a density network and a color network; initialize the input and network parameters of the network. The input is a 5D vector, including the 3D coordinate position x = (x, y, z) of a spatial point and the direction d = (θ, φ); initialize the model parameters using random initialization.

[0011] Step 2. Train the NeRF model to learn the continuous volume representation of the scene and continuously optimize the loss until the MLP network training is completed, which includes the following 2 steps:

[0012] Step 2.1 Perform positional encoding on the input, encoding the input 3D coordinates and viewing directions into a form that can be processed by the neural network to retain spatial information and improve the performance of the model. For the input 3D coordinates, calculate their hash values at different resolution levels, and then use these hash values to look up the corresponding trainable feature vectors in the hash table. Linearly interpolate the feature vectors of the voxels around each input coordinate to generate a feature representation that matches the input coordinate. Finally, concatenate the interpolated feature vectors at different resolution levels to form the final feature input. Positional encoding effectively reduces parameter storage while improving the expressive power of features.

[0013] Step 2.2 Use the zero-level set of the neural implicit signed distance function to represent the surface. The signed distance field implicitly represents the 3D scene using two functions, namely the SDF function f: R 3 →R, which maps the spatial position to the signed distance from that position to the object surface, and the color function c: R 3 ×S 2 →R 3 , which maps the spatial position and viewing direction to the RGB color associated with the input. Both of these functions are encoded by a multi-layer perceptron (MLP). The surface of the object is represented by the zero-level set of its SDF, that is:

[0014] S = {x ∈ R 3 |f(x) = 0}

[0015] Step 3. Sample the volume density and color, and render the image according to the density and color of the sampled points;

[0016] In volume rendering, first define the rays starting from the camera. Each ray has a far boundary t f and a near boundary t nDefinition: These two boundaries determine where the ray starts and ends in 3D space. Along each ray, density values σ and color values c(t) seen from the viewing direction d are calculated at multiple sampling points r(t) using a neural network. These sampling points are uniformly distributed or distributed according to a certain strategy (such as stratified sampling) to ensure that the entire ray path is fully considered. During the volume rendering process, the ray transfer equation describes how light interacts with the medium as it propagates through the medium. Specifically, for a ray, its color can be expressed as:

[0017]

[0018] where c(t) is the color of the sampling point and σ(t) is the density value of the sampling point.

[0019] Step 4. Based on the calculated SDF values, use the Marching Cubes algorithm to extract a relatively rough mesh, which specifically includes the following 3 steps:

[0020] Step 4.1 Preprocessing and cell extraction:

[0021] Read the original data into a specified array, extract a cell from the grid data volume, and obtain all information of the cell at the same time. Compare the SDF values of the 8 vertices of the current cell with the given isosurface value C to obtain the status table of the cell.

[0022] Step 4.2 Intersection point and normal vector calculation:

[0023] According to the status table index of the current cell, find the cell edges that intersect the isosurface, and use linear interpolation to calculate the position coordinates of each intersection point. Use the central difference method to find the normal vectors of the 8 vertices of the current cell, and then use linear interpolation to obtain the normals of each vertex of the triangular patch.

[0024] Step 4.3 Image rendering:

[0025] Render the isosurface image based on the coordinates and vertex normals of each vertex of the triangular patch.

[0026] Step 5. Generate a refined mesh;

[0027] Use the rough mesh generated in the first stage as input, further refine it, optimize it into a fine mesh with a more accurate surface and adaptive surface density, and perform texture editing;

[0028] Preferably, when the algorithm uses a loss function to train and optimize model parameters, a photometric consistency loss is introduced. For each sampled block, the best four NCC (Normalized Cross-Correlation) scores are found, and these scores are used to calculate the photometric consistency loss for the corresponding view. The method for calculating the photometric consistency loss for the corresponding view using the NCC scores is as follows:

[0029]

[0030] Preferably, the base resolution level of the density grid for multi-resolution hash encoding is 16, and the hash table size is 2 19 。

[0031] Preferably, the batch size is set to 100 during the training of the algorithm.

[0032] Preferably, the algorithm is trained for a total of 300 epochs.

[0033] Preferably, during the training of the algorithm, an exponentially decaying learning rate schedule is used, and the learning rate decays gradually from an initial 5×10 -3 to 5×10 -4 。

[0034] The present invention discloses a two-stage 3D reconstruction method based on the signed distance function for neural radiance fields, which generates a more accurate and realistic object surface mesh from 2D images. The above method includes: dividing the reconstruction process into two stages, using two different networks to model the object color and volume density respectively; in the first stage, using the zero-level set of the signed distance field to represent the object surface and performing a preliminary extraction of the rough mesh; in the second stage, continuously optimizing the loss and gradually adjusting the vertex positions and surface densities to refine the mesh surface; using multi-resolution hash encoding to accelerate the training; adding 3D position and view vector information to each layer of the multi-layer perceptron, and displaying and supervising the network training through multi-view geometric constraints to improve the reconstruction quality. The method disclosed in the present invention divides the 3D reconstruction process into two stages, maintains a good view synthesis effect by using implicit representation and geometric constraints, and improves the accuracy and speed of model training and reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0036] Figure 1Schematic diagram of the two-stage 3D reconstruction model structure of the neural radiance field based on the signed distance function provided by the embodiments of the present invention;

[0037] Figure 2 Flowchart of the two-stage 3D reconstruction method of the neural radiance field based on the signed distance function provided by the embodiments of the present invention;

[0038] Figure 3 Schematic diagram of the network structures of the density network, specular light color network, and diffuse reflection color network provided by the embodiments of the present invention;

[0039] Figure 4 Schematic diagram of the signed distance function in the two-dimensional case provided by the embodiments of the present invention;

[0040] Figure 5 Schematic diagram of the volume rendering process in the two-stage 3D reconstruction method of the neural radiance field based on the signed distance function provided by the embodiments of the present invention;

[0041] Figure 6 Schematic diagram of all possible states of the marching cube cells provided by the embodiments of the present invention. Detailed implementation manners

[0042] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0043] The 3D reconstruction technology based on the neural radiance field aims to reconstruct a high-quality 3D scene from multi-view images. By training a neural network to capture the complex geometric structure and appearance information of the scene, a coarse-to-fine 3D reconstruction is achieved. The model of the two-stage 3D reconstruction network of the neural radiance field based on the signed distance function proposed by the present invention is divided into two stages, as Figure 1 shown. In the first stage, the multi-view image information in the scene is learned through the density network, and after training, the density value of the scene is obtained, and a rough grid is derived. In the second stage, through the combination of differentiable rendering technology and iterative surface refinement algorithm, the grid derived in the first stage is continuously refined. The color network is trained in both stages, and two MLP networks are respectively used to decompose the appearance into view-independent diffuse reflection colors and view-dependent specular colors.

[0044] The embodiments of the present invention disclose a two-stage 3D reconstruction method of the neural radiance field based on the signed distance function to achieve the surface reconstruction of an object. Refer to Figure 2 , the above method at least includes the following 5 steps:

[0045] Step 1. Initialize the network parameters and load the scene image for reconstructing the 3D mesh.

[0046] The present invention constructs a 3D reconstruction model including a density network and a color network. As Figure 3 shown, the density network consists of a fully-connected multi-layer perceptron with 2 layers and 32 neurons in each layer; the color network consists of two networks. The specular color network is a fully-connected multi-layer perceptron with 3 layers and 64 neurons in each layer, and the diffuse color network is a fully-connected multi-layer perceptron with 2 layers and 32 neurons in each layer.

[0047] The network input is a 5D vector, including the 3D coordinate position x = (x, y, z) of a spatial point and the direction d = (θ, φ); the hyperparameters set in the present invention include the initial learning rate, the number of iterations, the density grid occupancy threshold, etc. An exponential decay learning rate schedule is adopted, and the learning rate decays from the initial 5×10 -3 to 5×10 -4 , with a total of 30,000 iterations. When the density value of a voxel is greater than or equal to the threshold 10, it is considered that the voxel is occupied by the scene.

[0048] Step 2. Train the NeRF model to learn the continuous volume representation of the scene and continuously optimize the loss until the MLP network training is completed, including the following 2 steps:

[0049] Step 2.1 Perform position encoding on the input, encoding the input 3D coordinates and viewing direction into a form that can be processed by the neural network to retain spatial information and improve the performance of the model.

[0050] In the neural radiance field, multi-resolution hash encoding can be used to accelerate the training process. It can select appropriate levels for sampling according to needs, thereby improving the sampling efficiency and maintaining the accuracy of the scene. The present invention combines multi-resolution hash encoding to represent the signed distance field of the scene. The voxel grid corresponding to the sampling point is calculated through the 3D position information P, and then through the hash algorithm, it is mapped to obtain the corresponding feature vector for each resolution level in the hash table. For each 3D position x, the learnable hash table entry Ω is used to map x to the multi-resolution hash encoding h Ω (x). x and h Ω (x) are used as the inputs of the density network and the color network.

[0051] The spatial information is expressed by grids with different resolutions, and the size of the hash table is set to a fixed value. The obtained multiple hash tables of the same size are multi-space hash tables. When the resolution exceeds a certain value, conflicts will occur, and the conflicts are eliminated through the learning of the neural network.

[0052] Step 2.2 Use the zero-level set of the neural implicit signed distance function to represent the surface.

[0053] The signed distance field implicitly represents a 3D scene using two functions, namely the SDF function f: R 3 →R, which maps a spatial position to the signed distance from that position to the object surface, and the color function c: R 3 ×S 2 →R 3 , which maps a spatial position and a viewing direction to the RGB color associated with the input. Both of these functions are encoded by a multi-layer perceptron (MLP).

[0054] The signed distance function (SDF) is an implicit function used to represent object surfaces in 3D space. It represents by calculating the distance from any point in space to the nearest object surface and assigning a sign to that distance. The SDF representation of a scene is as Figure 4 shown. For a point x in space, the value of the SDF, SDF(x), represents the distance from that point to the nearest object surface. If the point is inside the object surface, the distance is negative; if the point is outside the object surface, the distance is positive; if the point is exactly on the object surface, the distance is zero. The zero isosurface of the SDF (i.e., the set of points where the SDF value is zero) defines the object surface, that is: when the SDF value is zero, it means the point is exactly on the object surface. The object surface is represented by the zero-level set of its SDF:

[0055] S = {x ∈ R 3 | f(x) = 0} (1)

[0056] Step 3. Sample volume density and color, and render an image according to the density and color of the sampled points using the volume rendering principle;

[0057] In volume rendering, first define the rays starting from the camera, as Figure 5 shown. Each ray is defined by a far boundary t f and a near boundary t n , which determine where the ray starts and ends in 3D space. Along each ray, use a neural network to calculate the density value σ and the color value c(t) seen from the viewing direction d at multiple sampled points r(t). These sampled points are distributed according to a stratified sampling strategy to ensure that the entire ray path is fully considered. A series of uniformly spaced sampled points are generated between the near and far ends of the ray. At the same time, to reduce noise in the rendering process, a random perturbation is added to each sampled point. Specifically, for each sampled point, calculate the midpoint of its adjacent sampled points, and then randomly select a position between these midpoints as the final sampled position.

[0058] During the volume rendering process, the light transport equation describes how light interacts with the medium as it propagates through the medium. Specifically, for a ray, at each sampling point, the corresponding density value σ and color value c are calculated. The density value describes the occupancy of that point in the volume, and the color value represents the color information of that point. Then, according to the ray marching algorithm, the colors and density values of these sampling points are accumulated to generate the final pixel color. The accumulation process is usually expressed by the following formula:

[0059]

[0060] where c(t) is the color of the sampling point, σ(t) is the density value of the sampling point, C(r) is the final pixel color. T(t) is the cumulative transparency of all previous sampling points on the path. When accumulating, it is multiplied by the distance Δt between each volume density value and adjacent sampling points, representing the remaining light intensity when the light travels from the camera to the current sampling point t. The cumulative transmittance T(t) is a value ranging from 1 (fully transparent) to 0 (fully absorbed), indicating that the light gradually weakens during propagation due to absorption and scattering of the medium. The calculation formula of T(t) can be expressed as:

[0061]

[0062] Step 4. According to the calculated SDF values, use the Marching Cubes algorithm to extract a relatively rough mesh, which specifically includes the following 3 steps:

[0063] Step 4.1 Preprocessing and cell extraction:

[0064] Read the original data into a specified array, extract a cell from the mesh data volume, and obtain all information of the cell at the same time. Compare the SDF values of the 8 vertices of the current cell with the given isosurface value C to obtain the state table of the cell. There are 256 possible states, but considering rotational symmetry, there are actually only 15 basic patterns, such as Figure 6 shown

[0065] Step 4.2 Intersection point and normal vector calculation:

[0066] According to the state table index of the current cell, find the cell edges that intersect with the isosurface, and use linear interpolation to calculate the position coordinates of each intersection point. Use the central difference method to find the normal vectors of the 8 vertices of the current cell, and then use linear interpolation to obtain the normals of each vertex of the triangular patch.

[0067] Step 4.3 Image rendering:

[0068] The isosurface image is drawn based on the coordinates and vertex normal vectors of the vertices of each triangular patch.

[0069] Step 5. Generate a refined mesh;

[0070] Take the coarse mesh generated in the first stage as the input, further refine it, optimize it into a fine mesh with a more accurate surface and adaptive face density, and perform texture editing;

[0071] The present invention trains and optimizes the model parameters by minimizing the rendering loss and the geometric loss. For the rendering loss, minimize the loss between the predicted color and the true color of each pixel, that is:

[0072]

[0073] In order to more accurately simulate the optical properties of the object surface, improve the visual effect and rendering efficiency of the rendering, calculate the diffuse color and the specular light color separately. Separate the diffuse and specular terms by applying L2 regularization to the specular color:

[0074]

[0075] For the geometric loss, in order to effectively handle the occlusion problem in image matching and 3D reconstruction, find the best four NCC (Normalized Cross-Correlation) scores for each sampling block, and use these scores to calculate the photometric consistency loss of the corresponding view. The method of calculating the photometric consistency loss of the corresponding view using the NCC scores is as follows:

[0076]

[0077] The present invention verifies the performance of the disclosed method on the NeRF Synthetic dataset. The NeRF Synthetic dataset is a synthetic dataset for training and evaluating the NeRF algorithm, containing eight complex non-Lambertian scenes finely modeled by Blender software: Chair, Drums, Ficus, Hotdog, Lego, Materials, Mic, and Ship. The dataset is divided into training, validation, and test sets, containing 100, 100, and 200 images of 800×800 pixels respectively.

[0078] The present invention uses Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index (SSIM), and Learned Perceptual Image Patch Similarity (LPIPS) as quantitative evaluation metrics for view synthesis quality, and uses Chamfer Distance (CD) as a quantitative evaluation metric for the reconstructed 3D model. PSNR evaluates the fidelity of an image by calculating the mean squared error between the original image and the distorted image, and a higher value indicates better image quality. SSIM evaluates image quality by comparing the similarity in brightness, contrast, and structure between the original image and the distorted image, and a value closer to 1 indicates better image quality. LPIPS measures the perceptual difference between images by comparing the feature representations of the images, and a lower value indicates better image quality. The Chamfer Distance evaluates the matching degree between two point sets by calculating the sum of the closest point distances between them, and a lower value indicates a higher similarity between the point cloud or mesh. Tables 1, 2, 3, and 4 are the PSNR, SSIM, LPIPS, and CD comparison data of the method disclosed in the present invention and algorithms such as NeRF2Mesh for the reconstruction results of each scene on the NeRF Synthetic dataset. It can be seen that the method disclosed in the present invention achieves better reconstructed mesh quality while ensuring the rendering quality.

[0079] Table 1 PSNR Comparison on NeRF Synthetic Dataset

[0080]

[0081]

[0082] Table 2 SSIM Comparison on NeRF Synthetic Dataset

[0083]

[0084] Table 3 LPIPS Comparison on NeRF Synthetic Dataset

[0085]

[0086] Table 4 CD Comparison on NeRF Synthetic Dataset

[0087]

[0088] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A two-stage 3D reconstruction method of neural radiation field based on signed distance function, characterized in that: The following steps are involved: Step 1. Initialize the network parameters and load the scene image for reconstructing the 3D mesh; Construct a 3D reconstruction model including a density network and a color network; initialize the network input and network parameters, the input is a 5-dimensional vector, including the 3D coordinate position x = (x, y, z) and direction d = (θ, φ) of a spatial point; use random initialization to initialize the model parameters; Step 2. Train the NeRF model to learn the continuous volume representation of the scene and continuously optimize the loss until the MLP network training is completed. It includes the following two steps: Step 2.1 performs position encoding on the input, encoding the input 3D coordinates and viewing direction into a form that can be processed by the neural network to retain spatial information and improve the performance of the model; for the input 3D coordinates, calculate their hash values ​​at different resolution levels, and then use these hash values ​​to find the corresponding trainable feature vectors in the hash table; linearly interpolate the feature vectors of the surrounding voxels of each input coordinate to generate a feature representation that matches the input coordinate; finally, splice the interpolated feature vectors at different resolution levels together to form the final feature input; Step 2.2 uses the zero-order set of neural implicit signed distance functions to represent the surface; the signed distance field implicitly represents the three-dimensional scene using two functions, namely the SDF function f: R 3 →R, which maps the spatial position to the signed distance from that position to the surface of the object, and the color function c: R 3 ×S 2 →R 3 , mapping the spatial position and viewing direction to the RGB color associated with the input, both functions are encoded by the multi-layer perceptron MLP; the surface of the object is represented by the zero-level set of its SDF, namely: S={x∈R 3 |f(x)=0} Step 3. Sample volume density and color, and render the image using volume rendering principle according to the density and color of the sampling points; In volume rendering, we first define rays starting from the camera. Each ray is connected by the far boundary t f and the proximal boundary t n Definition: These two boundaries determine where the ray starts and ends in three-dimensional space; along each ray, a neural network is used to calculate the density value σ and the color value c(t) seen from the viewing direction d at multiple sampling points r(t); these sampling points are distributed uniformly or according to a certain strategy to ensure that the entire ray path is fully considered; during volume rendering, the light transport equation describes how light interacts with the medium as it propagates in the medium; specifically, for a ray, its color is expressed as: Among them, c(t) is the color of the sampling point, σ(t) is the density value of the sampling point; Step 4. Based on the calculated SDF value, use the marching cubes algorithm to extract a coarser grid, which includes the following three steps: Step 4.1 Preprocessing and unit cell extraction: Read the original data into the specified array, extract a unit cell from the mesh data volume, and obtain all the information of the unit cell; compare the SDF values ​​of the 8 vertices of the current unit cell with the given isosurface value C to obtain the state table of the unit cell; Step 4.2 Intersection point and normal vector calculation: According to the state table index of the current unit body, find the unit body edge that intersects with the isosurface, and use the linear interpolation method to calculate the position coordinates of each intersection point; use the central difference method to calculate the normal vectors of the 8 vertices of the current unit body, and then use the linear interpolation method to get the normal direction of each vertex of the triangle patch; Step 4.3 Image drawing: Draw the isosurface image according to the coordinates of each triangle vertex and vertex normal vector; Step 5. Generate a refined mesh; The coarse mesh generated in the first stage is taken as input and further refined to be optimized into a fine mesh with more accurate surface and adaptive surface density, and texture editing is performed.

2. The two-stage three-dimensional reconstruction method of neural radiation field based on signed distance function according to claim 1, characterized in that: This method introduces photometric consistency loss when using the loss function to train and optimize the model parameters; find the best four normalized cross-correlation NCC scores for each sampling block, and use these scores to calculate the photometric consistency loss of the corresponding view; the method of using the NCC score to calculate the photometric consistency loss of the corresponding view is as follows:

3. The two-stage three-dimensional reconstruction method of neural radiation field based on signed distance function according to claim 1, characterized in that: Multiresolution hash-coded density grid base resolution level is 16, hash table size is 2 19 .

4. The two-stage three-dimensional reconstruction method of neural radiation field based on signed distance function according to claim 1, characterized in that: The batch size of this method is set to 100.

5. The two-stage three-dimensional reconstruction method of neural radiation field based on signed distance function according to claim 1, characterized in that: This method iterates for 300 epochs in total.

6. The two-stage three-dimensional reconstruction method of neural radiation field based on signed distance function according to claim 1, characterized in that: This method uses exponential decay learning rate scheduling, and the learning rate is 5×10 -3 Gradually decay to 5×10 -4 .

Citation Information

Cited By

  • Three-dimensional thermal environment neural field reconstruction method and system for building facade defect diagnosis

    CN120852695A