Multi-view neural implicit surface reconstruction method based on dynamic sampling range

By adopting dynamic sampling range and hierarchical volume sampling strategies in multi-view neural implicit surface reconstruction technology, the problem of low sampling efficiency in the existing technology is solved, and efficient and accurate three-dimensional geometric grid reconstruction is achieved.

CN119722908BActive Publication Date: 2025-05-06CHENGDU UNIV OF INFORMATION TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510222983.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-06
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

The existing multi-view neural implicit surface reconstruction technology has low accuracy and efficiency in reconstructing objects due to low sampling efficiency, which requires long-term training or additional information input, limiting the application of this technology in the downstream market.

Method used

The multi-view neural implicit surface reconstruction method based on dynamic sampling range is adopted. By constructing a three-dimensional uniform grid, generating a dynamic sampling range, and performing layered volume sampling in the core sampling area, sampling points are obtained, and the symbol distance function value and color value are predicted through the geometric network. Finally, the geometric surface model is extracted using the Marching Cube algorithm.

Benefits of technology

It reduces sampling in invalid space, improves the utilization efficiency of light and sampling points, reduces the weight deviation during rendering, and can obtain a three-dimensional geometric grid with smooth surface and rich details in a short time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119722908B_ABST
    Figure CN119722908B_ABST
Patent Text Reader

Abstract

The invention discloses a multi-view neural implicit surface reconstruction method based on dynamic sampling range, which relates to the field of computer technology, including S1, constructing a uniform grid; S2, selecting the position of voxels in the grid and inputting it into a geometric network, calculating the opacity value and occupancy information, and marking whether the corresponding voxels are occupied by an object; S3, generating N rays; S4, analyzing the core sampling area; S5, obtaining sampling points through layered volume sampling; S6, predicting the signed distance function value and the color value; S7, analyzing the pixel color; S8, jointly supervising the prediction information and the pixel color; S9, whether the number of iterations is reached, if yes, proceed to S10; otherwise, return to S2; S10, extracting a geometric surface model; generating a gradually reduced dynamic sampling range based on the occupied grid, reducing the sampling in the invalid space, and improving the utilization efficiency of the light and sampling points. By using the layered volume sampling strategy within the dynamic sampling range, the sampling points are concentrated near the surface of the object, and the sampling deviation is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a multi-view neural implicit surface reconstruction method based on dynamic sampling range. Background Art

[0002] With the deep integration of computer vision and computer graphics with deep learning technology, the reconstruction technology of restoring three-dimensional scene structure from two-dimensional images has become a key technology in frontier fields such as virtual reality, augmented reality, and digital twins. In particular, multi-view neural implicit surface reconstruction technology has attracted widespread attention because of its ability to finely model the surface of objects. This technology uses a neural network composed of multi-layer perceptrons to implicitly represent the geometric properties of objects, and then achieves differentiable reconstruction of the surface of objects through inverse rendering pipeline and volume rendering technology. However, volume rendering calculations rely on the accumulation of colors of multiple sampling points in space, resulting in inevitable rendering deviations, thereby reducing the quality of reconstruction. At the same time, there are a large number of empty areas without foreground objects in real scenes, and invalid sampling in these areas will also affect the quality and efficiency of reconstruction. Due to the low sampling efficiency mentioned above, the existing neural implicit surface reconstruction technology requires a long time of training to obtain a high-quality geometric surface model of the object, or requires the input of additional information to improve the quality of reconstruction, which limits the application of this technology in the downstream market. Therefore, how to improve the effectiveness of sampling points to improve the accuracy and efficiency of reconstructed objects is an urgent problem to be solved in the current field of multi-view neural implicit reconstruction technology. Summary of the invention

[0003] The purpose of the present invention is to design a multi-view neural implicit surface reconstruction method based on dynamic sampling range in order to solve the above problems.

[0004] The present invention achieves the above-mentioned purpose through the following technical solutions:

[0005] Multi-view neural implicit surface reconstruction method based on dynamic sampling range, including:

[0006] S1. Construct a three-dimensional uniform grid, where each voxel in the grid caches its own binary occupancy information and opacity value;

[0007] S2, select any spatial position of the voxel in the grid and input it into the geometric network, calculate the opacity value, calculate the occupancy information according to the threshold, and mark whether the corresponding voxel is occupied by the object;

[0008] S3, select N pixels from N images, and generate N rays along the direction of the selected pixels into the grid;

[0009] S4, analyzing the core sampling area of ​​each ray according to the occupancy information;

[0010] S5. Perform stratified volume sampling in the core sampling area to obtain sampling points;

[0011] S6, predicting the signed distance function value SDF and color value of the sampling point;

[0012] S7, integrating the prediction information of all sampling points along the entire light ray, and analyzing the pixel colors of the corresponding viewing angle image;

[0013] S8, using color loss 、Chenghan loss , Mask loss and sampling loss Jointly supervise prediction information and pixel color;

[0014] S9, determine whether the preset number of iterations has been reached, if so, proceed to S10; otherwise, return to S2;

[0015] S10. The Marching Cube algorithm is used to extract the geometric surface model of the object from the zero level of the signed distance function value SDF.

[0016] The beneficial effects of the present invention are as follows: based on the occupation grid generation, a gradually reduced dynamic sampling range is generated, which reduces the sampling in the invalid space and improves the utilization efficiency of light and sampling points. By using a layered volume sampling strategy within the dynamic sampling range, the sampling points are further concentrated near the surface of the object, reducing the sampling deviation. The introduced sampling point regularization constraint reduces the weight deviation during rendering. Using only a two-dimensional image as input, the method of multi-view neural implicit surface reconstruction based on the dynamic sampling range of the present invention can obtain a three-dimensional geometric mesh with a smooth surface and rich details in a relatively short time. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a schematic diagram of the overall structure of the present invention;

[0018] Figure 2 It is a schematic diagram of the process of grid construction and occupancy information marking of the present invention;

[0019] Figure 3 It is a schematic diagram of a process of generating a sampling range according to occupancy information of the present invention;

[0020] Figure 4 It is a schematic diagram of generating sampling points in the core sampling area of ​​the present invention. DETAILED DESCRIPTION

[0021] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0022] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0023] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.

[0024] In the description of the present invention, it should be understood that the terms "upper", "lower", "inside", "outside", "left", "right", etc. indicate directions or positional relationships based on the directions or positional relationships shown in the accompanying drawings, or are directions or positional relationships in which the product of the invention is usually placed when in use, or are directions or positional relationships commonly understood by those skilled in the art. These directions or positional relationships are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore should not be understood as a limitation on the present invention.

[0025] Furthermore, the terms “first”, “second”, etc. are merely used for distinguishing descriptions and should not be understood as indicating or implying relative importance.

[0026] In the description of the present invention, it is also necessary to explain that, unless otherwise clearly specified and limited, the terms such as "setting" and "connection" should be understood in a broad sense. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be the internal communication of two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0027] The specific implementation modes of the present invention are described in detail below in conjunction with the accompanying drawings.

[0028] like Figure 1 As shown, a multi-view neural implicit surface reconstruction method based on dynamic sampling range includes:

[0029] S1. Construct a three-dimensional uniform grid, where each voxel in the grid caches its own binary occupancy information and opacity value.

[0030] S2, select any spatial position of the voxel in the grid and input it into the geometric network, calculate the opacity value, calculate the occupancy information according to the threshold, and mark whether the corresponding voxel is occupied by an object; Figure 2 As shown, specifically including:

[0031] S201, select any spatial position of voxels in the grid Input the geometric network and predict the signed distance function value SDF;

[0032] S202, any spatial position The signed distance function value SDF is converted into opacity , expressed as:

[0033] ;

[0034] in, Indicates location The signed distance function value SDF, Indicates processing through the sigmoid function;

[0035] S203, update the opacity value of the current voxel cache , expressed as:

[0036] ;

[0037] in, is the attenuation coefficient, which is set to 0.95 in this embodiment;

[0038] S204. Utilize the updated opacity Determine the occupancy of the current voxel, expressed as:

[0039] ;

[0040] When the cached opacity value is greater than the set threshold , marks the current voxel as "occupied" when , otherwise marks it as "unoccupied", In this embodiment, it is set to 0.001.

[0041] S3, select N pixels from N images, and generate N rays along the direction of the selected pixels toward the grid. Specifically, for K images of a scene, randomly select N pixels from N images in batches. From the camera center position corresponding to the selected image Start along the selected pixel direction Generates N rays into the grid.

[0042] S4, generating evenly spaced sampling points on each ray according to the set step value, retaining only the sampling points falling within the voxels occupied by the object according to the occupancy information, and calculating and generating the core sampling area of ​​the current ray according to the filtered sampling point positions; Figure 3 As shown, specifically including:

[0043] S401: Set the step value of dense sampling points , expressed as:

[0044] ;

[0045] in, To occupy the side length of the grid, in this embodiment, it is set to 2. is the grid resolution, which is set to 128 in this embodiment. represents the scaling factor, which is set to 8 in this embodiment;

[0046] S402: Generate evenly spaced sampling points on each ray according to the step value, and retain the sampling points that fall within the voxels marked as "occupied";

[0047] S403: According to the filtered sampling point positions, the core sampling area of ​​the current light is calculated and generated, which is expressed as:

[0048] ;

[0049] in, , are the starting and ending positions of the generated sampling area, near and far represent the nearest and farthest values ​​of the intersection of the ray and the grid, respectively. and are the first and last sampling points after filtering, respectively. In order to improve robustness, when generating the sampling area, and Average expansion interval , while ensuring that the generated sampling area falls within the range between near and far.

[0050] S5. Perform stratified volume sampling in the core sampling area to obtain sampling points, such as Figure 4 As shown in Figure 2, layered volume sampling includes a coarse sampling solution phase and a fine sampling phase, specifically:

[0051] Coarse sampling stage: at the starting position of the sampling area and end position Generate sparse and uniform Sampling points ;

[0052] Fine sampling stage: calculate all sampling points Weight After normalization, the probability density function is generated, and the inverse transformation sampling is performed according to the cumulative distribution function to obtain Fine sampling points , repeat the fine sampling stage h times, the total number of sampling points on each ray is: , in this embodiment =64, =16, h=4.

[0053] S6, the sampling point position and direction vector are hash coded and spherical harmonic coded respectively, and the coded data are input into the geometric network and color network composed of the multi-layer perceptron to predict the signed distance function value SDF and color value of the sampling point; specifically including:

[0054] S601, analyzing the position coordinates of the sampling points in three-dimensional space, expressed as:

[0055] ;

[0056] in, is the distance between the sampling point and the camera origin The distance is the light direction;

[0057] S602: The position coordinates of the sampling points After obtaining a 32-dimensional vector through hash grid encoding, together with the position coordinates of the sampling point Input them into the geometric network together to predict a 16-dimensional geometric feature vector, where the first dimension is the signed distance function value SDF;

[0058] S603, analyzing the gradient of the signed distance function value SDF as the normal vector of the sampling point;

[0059] S604: Obtain a 16-dimensional vector from the direction vector of the light through spherical harmonic encoding together with the geometric feature vector, normal vector, and position coordinates of the sampling point. Input them together into the color network to predict the color value.

[0060] S7. The predicted information of all sampling points along the entire ray is integrated through a volume rendering scheme to obtain the pixel color of the corresponding viewing angle image. The volume rendering is expressed as:

[0061] ,in ;

[0062] in, is the cumulative transmittance, which indicates the proportion of light reaching the camera, is the number of sampling points of the light, and is the predicted color and opacity value of the i-th sampling point, is the predicted color value, is the opacity converted according to the predicted signed distance function value SDF.

[0063] S8, using color loss 、Chenghan loss , Mask loss and sampling loss Jointly supervise prediction information and pixel color;

[0064] Joint supervision is expressed as: ;

[0065] Color Loss The predicted rendering color of the current light k is constrained in the form of L1 norm The actual pixel color on the image The difference between is expressed as: , where m is the batch size;

[0066] Chenghan Loss It is expressed as: ,in Indicates sampling point The signed distance function value SDF gradient of , where n is the number of sampling points and m is the batch size;

[0067] Mask loss It is expressed as: ,in is the binary cross entropy, For the The background mask value of the ray, is the first value calculated after network prediction The transparency of the ray, which is the sum of the weights of the sampling points;

[0068] Sampling loss It is expressed as: ;

[0069] in, , , Chenghan loss , Mask loss , sampling loss In this embodiment, and All are set to 0.01. It starts from 0 and gradually increases to 0.001 with the number of iterations.

[0070] S9, determine whether the preset number of iterations has been reached, if so, proceed to S10; otherwise, return to S2; the preset number of iterations is 20,000 times.

[0071] S10. The Marching Cube algorithm is used to extract the geometric surface model of the object from the zero level of the signed distance function value SDF.

[0072] The network training of the embodiment is performed on a single GeForce RTX3090 GPU. The network structure specifically includes the following components: a 16-level multi-resolution hash grid, each level has 2 feature points, and the resolution ranges from 2 5 Gradually expanded to 2 16 ; A geometric network consisting of one multi-layer perceptron has a hidden layer dimension of 64; an appearance network consisting of two multi-layer perceptrons also has a hidden layer dimension of 64.

[0073] The present invention generates a gradually reduced dynamic sampling range based on the occupation grid, reduces sampling in invalid space, and improves the utilization efficiency of light and sampling points. By using a layered volume sampling strategy within the dynamic sampling range, the sampling points are further concentrated near the surface of the object, reducing the sampling deviation. The introduced sampling point regularization constraint reduces the weight deviation during rendering. Using only two-dimensional images as input, the method of multi-view neural implicit surface reconstruction based on dynamic sampling range of the present invention can obtain a three-dimensional geometric mesh with smooth surface and rich details in a relatively short time.

[0074] The technical solution of the present invention is not limited to the above-mentioned specific embodiments. All technical variations made according to the technical solution of the present invention fall within the protection scope of the present invention.

Claims

1. A multi-view neural implicit surface reconstruction method based on dynamic sampling range, characterized in that: include: S1. Construct a three-dimensional uniform grid, where each voxel in the grid caches its own binary occupancy information and opacity value; S2, select any spatial position of the voxel in the grid and input it into the geometric network, calculate the opacity value, calculate the occupancy information according to the threshold, and mark whether the corresponding voxel is occupied by the object; S3, select N pixels from N images, and generate N rays along the direction of the selected pixels into the grid; S4, analyzing the core sampling area of ​​each ray according to the occupancy information; generating evenly spaced sampling points on each ray according to the set step value, retaining only the sampling points falling within the voxels occupied by the object according to the occupancy information, and calculating and generating the core sampling area of ​​the current ray according to the filtered sampling point positions; specifically including: S401: Set the step value of dense sampling points , expressed as: ; Among them, d is the side length of the grid, Gr is the resolution of the grid, and k represents the scaling factor; S402: Generate evenly spaced sampling points on each ray according to the step value, and retain the sampling points that fall within the voxels marked as "occupied"; S403: According to the filtered sampling point positions, the core sampling area of ​​the current light is calculated and generated, which is expressed as: ; in, , are the starting and ending positions of the generated sampling area, near and far represent the nearest and farthest values ​​of the intersection of the ray and the grid, respectively. and They are the first and last sampling points after filtering respectively; S5. Perform stratified volume sampling in the core sampling area to obtain sampling points; S6, predicting the signed distance function value SDF and color value of the sampling point; specifically: performing hash coding and spherical harmonic coding on the position and direction vector of the sampling point, respectively, and inputting the coded data into the geometric network and color network composed of the multi-layer perceptron, and predicting the signed distance function value SDF and color value of the sampling point; S7, integrating the prediction information of all sampling points of the entire light ray, and analyzing the pixel color of the corresponding view image; specifically: integrating the prediction information of all sampling points of the entire light ray through a volume rendering scheme to obtain the pixel color of the corresponding view image, and the volume rendering is expressed as: ,in ; in, is the cumulative transmittance, which indicates the proportion of light reaching the camera, is the number of sampling points of the light, is the predicted color value, is the opacity converted according to the predicted signed distance function value SDF; S8, using color loss 、Chenghan loss , Mask loss and sampling loss Jointly supervise prediction information and pixel color; S9, determine whether the preset number of iterations has been reached, if so, proceed to S10; otherwise, return to S2; S10. The Marching Cube algorithm is used to extract the geometric surface model of the object from the zero level of the signed distance function value SDF.

2. The multi-view neural implicit surface reconstruction method based on dynamic sampling range according to claim 1 is characterized in that: Included in S2: S201, select any spatial position of voxels in the grid Input the geometric network and predict the signed distance function value SDF; S202, any spatial position The signed distance function value SDF is converted into opacity , expressed as: ; in, Indicates location The signed distance function value SDF, Indicates processing through the sigmoid function; S203, update the opacity value of the current voxel cache , expressed as: ; Where, μ is the attenuation coefficient; S204. Utilize the updated opacity Determine the current voxel occupancy, expressed as: ; When the cached opacity value is greater than the set threshold The current voxel is marked as "occupied" when , otherwise it is marked as "unoccupied".

3. The multi-view neural implicit surface reconstruction method based on dynamic sampling range according to claim 1, characterized in that: In S5, stratified volume sampling includes coarse sampling and fine sampling, specifically: Coarse sampling stage: at the starting position of the sampling area and end position Generate sparse and uniform Sampling points ; Fine sampling stage: calculate all sampling points Weight After normalization, the probability density function is generated, and the inverse transformation sampling is performed according to the cumulative distribution function to obtain Fine sampling points , repeat the fine sampling stage h times, the total number of sampling points on each ray is: .

4. The multi-view neural implicit surface reconstruction method based on dynamic sampling range according to claim 1, characterized in that: Specifically included in S6: S601, analyzing the position coordinates of the sampling points in three-dimensional space, expressed as: ; in, is the distance between the sampling point and the camera origin The distance is the pixel direction; S602: The position coordinates of the sampling points After obtaining a 32-dimensional vector through hash grid encoding, together with the position coordinates of the sampling point Input them into the geometric network together to predict a 16-dimensional geometric feature vector, where the first dimension is the signed distance function value SDF; S603, analyzing the gradient of the signed distance function value SDF as the normal vector of the sampling point; S604: Obtain a 16-dimensional vector from the direction vector of the light through spherical harmonic encoding, together with the geometric feature vector, normal vector, and position coordinates of the sampling point. Input them together into the color network to predict the color value.

5. The multi-view neural implicit surface reconstruction method based on dynamic sampling range according to claim 1, characterized in that: In S8, joint supervision is expressed as: ; Color Loss The predicted rendering color of the current light k is constrained in the form of L1 norm The actual pixel color on the image The difference between is expressed as: , where m is the batch size; Chenghan Loss It is expressed as: ,in Indicates sampling point The signed distance function value SDF gradient of , where n is the number of sampling points and m is the batch size; Mask loss It is expressed as: ,in is the binary cross entropy, is the background mask value of the kth ray, is the transparency of the kth ray calculated after network prediction, that is, the sum of the weights of the sampling points; Sampling loss It is expressed as: ; Among them, λ1, λ2, and λ3 are the process function losses respectively. , Mask loss , sampling loss The weight coefficient of .

Citation Information

Patent Citations

  • Monocular hidden nerve mapping method based on multi-modal depth estimation guidance

    CN117765187A

  • Multi-view nerve implicit surface reconstruction method based on prior driving

    CN118037989A