Refractive index self-adaptive transparent object reconstruction method
Through the refractive index adaptation method, SAM, multi-view geometry and DMTet algorithms are used to perform mask extraction, pose recovery and surface reconstruction of transparent objects, and neural radiation field training is combined with ray tracing, which solves the problem of inverse rendering of transparent objects in the existing technology, and achieves efficient reconstruction of transparent objects and refractive index optimization.
Patent Information
- Application Number
- CN202510181752.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-03
AI Technical Summary
The prior art is difficult to effectively reverse render transparent objects, especially due to the non-Lambertian nature and strong refractive and reflective properties of transparent objects, which lead to the failure of common algorithms, and most methods require specific equipment and prior information.
The transparent object reconstruction method with adaptive refractive index is adopted, and the mask information is extracted through SAM model, the multi-view geometry algorithm is used to restore the camera position, and the DMTet algorithm is used to perform surface reconstruction. The neural radiation field is trained and rendered by ray tracing. During the radiation field training process, the refractive index of transparent objects is optimized through iterative search.
It realizes transparent object reconstruction that ensures the accuracy of radiation field training without the need for accurate refractive index values, reduces training and rendering time, and supports a variety of downstream tasks such as object editing and refractive index editing.
Smart Images

Figure CN120088406A_ABST
Abstract
Description
Background Art
[0001] Transparent objects are extremely common in daily life, such as glass objects, plastics, or water bodies, etc., and play an important role in our daily life. Rendering scenes containing transparent objects is of great significance in fields such as design, animation, and scientific research. To achieve the rendering of transparent object scenes, it is first necessary to perform inverse rendering on the scene to obtain information such as the geometry and radiation of the scene, and then use the forward rendering method to synthesize images from new perspectives. However, the non-Lambertian characteristics of transparent objects make the inverse rendering of transparent objects a difficult point. The commonly used three-dimensional reconstruction and radiation field reconstruction algorithms in computer vision are designed based on the assumption of the straight-line propagation of light, and the color of transparent objects is completely determined by the transmitted or reflected ambient light, resulting in the failure of these common algorithms. In addition, due to the strong refraction and reflection properties of transparent objects, general detection methods such as structured light are also difficult to play a role in such scenes.
[0002] Currently, many scholars have carried out research on the inverse rendering and forward rendering tasks of transparent objects. Inverse rendering usually includes two parts: three-dimensional reconstruction in terms of geometry and radiation field reconstruction in terms of radiation.
[0003] For scene reconstruction, the most common and convenient method is to recover the geometry, material, and lighting information of the scene from several groups of photos. Although the reconstruction algorithms for general scenes are relatively mature, for transparent objects, the influence of refraction will render many existing algorithms ineffective. Refraction causes the surface color of transparent objects to be related to the viewing angle, resulting in failed feature matching between views and causing reconstruction failure. One solution is to apply substances such as paint on the surface of transparent objects to ensure accurate feature matching by avoiding refraction. Alternatively, after immersing the lens object in a liquid, the Archimedes principle or tomography can be used for reconstruction. However, such methods may damage the transparent objects, limiting the scope of application. Non-contact algorithms can avoid causing damage, but they often require relatively complex equipment. Some algorithms directly use a scanning-like method for reconstruction without relying on the refraction characteristics of transparent objects, such as special line laser plane mirrors, lasers, and scanners in the thermal infrared or ultraviolet bands. There are also many algorithms that encode the scene and reconstruct transparent objects based on the encoded values that have been distorted after refraction and reflection. There is a method that combines a fringe screen for position encoding and a multi-layer perceptron for transparent object reconstruction. Many algorithms reconstruct based on the normal consistency of the same point on the object surface, such as the measurement algorithm based on a moving display screen or an improved triangulation reconstruction algorithm, which reconstructs multiple corresponding three-dimensional coordinate points at each pixel point to fit the possible refraction path. The environmental matting technique solves the distortion of the background by transparent objects, enabling it to be composited onto different backgrounds. However, the above algorithms have complex processes and all require specific equipment, such as specific display screens or turntables to assist the algorithms, and most methods require a known refractive index, have a large number of assumptions, and high computational complexity.
[0004] The reconstruction of the radiation field is mainly divided into three types according to the degree of recovery from low to high: view-independent, view-dependent radiation field reconstruction, and illumination-material reconstruction. The first type of algorithm faithfully restores the radiance at each position of the scene at the time of photographing. The view-dependent radiation field reconstruction takes into account the radiance changes caused by the viewing angle change and obtains a new view image through the rendering equation. The illumination-material reconstruction decouples the radiation field and reconstructs the illumination and the material of the object respectively. This algorithm can achieve the editing of illumination and material and has a more flexible application scenario.
[0005] In recent years, methods based on differentiable rendering have developed rapidly. Currently, the most commonly used algorithm is the Neural Radiance Field (NeRF) series of methods. In this method, the entire scene is modeled as a radiance field. The color and volume density at a certain position in the radiance field depend on the position and the observation perspective to achieve an anisotropic rendering effect. Then, forward rendering is performed according to the radiative transfer equation to obtain the rendered image. The parameters of the neural network are updated according to the gradients calculated from the difference between the rendered image and the real image. On the one hand, such methods only require visible light images and camera poses for learning.
[0006] To accurately represent the geometric shape of the object surface, some methods use Signed Distance Function (SDF) to store the scene, fit the SDF at each position using Neural Signed Distance Field (NeuS), and deduce the correspondence between the SDF value at a certain position in space and the volume density in NeRF according to the rendering equation, so as to use the same radiative transfer equation as NeRF for rendering. There are many studies on inverse rendering based on NeuS. Most of these methods are similar to the inverse rendering process based on NeRF, that is, neural networks are used to estimate SDF, material, indirect illumination, and direct illumination. The only difference is that SDF can provide more accurate normal vectors and more direct occlusion information, which is convenient for obtaining more accurate normal vectors and direct light.
[0007] The above algorithms are also based on the straight-line propagation assumption of light. Although some algorithms consider the reflection effect, these algorithms are not designed for the refraction characteristics of transparent objects and will have serious blurring and artifacts in the reconstruction process of transparent objects. For the radiance field reconstruction of transparent object scenes, some methods introduce the eikonal equation to optimize the refractive index field on the light propagation path, and some methods propose a two-stage method for three-dimensional reconstruction of nested transparent objects. In the first stage, the environmental radiance field and the outer surface are reconstructed, and the refracted light is approximated using a Multi-Layer Perceptron (MLP). In the second stage, ray tracing is used to reconstruct the inner surface. However, this method requires manual specification of the transparent parts in the three-dimensional scene and prior information of the refractive index. Neural Refractive Radiance Field (NeRRF) decomposes the rendering problem into two parts: geometric reconstruction and ray tracing rendering. First, after three-dimensional reconstruction of the object using the Deep Marching Tetrahedra (DMTet) algorithm combined with a mask, ray tracing is performed according to the refractive index, and sampling and training of the radiance field are carried out on the modified light path. However, this algorithm requires ray tracing for each outgoing ray, and the training and new view synthesis speeds are extremely slow. Moreover, in actual situations, it is difficult to obtain the accurate refractive index of transparent objects, which brings difficulties to accurate ray tracing and limits the use of this algorithm. Summary of the Invention
[0008] To solve the above problems, the present application provides a method for reconstructing a transparent object with refractive index self - adaptation.
[0009] The technical solution of the present invention is as follows:
[0010] A method for reconstructing a transparent object with refractive index self - adaptation, comprising the following steps:
[0011] Based on SAM, perform mask extraction on the captured image to obtain the mask information of the transparent object;
[0012] Based on the multi - view geometry algorithm, perform sparse reconstruction on the captured image to restore the camera pose corresponding to the image;
[0013] After obtaining the camera pose and the mask, use the DMTet algorithm to perform surface reconstruction on the transparent object to generate the surface shape of the transparent object;
[0014] After obtaining the surface shape of the transparent object, combine ray tracing for the training and rendering of the neural radiance field;
[0015] During the process of radiance field training, optimize the refractive index of the transparent object through iterative search.
[0016] Preferably, the step of performing mask extraction on the captured image based on SAM to obtain the mask information of the transparent object includes:
[0017] Use the pre - trained SAM model to extract all masks in the captured image;
[0018] Manually select the mask required for any one picture, and automatically screen out several masks that are most matched with the selected mask through position information, color histogram, and mask size;
[0019] Take the intersection of the screened masks, extract the largest connected region by area, and eliminate the burrs in the mask through opening operation.
[0020] Preferably, the step of performing sparse reconstruction on the captured image based on multi - view geometry to restore the camera pose corresponding to the image includes:
[0021] Use the multi - view stereo vision algorithm to perform sparse reconstruction on the input image to generate a sparse point cloud of the scene;
[0022] Optimize the camera pose and the position of the sparse point cloud through bundle adjustment.
[0023] Preferably, the step of using the DMTet algorithm to perform surface reconstruction on the transparent object to generate the surface shape of the transparent object includes:
[0024] Generate an initial mesh according to the mask and the camera pose;
[0025] Use a multi-layer perceptron (MLP) to process the vertex data of the initial mesh, generating the signed distance field (SDF) value and deformation value for each vertex;
[0026] Update the vertex positions of the initial mesh according to the SDF values and deformation values;
[0027] Use the DMTet algorithm to extract an explicit mesh from the updated mesh.
[0028] Preferably, after obtaining the surface shape of the transparent object, the steps of training and rendering the neural radiance field in combination with ray tracing include:
[0029] For each ray emitted from the camera, determine whether the ray intersects with the transparent object;
[0030] If the ray does not intersect with the transparent object, directly use the neural radiance field (NeRF) to calculate the color value of the ray without considering reflection and refraction;
[0031] If the ray intersects with the transparent object, calculate the refraction direction and reflection direction according to the law of refraction, and calculate the radiance ratio of the reflected ray and the radiance ratio of the refracted ray according to the Fresnel law;
[0032] Perform path sampling along the reflection direction and calculate the color of the reflected ray through the neural radiance field (NeRF);
[0033] Perform path sampling along the refraction direction and calculate the color of the refracted ray through the neural radiance field (NeRF);
[0034] Synthesize the final pixel color of the ray according to the color of the reflected ray, the radiance ratio of the reflected ray, the color of the refracted ray, and the radiance ratio of the refracted ray.
[0035] Preferably, the steps of directly using the neural radiance field (NeRF) to calculate the color value of the ray include:
[0036] Use the neural radiance field (NeRF) to represent the radiance field of the scene;
[0037] Divide the propagation path of the ray from the camera position o to the intersection point of the ray with the surface of the object in the scene into N equal parts, randomly coarsely sample a sampling point in each interval to obtain the coarsely sampled color value and volume density information at each sampling point;
[0038] According to the volume density information at each sampling point, perform denser fine sampling near the object to obtain the fine sampled color value;
[0039] Accumulate all the coarsely sampled color values and fine sampled color values to obtain the cumulative color value of the ray in the scene.
[0040] Preferably, during the radiation field training, the steps of optimizing the refractive index of the transparent object by iterative search include:
[0041] After the background radiation field training is completed, search within a preset refractive index range;
[0042] For each candidate refractive index, calculate the mean square error between the foreground rendered image and the real image;
[0043] Select the refractive index that minimizes the mean square error as the optimal refractive index;
[0044] Under the optimal refractive index, jointly train the radiation fields of the foreground and the background.
[0045] The beneficial effects of the present invention are:
[0046] The present invention can realize the automatic search of the refractive index during the radiation field training, and ensure the accuracy of the subsequent radiation field training without the premise of accurate refractive index values; based on the mask-based ray tracing acceleration, it can reduce the training and rendering time by about 1 / 2; perform a complete inverse rendering of the scene to obtain relatively accurate ambient light, transparent object geometry, and refractive index, and can perform various downstream tasks such as object editing, environment editing, and refractive index editing. Description of the Drawings
[0047] Figure 1 It is a schematic diagram of the geometric reconstruction result in the embodiment of the present application;
[0048] Figure 2 It is a schematic diagram of the reconstruction result of the Metashape glass ball scene in the embodiment of the present application;
[0049] Figure 3 It is a schematic diagram of the reconstruction result of the NeuS glass ball scene in the embodiment of the present application;
[0050] Figure 4 It is a schematic diagram of the comparison between the algorithm effect (a) and Instant-NGP (b), NeRRF (c), and the ground truth (d) in the embodiment of the present application;
[0051] Figure 5 It is a schematic diagram of the comparison of the foreground partial rendering effects in the embodiment of the present application;
[0052] Figure 6 It is a schematic diagram of the reconstruction result of the environmental radiation field in the embodiment of the present application;
[0053] Figure 7 It is a schematic diagram of the scene editing rendering result in the embodiment of the present application;
[0054] Figure 8 It is a schematic diagram of the rendering effects under different refractive indices in the embodiment of the present application;
[0055] Figure 9 This is a schematic flowchart of the refractive index adaptive transparent object reconstruction method in the embodiments of the present application. Detailed implementation manners
[0056] First, refer to Figure 9 , the embodiments of the present application provide a refractive index adaptive transparent object reconstruction method, which follows a series of steps, including data preprocessing, dense reconstruction based on DMTet, and radiation field training and refractive index optimization based on ray tracing in four steps.
[0057] Among them, the data preprocessing step includes:
[0058] S1: Extract the mask for the captured image based on the SAM model (SegmentAnythingModel);
[0059] S2: Then perform sparse reconstruction on the captured image based on the multi-view geometry algorithm to restore the camera pose corresponding to the image.
[0060] The dense reconstruction step based on the DMTet algorithm includes:
[0061] S3: After obtaining the camera pose and the mask, use the DMTet algorithm to perform surface reconstruction on the transparent object to generate the surface shape of the transparent object.
[0062] The steps of radiation field training based on ray tracing include:
[0063] S4: After obtaining the surface shape of the transparent object, combine ray tracing to perform training and rendering of the neural radiation field.
[0064] The steps of refractive index optimization include:
[0065] S5: During the radiation field training process, optimize the refractive index of the transparent object through iterative search.
[0066] Among them, the data preprocessing steps specifically include:
[0067] Mask extraction: Use the pre-trained SAM (SegmentAnythingModel) model to extract the object mask in the image, and perform sparse reconstruction through multi-view geometry to restore the camera pose.
[0068] The steps of using the pre-trained SAM (SegmentAnythingModel) model to extract the object mask in the image include:
[0069] The training of DMTet and NeRF requires the object mask and the camera poses of the images as supervision. First, the object mask is extracted based on the SAM model. First, use the pre-trained SAM to obtain all the masks of the captured images, then manually select the mask required for any one image, and then automatically select the top-k matching masks for the remaining images based on position information, color histogram, mask size, etc. and store them. Take the intersection of the above k masks, extract the largest connected region by area, and then eliminate the burrs in the mask through opening operation.
[0070] The recovery of camera poses is completed by the open-source software COLMAP. COLMAP is a commonly used multi-view stereo vision (Structure-from-Motion, SfM) tool, and its workflow can be summarized as: optimizing the camera poses and the positions of the sparse point cloud through bundle adjustment.
[0071] Optimizing the camera poses and the positions of the sparse point cloud through bundle adjustment specifically recovers the camera poses and the sparse point cloud through the following steps:
[0072] Feature extraction and matching: Extract feature points (such as SIFT features) from the captured images and match the feature points in different captured images.
[0073] Sparse reconstruction: Generate a sparse point cloud using the triangulation method through the features in different captured images that are matched, and initially estimate the camera poses.
[0074] Bundle Adjustment: Optimize the camera poses and the positions of the sparse point cloud simultaneously through an optimization algorithm (such as non-linear least squares method) to minimize the reprojection error.
[0075] After obtaining the mask and the camera poses, use the DMTet algorithm to perform surface reconstruction on the transparent object.
[0076] First, generate the initial mesh T according to the mask and the camera poses. The steps of using the DMTet algorithm to perform mesh reconstruction of the surface shape of the object include: DMTet processes the vertex data of the initial mesh T through a multi-layer perceptron (MLP) to generate the signed distance field value (i.e., SDF value) and the deformation value of each vertex, and finally extracts the surface shape of the object.
[0077] Specifically, combined with Figure 1 , use the multi-layer perceptron MLP to process the vertex data on the initial mesh T to generate each vertex v i -related SDF value S i and deformation value Δv i , S i ∈R, Δv i ∈R 3 ; R is the set of real numbers, indicating Si The value range of is real numbers. Specifically, the MLP receives the vertices of the initial mesh T as input. These vertices are the basic geometric elements that make up the mesh, and each vertex has a position in three-dimensional space. Through its internal multi-layer neural network structure, the MLP performs a non-linear transformation on the input vertex information to predict the SDF value S i and the deformation value Δv i of each vertex v i . The SDF value S i represents the distance from the vertex v i to the surface of the transparent object, while the deformation value Δv i represents the displacement of the vertex v i in space.
[0078] Then use the DMTet algorithm to update the initial mesh T. The DMTet algorithm defines the geometry of the object as a signed distance field (SDF) on the mesh T. Among them, each vertex v i on the initial mesh T is characterized by the SDF value S i ∈R and the deformation value Δv i ∈R 3 . Each time the DMTet algorithm updates the mesh, the vertices of the updated mesh will be deformed into v i ′, and v i ′ = v i +Δv i . In order to obtain an explicit Mesh, the marching tetrahedra algorithm is used to extract the surface shape of the object from the updated mesh T.
[0079] The surface shape Mesh generated by the DMTet algorithm provides a geometric basis for subsequent ray tracing and neural radiance field training. Specifically, the surface shape is used to: determine the intersection points of rays and objects; calculate the refraction and reflection directions of rays; provide geometric constraints to ensure the consistency of the radiance field with the object geometry.
[0080] After obtaining the surface shape (Mesh) of the object, ray tracing and neural radiance field training can be performed based on the surface shape.
[0081] In the embodiments of this application, the neural radiance field NeRF is used for the expression and reconstruction of the radiance field.
[0082] The neural radiance field NeRF represents the neural radiance field using a single function. Specifically, NeRF represents the radiance field of the scene as a five-dimensional function, that is, F θ (x, d) → (σ, c), where:
[0083] x = (x, y, z), representing the three-dimensional position of the camera observation point;
[0084] d = (θ, φ) represents the viewing direction; θ is the zenith angle (the angle from the vertically downward direction to the viewing direction); φ is the azimuth angle (the angle from the positive x-axis to the viewing direction).
[0085] σ represents the volume density at the camera observation point.
[0086] c represents the color value corresponding to the viewing direction at the camera observation point, usually represented by the radiance in the three R, G, B bands.
[0087] NeRF renders according to the rendering equation. For a ray r = o + td emitted from the camera position, the color value c(r) of the corresponding pixel is given by the following integral:
[0088]
[0089] Where:
[0090] o represents the starting point of the ray (i.e., the camera position), determined by the camera pose P = {R, t};
[0091] t represents the distance parameter of the ray propagation, determined by stratified sampling;
[0092] d represents the direction vector of the ray, determined by the camera pose and the pixel position;
[0093] is the cumulative transmittance of the ray r from the starting point t n to the current point t; σ(r(t)) represents the volume density at the position r(t) on the ray path;
[0094] c(r(t), d) represents the color along the viewing direction d at the position r(t) on the ray path; σ(r(s) represents the volume density at the position r(s) on the ray path;
[0095] t n and t f respectively represent the parameter values when the ray enters and exits the transparent object, used to define the upper and lower limits of the integral.
[0096] In order to calculate this integral c(r), it is necessary to discretize the rendering equation by stratified sampling.
[0097] The propagation path of the ray from the camera position o to the intersection point of the ray and the surface of the transparent object in the scene is divided into N equal parts.
[0098] The stratified sampling formula is:
[0099]
[0100] Here, t iis a point randomly sampled uniformly within the i-th interval, t n and t f represent the parameter values of the light entering and leaving the object respectively, and U represents a uniform distribution.
[0101] Next, a point is randomly sampled uniformly in each interval, and the color integral along the light path is calculated. The integral formula is:
[0102]
[0103] where σ i is the volume density at the i-th sampling point; δ j = t i - t i-1 , which is the length of the i-th interval; c i is the color at the i-th sampling point. represents the cumulative transmittance from the light entering the object to the i-th sampling point; σ j is the volume density at the j-th sampling point; σ j is the volume density at the j-th sampling point, and δ j is the length of the j-th interval. To enable the network to learn more high-frequency information, positional encoding is used to map the original input to a high-dimensional space. In the original NeRF, a trigonometric encoding method was adopted. Barron et al. proposed Instant-NGP, a positional encoding method of multi-resolution hashing encoding, which reduces the number of network parameters while improving the expression ability of the scene and improves the rendering speed. In the embodiments of the present application, the same method is also used for positional encoding.
[0104] In the initial training, since the position of the transparent object is unknown and the weights of each position for the color during the light propagation process cannot be determined, a uniform sampling method is adopted, which is called coarse sampling. When the coarse sampling is completed, the positions where the objects exist in the scene have been roughly determined. On this basis, NeRF performs fine sampling based on the volume density of each position on the light path as the weight to improve the detail performance. The final loss includes the color loss of the coarse sampling and the color loss L of the fine sampling:
[0105]
[0106] R represents the set of all sampled rays;
[0107] represents the predicted color obtained by coarse sampling along the ray r;
[0108] represents the predicted color obtained by fine sampling along the ray r;
[0109] c(r) represents the true color of ray r.
[0110] After obtaining the surface shape of the object, ray tracing can be performed pixel by pixel on the transparent object image to modify the light path.
[0111] First, calculate whether each ray intersects with the transparent object. For rays that do not intersect with the object, there is no need to modify the light path, and color rendering can be directly performed. When training the radiation field using a simulation dataset based on Blender, the ambient light can be regarded as infinitely far away, and the color is only related to the observation direction. At this time, the color and volume density of each sampling point in space are represented as F σ,c (d) → (σ, c).
[0112] For rays that intersect with the transparent object, since the transparent object itself is colorless, the color of the transparent object only depends on the color of the outgoing ray. Therefore, the ray will refract according to Snell's law. For a ray r = o + td that intersects with the object i , assuming the normal vector of the surface shape is n, the refractive index of air is approximately 1, and the refractive index of the object is η, the incident direction is d i , the incident angle is θ i in the case of, the cosine value cosθ t of the refraction angle θ t is:
[0113]
[0114] The refraction direction d t is:
[0115] d t = -ηd i +(η(d i ·n) - cos(θ t )n
[0116] The reflection direction d r is:
[0117] d r = 2n(d i ·n) - d i
[0118] Assume that the ray intersects the object at point p, then the reflected ray is r r = p + td r , and the refracted ray is r t = p + td t . When the ray radiance is L, the radiance ratio L r of the reflected ray L r and the radiance ratio L t of the refracted ray L t are given by Fresnel's law:
[0119]
[0120] L r = FL,L t = (1 - F)L
[0121] By calculating the first incidence and the first exit of each ray, calculating the proportion of the reflected radiance at incidence and the ratio of the refracted radiance at exit, finally the color c(r) of this pixel is determined by the color c r reflected into the camera and the color of the ray refracted into the camera c t simultaneously:
[0122] c(r) = Fc r (r r ) + (1 - F)c t (r t )
[0123] The training simultaneously optimizes the color values of the coarse sampling and the fine sampling.
[0124] During the ray tracing process, it is necessary to traverse each ray to find intersections with objects, and this step is a very time-consuming process. In the embodiments of the present application, mask information is added during the ray tracing process. Since only the rays emitted by the foreground pixels will intersect with the objects, other pixels will surely not intersect with the objects. Therefore, during the training and rendering processes, only the pixels with a mask value of 1 are traced.
[0125] In practical applications, it is difficult to obtain the accurate refractive index of a transparent object. Research has shown that the background radiation field is basically not affected by refraction and can provide a reference for the refractive index. Therefore, in the embodiments of the present application, the optimal refractive index is iteratively obtained during the training process to achieve adaptive optimization of the refractive index.
[0126] The adaptive optimization of the refractive index mainly includes three steps: background radiation field training, refractive index search, and complete radiation field training. First, the background radiation field is trained, that is, the colors of the pixels with a mask value of 0 are trained. After several rounds of training, we obtain a basically correct background radiation field. Since the color of a transparent object only depends on the ambient light, we have reason to believe that the correct rendering effect can be obtained only by virtue of the background radiation field and the correct refractive index.
[0127] Subsequently, a search is performed within a possible refractive index range [ηmin, ηmax]. The specific steps include: dividing the range evenly into n equal parts to obtain candidate refractive indices η1, η2, …, ηn; for each candidate refractive index ηk, use the current surface shape (Mesh) and the background radiation field to render the foreground image; calculate the mean square error (MSE) between the rendered foreground image and the real foreground image, and select the refractive index that minimizes the MSE as the current optimal refractive index ηopt.
[0128] The refractive index range [ηmin, ηmax] is usually set according to the physical properties of the transparent object.
[0129] Finally, at the current optimal refractive index η opt perform simultaneous training on the radiation fields of the foreground and the background to train for a few rays that only appear in the foreground.
[0130] Figure 1 Shows the reconstruction effects of the algorithm on different scene datasets. It can be seen from the figure that the surface geometry reconstruction based on DMTet has achieved relatively stable effects on different datasets. At the same time, the geometry reconstruction algorithm based only on Mask has low requirements for input data, and the Mask is not affected by refraction and reflection, improving the practicability of the algorithm.
[0131] Figure 2 Shows the 3D reconstruction effect of Metashape in a real scene. It can be seen that the refractive property of the glass causes the positions of corresponding points in the foreground to change with the view, resulting in the failure of feature matching. Figure 3 Shows the 3D reconstruction effect of NeuS. It can be seen that refraction causes the network to learn the wrong SDF.
[0132] In this embodiment, the effects of the above method in the novel view synthesis task are tested on the test sets of the blender dataset and the real scene dataset, and a comparison is made with the Instant-NGP method.
[0133] Figure 4 Shows the comparison between Instant-NGP, the original NeRRF, the algorithm of this article, and the ground truth. It can be seen that due to refraction, Instant-NGP based on the assumption of straight-line light propagation will cause a large deviation in the radiation field training on the dataset of transparent objects. Specifically, artifacts appear at multiple incorrect positions of the transparent object, which is caused by the network wrongly learning the radiation field in that direction due to the refraction of light. On the contrary, NeRRF and the above method of this embodiment effectively avoid this problem. Figure 5It shows the prospects of the above method in this embodiment, that is, the comparison of the transparent object part with NeRF and the ground truth. It can be seen that the algorithm proposed in this embodiment effectively models the refraction phenomenon and achieves good rendering results. At the same time, it can also be observed that due to some noise in the surface reconstruction process of the DMTet algorithm and the lack of strict geometric supervision, large deviations occur in the ray tracing of some pixels in the foreground, which in turn brings some noise.
[0134] The refractive index adaptive refraction radiation field proposed in this embodiment can still obtain results similar to those of the original refraction radiation field when the initial refractive index is unknown. The estimated refractive index is compared with the true refractive index as shown in Table 1. It can be seen that the refractive index estimation algorithm proposed in this embodiment can obtain a relatively accurate refractive index, and the selected refractive index is the one that maximizes the PSNR. Although the refractive index may deviate slightly, the rendering effect is still guaranteed.
[0135] Table 1 Refractive index estimation results
[0136]
[0137] Table 2 gives a quantitative comparison of the novel view synthesis results of the method in this paper with Instant-NGP, the original NeRRF, and the algorithm NeRFRO designed for the refraction object scene on the synthetic dataset. This comparison includes the complete image and the foreground. It can be seen that the algorithm proposed in this paper has advantages compared with the traditional neural radiation field, and compared with NeRRF, the algorithm proposed in this paper does not require the prior knowledge of the initial refractive index. Among them, CI represents the index calculated using the complete image (CompleteImage), and FI represents the index calculated using the foreground (ForegroundImage) image.
[0138] Table 2 Comparison of different algorithm metrics under the Blender dataset
[0139]
[0140] Table 3 gives a quantitative comparison of the novel view synthesis results of the method in this paper with the current relatively advanced algorithms on the real dataset. This comparison includes the complete image and the foreground, and the top three metrics are marked with gold, silver, and copper. It can be seen that the algorithm proposed in this paper is also comparable in terms of metrics.
[0141] Table 3 Comparison of different algorithm metrics under the real scene data
[0142]
[0143]
[0144] In addition to novel view synthesis, this paper can also perform scene editing, including object replacement and relighting, because it can obtain the environmental radiation field and object geometry. As Figure 6 , the results of ambient light estimation are given. It can be seen that the method in this paper reconstructs the ambient light of the visible light path in the dataset more accurately. The poor reconstruction effect of the upper ambient light is mainly due to the perspective limitation of the dataset, and there is no upward shooting perspective in the dataset.
[0145] Figure 7 The scene editing results are given. This operation can be regarded as either object replacement under the existing ambient light or relighting of the existing object.
[0146] Figure 8 The rendering effects of the Ball dataset at different refractive indices are given. It can be seen that the algorithm reconstructs the complete radiation field more accurately and decouples the radiation field from the refractive index, making it possible to edit the refractive index of the object.
[0147] This paper proposes a method for reconstructing transparent objects with adaptive refractive index. This method includes three main steps: geometric reconstruction, radiation field training, and refractive index optimization, and can achieve novel view synthesis and refractive index estimation without prior knowledge of the refractive index. On the basis of retaining the characteristics of NeRRF, this method realizes the adaptive optimization of the refractive index, and is similar to the current best method in terms of indicators, and can be conveniently applied to many fields. Table 4 shows the characteristic comparison of different algorithms. It can be seen that our algorithm has certain applicability advantages compared with other methods. In the future, a possible improvement method is to optimize the geometric estimation and radiation field estimation simultaneously. In addition, for the background radiation field, a more powerful radiation field representation method, such as Mip-NeRF, Zip-MeRF (Barron et al., 2023), etc., can be used for representation. Finally, ray tracing inside the object can be tried to achieve the reconstruction and rendering of nested objects and colored transparent objects.
[0148] Table 4 Characteristic Comparison of Different Methods
[0149]
[0150] It should be noted that each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.
[0151] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.
[0152] It should also be noted that in this text, the orientation or positional relationship indicated by terms such as "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. This is for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or component referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. In addition, relative terms such as "first" and "second" are used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations, nor can they be construed as indicating or implying relative importance. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements does not include those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the element.
[0153] The technical solutions provided by the present invention have been introduced in detail above. Specific examples are used in this text to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only for helping to understand the present invention, and the content of this specification should not be construed as a limitation to the present invention. At the same time, for those of ordinary skill in the art, according to the present invention, there will be various forms of changes in the specific implementation manners and application scopes. It is not necessary and impossible to enumerate all the implementation manners here, and the obvious changes or variations derived therefrom are still within the protection scope of the present invention.
Claims
1. A method for refractive index adaptive transparent object reconstruction, characterized in that: The following steps are involved: Perform mask extraction on the captured image based on SAM to obtain the mask information of the transparent object; Sparsely reconstruct the captured image based on the multi-view geometry algorithm to restore the camera pose corresponding to the image; After obtaining the camera pose and mask, the DMTet algorithm is used to reconstruct the surface of the transparent object and generate the surface shape of the transparent object; After obtaining the surface shape of the transparent object, the neural radiation field is trained and rendered in combination with ray tracing; During the radiation field training process, the refractive index of the transparent object is optimized through iterative search.
2. The refractive index adaptive transparent object reconstruction method according to claim 1, characterized in that: The steps of performing mask extraction on the captured image based on SAM to obtain mask information of the transparent object include: Use the pre-trained SAM model to extract all masks in the captured image; Manually select the mask required for any image, and automatically select several masks that best match the selected mask based on position information, color histogram, and mask size; The selected masks are intersected and the connected region with the largest area is extracted. The burrs in the masks are eliminated through the opening operation.
3. The method for refractive index adaptive transparent object reconstruction according to claim 1, characterized in that: The steps of sparsely reconstructing the captured image based on multi-view geometry to restore the camera pose corresponding to the image include: Use a multi-view stereo vision algorithm to sparsely reconstruct the input image and generate a sparse point cloud of the scene; Optimize the camera pose and the position of the sparse point cloud through bundle adjustment.
4. The refractive index adaptive transparent object reconstruction method according to claim 1, characterized in that: The DMTet algorithm is used to reconstruct the surface of a transparent object. The steps of generating the surface shape of the transparent object include: Generate an initial mesh based on the mask and camera pose; Use a multi-layer perceptron MLP to process the vertex data of the initial mesh and generate the signed distance field SDF value and deformation value of each vertex; Update the vertex positions of the initial mesh according to the SDF value and deformation value; Extract the explicit mesh from the updated mesh using the DMTet algorithm.
5. The refractive index adaptive transparent object reconstruction method according to claim 1, characterized in that: After obtaining the surface shape of the transparent object, the steps of training and rendering the neural radiation field in combination with ray tracing include: For each ray emitted from the camera, determine whether the ray intersects with a transparent object; If the light does not intersect with a transparent object, the color value of the light is directly calculated using the neural radiation field NeRF without considering reflection and refraction; If the light intersects with a transparent object, the refraction direction and the reflection direction are calculated according to the law of refraction, and the radiance ratio of the reflected light and the refracted light are calculated according to the Fresnel law; Path sampling is performed along the reflection direction, and the color of the reflected light is calculated through the neural radiation field NeRF; Path sampling is performed along the refraction direction, and the color of the refracted light is calculated through the neural radiation field NeRF; The final pixel color of the light is synthesized according to the color of the reflected light, the radiance ratio of the reflected light, the color of the refracted light and the radiance ratio of the refracted light.
6. The refractive index adaptive transparent object reconstruction method according to claim 5, characterized in that: The steps to directly use the neural radiation field NeRF to calculate the color value of the light include: Use Neural Radiance Fields (NeRF) to represent the radiance field of the scene. Divide the propagation path of the light from the camera position o to the intersection point of the light and the surface of the object in the scene into N equal parts, randomly sample a sampling point in each interval, and obtain the coarse sampled color value and volume density information at each sampling point; According to the volume density information at each sampling point, more dense fine sampling is performed near the object to obtain the fine sampling color value; Accumulate all coarse sampling color values and fine sampling color values to get the cumulative color value of the light in the scene.
7. The refractive index adaptive transparent object reconstruction method according to claim 1, characterized in that: During the radiation field training process, the steps of optimizing the refractive index of a transparent object through iterative search include: After the background radiation field training is completed, search is performed within the preset refractive index range; For each candidate refractive index, the mean square error between the foreground rendered image and the true image is calculated; Select the refractive index that minimizes the mean square error as the optimal refractive index; The radiance fields of foreground and background are jointly trained under the optimal refractive index.
Citation Information
Cited By
Transparent object inverse rendering method and device based on three-dimensional Gaussian
CN120411324A