A high-precision stereo reconstruction method for complex scenes

By combining a polarization camera and an MLP network, the problem of inaccurate reconstruction of textureless and specular reflection areas in complex scenes was solved, achieving high-precision 3D reconstruction, especially accurate reconstruction of textureless and specular reflection areas, overcoming the technical shortcomings of multi-view methods.

CN119672208BActive Publication Date: 2026-02-10XIDIAN UNIV HANGZHOU RES INST
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411577097.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-06
Publication Date
2026-02-10
Estimated Expiration
2044-11-06

AI Technical Summary

Technical Problem

Existing 3D reconstruction methods suffer from inaccurate reconstruction of textureless areas, blurred specular reflections, and inaccurate estimation of occlusion and shadows in complex scenes. In particular, neural radiation field-based and multi-view stereo methods do not perform well in reconstructing textureless and specular reflection areas.

Method used

Multi-view and multi-angle polarized images are acquired using a polarization camera. The incident angle and azimuth angle are calculated using Stokes vectors and polarization normal vectors. The scene surface normal vectors are obtained by combining them with an MLP network. The polarization normal vectors are corrected using prior normal vectors. The results are then input into an improved NeuS network for volume rendering to obtain high-precision 3D information.

Benefits of technology

It improves the reconstruction accuracy of textureless and specular reflection areas in complex scenes, can handle occlusion and thin-structured objects, reduces depth scale blur in multi-view mode, and obtains smooth geometry and accurate 3D information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119672208B_ABST
    Figure CN119672208B_ABST
Patent Text Reader

Abstract

The application discloses a high-precision three-dimensional reconstruction method for a complex scene, which comprises the following steps: obtaining the position and orientation of a camera from multi-view images through polarization camera pose estimation; obtaining a scene surface normal vector through an MLP network; correcting an ambiguous polarization normal vector calculated by using a polarization image by using the scene surface normal vector, thereby obtaining a prior normal vector; inputting the prior normal vector into a pre-trained improved NeuS network to provide additional constraints, so as to alleviate the geometric blur problem of a textureless area and a mirror reflection area; and compared with depth prior, the network based on the prior normal vector can generate smooth geometry, can avoid the depth scale blur problem between multiple views, and can obtain more accurate three-dimensional information. The application can also perform high-precision reconstruction on an object with occlusion and an object with a thin structure, and can also perform better geometric reconstruction on a scene with sudden depth changes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of three-dimensional reconstruction technology, specifically relating to a high-precision stereo reconstruction method for complex scenes. Background Technology

[0002] Reconstructing 3D scenes from multiple input images is a crucial and challenging task in many practical applications such as robot navigation, virtual reality, and path planning. Recovering the 3D shape of a real-world scene is a fundamental problem in computer vision. Inverse rendering aims to recover scene parameters from a set of related images, and traditionally, inverse rendering methods rely on multi-view geometry, photometric measurements, and structured lighting for 3D reconstruction. Due to the ill-posed nature of inverse rendering, these methods often require simplifying assumptions about the scene, such as textured surfaces, Lambertian reflectivity, direct lighting, and simple geometry. Methods for generalized scene settings include incorporating scene priors, iterative scene optimization using differentiable rendering, and leveraging different properties of light, such as polarization, time of flight, and spectrum.

[0003] The emergence of neural implicit network representations in recent years has sparked interest in neural inverse rendering. Neural implicit representations use coordinate-based neural networks to represent visual signals such as images, videos, and 3D objects. These representations are powerful because the resolution of the underlying signal is limited only by the network capacity, not by the discretization of the signal. Much of the interest in the visual world stems from neural radiation fields (NeRF), which demonstrate that modeling radiation using implicit representations can achieve high-quality, novel view synthesis. However, NeRF encounters difficulties with textureless or specularly reflective surfaces. The lack of texture leads to fuzzy correspondences, while the presence of specular reflections violates the photoconsistency assumption. Real-world scenes often contain textureless and specularly highlighted areas, resulting in inaccurate reconstructions. Polarized 3D imaging, as an incoherent imaging technique, does not inherently require a modulated light source, is unaffected by propagation paths and coherence errors, is less sensitive to ambient light, offers high spatial resolution, and can perform real-time imaging, making it a promising practical application for 3D imaging of complex targets. However, polarization 3D imaging technology suffers from singularity issues when interpreting the polarization field of complex target surfaces, multivalued issues when using the polarization field to solve the normal vector field, and the reconstruction process is sensitive to polarization field errors. This is one of the main reasons why polarization 3D imaging technology cannot be widely applied to complex target scenes.

[0004] Currently, common methods for 3D reconstruction of large scenes include monocular depth estimation and neural radiation field-based 3D reconstruction methods that utilize images from multiple perspectives. However, monocular depth estimation suffers from inaccurate estimation in complex scenes and in areas with occlusion and shadows; while neural implicit 3D reconstruction methods using multiple perspectives suffer from inaccurate reconstruction of geometric surfaces in areas without texture or specular reflection. Summary of the Invention

[0005] To address the aforementioned problems in the existing technology, this invention provides a high-precision stereo reconstruction method for complex scenes. The technical problem to be solved by this invention is achieved through the following technical solution:

[0006] This invention provides a high-precision stereo reconstruction method for complex scenes, the method comprising:

[0007] Use a polarization camera to acquire multi-view and multi-angle polarization images of the target scene;

[0008] Multi-angle polarization images are processed to obtain polarization information and the degree of polarization of the target scene surface;

[0009] Based on the polarization information and the degree of polarization of the target scene surface, the azimuth and incident angle of the incident light on the target scene surface are obtained;

[0010] Based on the azimuth and incident angle, the polarization normal vector is obtained;

[0011] Based on the multi-view images, the spatial position and view orientation of the polarization camera are obtained; based on the spatial position and view orientation, the symbolic distance function (SDF) and color value of the target scene are obtained using an MLP network; based on the symbolic distance function (SDF), the scene surface normal vector is obtained; the scene surface normal vector is used as a guiding vector to correct the polarization normal vector, resulting in a disambiguated surface normal vector.

[0012] The disambiguated surface normal vector is used as the prior normal vector. A pre-trained improved NeuS network is used to perform volume rendering on the position and orientation of the polarization camera, the signed distance function SDF, the color value and the prior normal vector to obtain the predicted color value and normal vector. Based on the difference between the predicted normal vector and the prior normal vector, as well as the difference between the predicted color value and the color value of the target scene, the three-dimensional information of the target scene is obtained.

[0013] In one embodiment of the present invention, a polarization camera is used to acquire multi-view images and multi-angle polarization images of a target scene, including:

[0014] The target scene was captured using a FLIR polarization camera to obtain video frames;

[0015] Extract multi-view images from video frames and obtain multi-angle image data of the scene. Combine this with the polarizer on the polarization camera to obtain multi-angle polarized images I0 and I0. 45 I 90 I 135 .

[0016] In one embodiment of the present invention, processing a multi-angle polarization image to obtain polarization information and the degree of polarization of the target scene surface includes:

[0017] The multi-angle polarization image is processed to obtain polarization information; the polarization information is composed of Stokes vectors; the expression of the Stokes vectors is as follows:

[0018]

[0019] Among them, I 0° I represents the light intensity passing through a 0° polarizer. 90° I represents the light intensity passing through a 90° polarizer. +45° I represents the light intensity passing through a polarizer at a 45° angle to the positive horizontal direction. -45° I represents the light intensity obtained through a linear polarizer at a 135° angle to the horizontal. RH I represents the right-hand circularly polarized light component. LH This represents the left-hand circularly polarized light component.

[0020] Based on the Stokes vector, the degree of polarization of the target scene surface is obtained using the first formula; the expression of the first formula is as follows:

[0021]

[0022] In one embodiment of the present invention, based on the polarization information and the degree of polarization of the target scene surface, the azimuth angle and incident angle of the incident light on the target scene surface are obtained, including:

[0023] Based on the polarization information and the degree of polarization of the target scene surface, the azimuth angle of the incident light on the target scene surface is obtained using the second formula.

[0024] The second formula is as follows:

[0025]

[0026] in, S0 and S1 represent the Stokes vectors;

[0027] The angle of incidence is obtained based on the relationship between the degree of polarization and the angle of incidence.

[0028] The relationship between the degree of polarization and the angle of incidence is as follows:

[0029] For specular reflection:

[0030]

[0031] For diffuse reflection:

[0032]

[0033] Where P represents the degree of polarization, n represents the refractive index of the object's surface, and θ represents the angle of incidence.

[0034] In one embodiment of the present invention, the polarization normal vector is obtained based on the azimuth angle and the incident angle, including:

[0035] Based on the azimuth and incident angle, the polarization normal vector is obtained using the third formula.

[0036] The third formula is as follows:

[0037]

[0038] Where, n x express The component in the x-axis direction, n y n represents the component along the y-axis. z This represents the component along the z-axis. θ represents the azimuth angle, and θ represents the angle of incidence.

[0039] In one embodiment of the present invention, based on the multi-view image, obtaining the spatial position and view orientation of the polarization camera includes:

[0040] The COLMAP algorithm is used to process the multi-view images to obtain the position (x, y, z) and orientation of the polarization camera corresponding to the multi-view images.

[0041] The fourth formula is used to determine the position (x, y, z) and orientation of the polarization camera, respectively. Perform position encoding to obtain the corresponding spatial position γ(x,y,z) and view orientation.

[0042] The fourth formula is as follows:

[0043] γ(η)=[sin(20πη),cos(20πη),…,sin((2L-1)πη),cos((2L-1)πη)];

[0044] Among them, γ(η)=γ(x,y,z) or L represents preset parameters. When processing spatial position, η = x, y, z; when processing view direction...

[0045] In one embodiment of the present invention, based on the spatial location and view orientation, the symbolic distance function (SDF) and color values ​​of the target scene are obtained using an MLP network, including:

[0046] The spatial location is mapped to the signed distance on the surface to obtain the signed distance function SDF of the target scene;

[0047] The colors of the spatial location and view direction are encoded to obtain the color values ​​of the target scene.

[0048] In one embodiment of the present invention, the scene surface normal vector is obtained based on the symbolic distance function SDF, including:

[0049] The gradient of the symbolic distance function SDF is calculated using the fifth formula to obtain the scene surface normal vector.

[0050] The fifth formula is as follows:

[0051]

[0052] in, This indicates the gradient calculation operation, and f(p) represents the sign distance function SDF.

[0053] In one embodiment of the present invention, the scene surface normal vector is used as a guiding vector to correct the polarization normal vector, resulting in a disambiguated surface normal vector, including:

[0054] The polarization normal vector is corrected by minimizing the energy loss function to obtain the disambiguated surface normal vector;

[0055] The energy loss minimization function is as follows:

[0056]

[0057] Where u = (x, y), v = (x+1, y) and w = (x, y+1) represent pixels, and E pol (n polar (u) represents the polarization energy loss function, which is obtained based on the scene surface normal vector and the polarization normal vector. E smooth (n polar (u),n polar (v),n polar (w) represents the energy loss function of the smoothing term, n polar (u) represents the polarization normal vector at pixel u, n polar (v) represents the polarization normal vector at pixel v, n polar(w) represents the polarization normal vector at pixel w, and λ1 represents the preset weight.

[0058] In one embodiment of the present invention, the loss function of the improved NeuS network is as follows:

[0059] L all =L color +λ nomal L nomal +λ reg L reg ;

[0060] Among them, L color L represents the color loss function. nomal L represents the normal vector loss function. reg λ represents the loss of the process function. nomal and λ reg This indicates the corresponding weight.

[0061] The beneficial effects of this invention are:

[0062] The solution provided in this invention obtains the camera's position and orientation from multi-view images through polarization camera pose estimation, and acquires scene surface normal vectors through an MLP network. These scene surface normal vectors are then used to correct ambiguous polarization normal vectors obtained from polarization images, resulting in prior normal vectors. These prior normal vectors are input into a pre-trained improved NeuS network to provide additional constraints, mitigating geometric blurring in textureless and specular reflection regions. Furthermore, compared to depth priors, networks based on prior normal vectors can produce smoother geometry and avoid depth scale blurring between multiple views, obtaining more accurate 3D information. This invention can also perform high-precision reconstruction of occluded objects and objects with thin structures, and can achieve good geometric reconstruction even in scenes with sudden depth changes. Attached Figure Description

[0063] Figure 1 A schematic diagram illustrating the steps of a high-precision stereo reconstruction method for complex scenes provided in an embodiment of the present invention;

[0064] Figure 2 This is a flowchart illustrating a high-precision stereo reconstruction method for complex scenes provided in an embodiment of the present invention.

[0065] Figure 3 The surface coordinate map of the target scene in a high-precision stereo reconstruction method for complex scenes provided in an embodiment of the present invention. Detailed Implementation

[0066] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.

[0067] In existing technologies, Neural Radiance Fields (NeRF) is a computer vision technique used to generate high-quality 3D reconstruction models. It utilizes deep learning to extract the geometric shape and texture information of objects from images viewed from multiple perspectives, and then uses this information to generate a continuous 3D radiation field, thus presenting highly realistic 3D models at arbitrary angles and distances. NeRF technology combines input images from multiple perspectives and camera parameters to build a multilayer perceptron (MLP) model to represent the color and density of each point in the scene. The implementation structure of NeRF includes two main components: an encoder and a decoder.

[0068] The encoder is responsible for extracting the spatial location and viewpoint features of each point in the scene from multiple input viewpoint images and camera parameters. NeRF proposes a positional encoding method that can map the input data from a low-dimensional space to a high-dimensional space and extract more complex feature representations.

[0069] Decoders typically consist of multilayer perceptrons (MLPs) responsible for generating a continuous three-dimensional radiation field from the features extracted by the encoder. Specifically, the decoder takes the spatial location and viewpoint features of each point from the encoder as input and outputs the color and density values ​​of that point. Each MLP layer in the decoder can map the input data to another high-dimensional space and extract more complex feature representations.

[0070] During the training phase, NeRF uses a set of 2D images and corresponding camera parameters as input, and transforms the input data into a 3D scene using rendering equations. The rendering equations are used to calculate the intersection points of rays emitted from the camera position and direction with objects in the scene, and to determine the color and density values ​​of each point. NeRF then uses a neural network to approximate the rendering equations to minimize the difference between the generated scene and the real image.

[0071] However, most neural implicit network-based 3D reconstruction methods cannot reconstruct scenes with large, untextured areas, such as white walls, floors, and reflective scenes, because these areas do not contain enough visual features for pixel-level optimization. Furthermore, while NeRF's view-dependent appearance rendering may seem plausible at first glance, closer examination of specular highlights reveals artificial gloss artifacts that fade in and out between rendered views, rather than moving smoothly across the surface in a physically plausible manner. This is because NeRF can only accurately render the appearance of scene points from specific viewing directions observed in the training images, and its interpolation of glossy appearances from new viewpoints is poor; NeRF tends to use the isotropic properties within objects to "fake" specular reflections instead of view-dependent radiation emitted from surface points, resulting in objects having a translucent or "foggy" outer shell.

[0072] In existing technologies, deep learning-based multi-view stereo (MVSNet) methods take a reference image and multiple source images as input to predict a depth map for the reference image, rather than the entire 3D scene. The network first extracts features from the 2D images to obtain feature maps, then constructs a 3D cost volume based on the camera frustum of the reference view using a differentiable homography transformation. The cost volume is then regularized using 3D convolution, and regression is used to obtain an initial depth map; this initial depth map is then optimized using the reference image to obtain the final depth map.

[0073] However, deep learning-based multi-view stereo methods first estimate the depth map of the image separately, and then use additional filtering and depth fusion to reconstruct the scene. However, due to inconsistencies caused by individual depth map estimations, this method often suffers from incompleteness, surface noise, and scale blurring, and still struggles to handle textureless and repetitive texture regions, resulting in insufficient accuracy in the reconstruction results. Furthermore, this method requires 3D geometric information (ground truth) for training, which is often difficult to obtain in real-world scenes; if synthetic data is used, the trained network's generalization ability is insufficient.

[0074] To address the aforementioned problems, embodiments of the present invention provide a high-precision stereo reconstruction method for complex scenes, such as... Figure 1 As shown, it may include:

[0075] S1, using a polarization camera to acquire multi-view images and multi-angle polarization images of the target scene;

[0076] S2, process the multi-angle polarization image to obtain polarization information and the degree of polarization of the target scene surface;

[0077] S3, based on polarization information and the degree of polarization of the target scene surface, obtain the azimuth and incident angle of the incident light on the target scene surface;

[0078] S4, based on the azimuth and incident angle, obtain the polarization normal vector;

[0079] S5. Based on multi-view images, obtain the spatial position and view orientation of the polarization camera; based on the spatial position and view orientation, use an MLP network to obtain the symbolic distance function (SDF) and color value of the target scene; based on the symbolic distance function (SDF), obtain the scene surface normal vector; use the scene surface normal vector as a guide vector to correct the polarization normal vector and obtain the disambiguated surface normal vector.

[0080] S6 uses the disambiguated surface normal vector as the prior normal vector, performs volume rendering based on the position and orientation of the polarization camera, the signed distance function SDF, the color value, and the prior normal vector to obtain the predicted color value and normal vector; and obtains the three-dimensional information of the target scene based on the difference between the predicted normal vector and the prior normal vector, as well as the difference between the predicted color value and the color value of the target scene.

[0081] This invention provides a high-precision stereo reconstruction method for complex scenes. It uses camera pose estimation to obtain the camera's position and orientation as input, and an MLP network to obtain scene surface normal vectors. These scene surface normal vectors are then used to correct the polarization normal vectors derived from polarization images, resulting in prior normal vectors. These prior normal vectors are input into a pre-trained improved NeuS network to provide additional constraints, mitigating geometric blurring in textureless and specular reflection regions. Furthermore, compared to depth priors, networks based on prior normal vectors can produce smoother geometry and avoid depth scale blurring between multiple views, thus obtaining more accurate 3D information.

[0082] Please refer to the detailed flowchart. Figure 2 ,from Figure 2 As can be seen, after obtaining multi-view images and multi-angle polarized images of the target scene using a polarization camera in step S1, it can be divided into two parallel processing steps, simultaneously processing the multi-view images and multi-angle polarized images. Specifically, the multi-angle polarized images are processed based on Stokes vectors using steps S2-S4 to obtain polarization normal vectors; the multi-view images are processed using the corresponding steps in step S5 to obtain scene surface normal vectors. The polarization normal vectors are then corrected based on the scene surface normal vectors to obtain disambiguated surface normal vectors, which serve as prior rough normal vectors; thus, step S6 performs corresponding processing to obtain the three-dimensional information of the target scene. It is understood that since obtaining scene surface normal vectors in S2 and S5 can be done in parallel, the step numbers should not be construed as limitations on the steps in this embodiment of the invention.

[0083] For ease of understanding, the following describes each step of a high-precision stereo reconstruction method for complex scenes provided by an embodiment of the present invention.

[0084] For S1, it can include:

[0085] The target scene was captured using a FLIR polarization camera to obtain video frames;

[0086] Extract multi-view images from video frames and obtain multi-angle image data of the scene. Combine this with the polarizer on the polarization camera to obtain multi-angle polarized images I0 and I0. 45 I 90 I 135.

[0087] Specifically, the polarizer on the polarization camera is a four-way (0°, 45°, 90°, and 135°) polarizer on the camera sensor. Using this polarizer, the intensity and polarization angle of each pixel can be output, thus obtaining multi-angle polarized images I0, I... of the target scene from four angles. 45 I 90 I 135 .

[0088] For S2, it can include:

[0089] Multi-angle polarization images are processed to obtain polarization information; the polarization information is composed of Stokes vectors; the expression for the Stokes vectors is as follows:

[0090]

[0091] Among them, I 0° I represents the light intensity passing through a 0° polarizer. 90° I represents the light intensity passing through a 90° polarizer. +45° I represents the light intensity passing through a polarizer at a 45° angle to the positive horizontal direction. -45° I represents the light intensity obtained through a linear polarizer at a 135° angle to the horizontal. RH I represents the right-hand circularly polarized light component. LH This represents the left-hand circularly polarized light component.

[0092] Based on the Stokes vector, the degree of polarization of the target scene surface is obtained using the first formula; the expression of the first formula is as follows:

[0093]

[0094] Specifically, the polarization of light is described using two or more photons with different polarization states. The polarization states and their sum can be described by the Stokes vector.

[0095] Understandably, the degree of polarization (DOP) of any optical signal can be directly determined using the Stokes vector, which characterizes the proportion of polarization components.

[0096]

[0097] If the beam is linearly polarized, then the elliptic and circular polarization components are zero, leading to the first formula. P is generally called the degree of linear polarization, characterizing the linear polarization degree of the beam. Understandably, for a linearly polarized beam, the degree of polarization DOP can be represented by P.

[0098] For S3, it can include:

[0099] S31, based on the polarization information and the degree of polarization of the target scene surface, the azimuth angle of the incident light on the target scene surface is obtained using the second formula; for details, please refer to the surface coordinate diagram of the target scene. Figure 3 .

[0100] The second formula is as follows:

[0101]

[0102] in, S0 and S1 represent the azimuth angle and the Stokes vector.

[0103] S32, based on the relationship between the degree of polarization and the angle of incidence, the angle of incidence is obtained;

[0104] The relationship between the degree of polarization and the angle of incidence is as follows:

[0105] For specular reflection:

[0106]

[0107] For diffuse reflection:

[0108]

[0109] Where P represents the degree of polarization, n represents the refractive index of the object's surface, and θ represents the angle of incidence.

[0110] Optionally, the refractive index n of the object's surface can be set to 1.5. If the pixel is dominated by diffuse reflection, the phase angle is uniquely determined by the degree of polarization, and the azimuth angle is limited to two possibilities by the phase angle, resulting in two possible normal vector directions. If the pixel is dominated by specular reflection, the degree of polarization limits the viewing angle to two possibilities, and the azimuth angle is also limited to two possibilities, giving a total of four possible normal vector directions.

[0111] For S4, it can include:

[0112] The polarization normal vector is obtained using the third formula based on the azimuth and incident angle.

[0113] The third formula is as follows:

[0114]

[0115] Where, n x express The component in the x-axis direction, n y n represents the component along the y-axis. z This represents the component along the z-axis. θ represents the azimuth angle, and θ represents the angle of incidence.

[0116] Because a given light intensity can correspond to multiple polarizer rotation angles, calculating the azimuth angle using light intensity presents a 180° multivariability problem, leading to inaccurate acquisition of the polarization normal vector. Therefore, we need to correct the polarization normal vector before reconstruction.

[0117] For S5, it can include:

[0118] S51, based on multi-view images, acquires the spatial position and view orientation of the polarization camera, which may include:

[0119] The COLMAP algorithm is used to process multi-view images to obtain the position (x, y, z) and orientation of the polarization camera corresponding to the multi-view images.

[0120] The fourth formula was used to determine the position (x, y, z) and orientation of the polarization camera, respectively. Perform position encoding to obtain the corresponding spatial position γ(x,y,z) and view orientation.

[0121] The fourth formula is as follows:

[0122] γ(η)=[sin(20πη),cos(20πη),…,sin((2L-1)πη),cos((2L-1)πη)];

[0123] Among them, γ(η)=γ(x,y,z) or L represents preset parameters. When processing spatial position, η = x, y, z; when processing view direction...

[0124] Specifically, for the COLMAP algorithm, the point cloud is reconstructed using the true pose of the ground, and then a filtered Poisson surface reconstruction is used to obtain the mesh, which is then used to obtain the 3D coordinates and orientation of the camera corresponding to the image. Position encoding is then used to map these coordinates to high frequencies. The preset parameter L can be set according to requirements. Optionally, the preset parameter L can be set to 10.

[0125] S52, based on the spatial location and view orientation, using an MLP network to obtain the symbolic distance function (SDF) and color values ​​of the target scene, may include:

[0126] The signed distance between spatial locations and surfaces is mapped to obtain the signed distance function SDF of the target scene;

[0127] The color values ​​of the target scene are obtained by encoding the color of the spatial location and the view direction.

[0128] Specifically, the scene to be reconstructed can be encoded by two multilayer perceptrons (MPLs). The MLP encoding f can consist of eight hidden layers, each with a size of 256. The activation function used is the Softplus function with β = 100, which maps the spatial location γ(x,y,z) to the signed distance on the surface, thus obtaining the signed distance function SDF of the target scene. This function processes the three-dimensional data into one-dimensional data (R... 3 →R). The MLP of encoder c can consist of 6 hidden layers, each with a size of 256, used for spatial position γ(x,y,z) and view orientation. The colors are encoded to obtain the color values ​​of the target scene, which processes the 3D and 2D data into 3D data (R). 3 ×S 2 →R 3 ).

[0129] S53, based on the symbolic distance function SDF, obtains the scene surface normal vector, which may include:

[0130] The gradient of the signed distance function SDF is calculated using the fifth formula to obtain the scene surface normal vector.

[0131] The fifth formula is as follows:

[0132]

[0133] in, This indicates the gradient calculation operation, and f(p) represents the sign distance function SDF.

[0134] The obtained scene surface normal vector It can be used for subsequent correction of the polarization normal vector.

[0135] S54, using the scene surface normal vector as a guiding vector, correcting the polarization normal vector to obtain the disambiguated surface normal vector, which may include:

[0136] The polarization normal vector is corrected by minimizing the energy loss function to obtain the disambiguated surface normal vector;

[0137] The function to minimize energy loss is as follows:

[0138]

[0139] Where u = (x, y), v = (x+1, y), and w = (x, y+1) represent pixels, and E pol (n polar (u) represents the polarization energy loss function, which is obtained based on the scene surface normal vector and the polarization normal vector. E smooth(n polar (u),n polar (v),n polar (w) represents the energy loss function of the smoothing term, n polar (u) represents the polarization normal vector at pixel u, n polar (v) represents the polarization normal vector at pixel v, n polar (w) represents the polarization normal vector at pixel w, and λ1 represents the preset weight.

[0140] Understandably, the constraints in S4 limit the surface normal vectors at each pixel to six possible directions. A higher-order graphics model can be used for integrability-based disambiguation, specifically by minimizing the energy loss function to obtain the disambiguated surface normal vectors. The preset weight λ1 can be determined automatically by the computer during computation to minimize the energy loss function.

[0141] Specifically, in minimizing the energy loss function, the expression for the polarization energy loss function can be: The polarization energy loss function can represent the polarization normal vector n polar (u) and scene surface normal vector The angle difference between them.

[0142]

[0143] Where, f(u) = exp(-n polar (u)·n), f(u) represents the polarization normal vector n polar (u) and scene surface normal vector The cosine value between.

[0144] Based on the previous analysis, we know that n polar (u) can have a maximum of 6 solutions, where D represents the 2 solutions for diffuse reflection, J represents the 4 solutions for specular reflection, and L can be the initial specular reflection mask.

[0145] The expression for the energy loss function of the smoothing term is as follows:

[0146]

[0147] in, This represents the set of pixel triples (u, v, w).

[0148] The smoothing term energy loss function can represent the integrability constraint of the scene surface normal vector. For an integrable surface, the mixed second-order partial derivatives on the gradient field should be equal, i.e. p represents the gradient in the horizontal direction, and q represents the gradient in the vertical direction. However, in reality, due to noise and discrete pixel networks, the gradient field cannot be completely zero-curvature, so it is necessary to find the normal vector with minimum curl.

[0149]

[0150] ψ(n polar (u), n polar (v), n polar (w))=p(w)-p(u)-(q(v)-q(u)).

[0151] Understandably, the disambiguated surface normal vector at this point is the unambiguous n obtained by minimizing the energy function. polar (u). In this embodiment of the invention, higher-order belief propagation is used to minimize energy loss in order to obtain the disambiguated surface normal vector n′. However, the disambiguated surface normal vector still contains noise and is affected by low-frequency bias.

[0152] For S6, the disambiguated surface normal vector is used as the prior normal vector. Using a pre-trained improved NeuS network, the position and orientation of the polarization camera, the signed distance function SDF, the color value and the prior normal vector are used for volume rendering to obtain the predicted color value and normal vector. Based on the difference between the predicted normal vector and the prior normal vector, as well as the difference between the predicted color value and the color value of the target scene, the three-dimensional information of the target scene is obtained.

[0153] Specifically, for each pixel, a set of points is sampled along the corresponding emitted ray, sampling point p i =o+d i e, where o represents the center of the polarization camera, e represents the direction of the emitted light, and the color can be the accumulation of the colors of the sampling points along this direction, as shown in the following formula:

[0154]

[0155] Among them, T i This represents the discrete cumulative transmittance corresponding to the i-th sampling point. α i This represents the discrete opacity corresponding to the i-th sampling point. α i It can be related to the i-th sampling point p i The corresponding signed distance function f(p) i (t i Functions related to )).

[0156] The normal vector of a surface at a viewpoint can also be predicted through volume rendering accumulation. Specifically,

[0157]

[0158] Where, n i ′ represents the normal vector of the i-th sampling point.

[0159] During the training phase of the improved NeuS network, in space {c k ,o k ,e k A batch of pixels is randomly sampled from the target scene. The difference between the predicted normal vector and the prior normal vector, the difference between the predicted color value and the target scene color value, and the difference between the corresponding reference are used as the loss function of the improved NeuS network. Based on this loss function, the corresponding weights are learned to obtain accurate 3D information of the target scene. Where c k Represents the color value of a pixel, o k e represents the center position of the polarization camera. k The direction of the polarization camera is represented by h, the number of sampling points along each ray is h, and the number of sampling batches is m.

[0160] This invention utilizes polarization information to separate specular and diffuse reflections, solves for the normal vector, and uses it as a priori input in the improved NeuS network after disambiguation. This enables the reconstruction of specular reflections and textureless parts of the scene, resulting in a high-precision 3D scene shape.

[0161] Specifically, the loss function of the improved NeuS network is as follows:

[0162] L all =L color +λ nomal L nomal +λ reg L reg ;

[0163] Among them, L color L represents the color loss function. nomal L represents the normal vector loss function. reg λ represents the loss of the process function. nomal and λ reg This indicates the corresponding weight.

[0164] Specifically, C represents the predicted color value. k Represents the color value of the target scene. n′ represents the prior normal vector. Represents the predicted normal vector. This represents the gradient calculation operation, where h represents the number of sampling points, m represents the number of sampling batches, i represents the i-th sampling point, k represents the k-th pixel, and f(p) k,i Let ) represent the symbolic distance function corresponding to k,i.

[0165] The high-precision stereo reconstruction method for complex scenes provided in this invention introduces polarization information to perform three-dimensional reconstruction of complex large scenes. It can accurately reconstruct textureless areas and specular reflection areas in the scene, such as walls, floors, and furniture, effectively overcoming the technical defects of multi-view three-dimensional reconstruction methods. It also has the ability to reconstruct the fine details of occluded, complex, and thin-structured objects, overcoming the scale ambiguity problem of depth estimation-based reconstruction methods.

[0166] The key idea of ​​this invention lies in utilizing the differences in polarization information regarding specular and diffuse reflection, the robustness of polarization information in textureless regions, and the dependence of polarization on surface normal vectors. The disambiguated prior normal vectors are used as globally consistent geometric constraints to guide the optimization process. Specifically, the camera position and orientation are first obtained from multi-view images through polarization camera pose estimation, and scene surface normal vectors are obtained through an MLP network. These scene surface normal vectors are then used to correct the ambiguous polarization normal vectors obtained from the polarized images, thus obtaining prior normal vectors. These prior normal vectors are input into a pre-trained improved NeuS network to provide additional constraints, mitigating geometric blurring in textureless and specular regions. Furthermore, compared to depth priors, networks based on prior normal vectors can produce smoother geometry and avoid depth scale blurring between multiple views, resulting in more accurate 3D information. This invention can also perform high-precision reconstruction of occluded objects and objects with thin structures, and can also perform good geometric reconstruction of scenes with sudden depth changes.

[0167] It should be noted that, in the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0168] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. A high-precision stereo reconstruction method for complex scenes, characterized in that, include: Use a polarization camera to acquire multi-view and multi-angle polarization images of the target scene; Multi-angle polarization images are processed to obtain polarization information and the degree of polarization of the target scene surface; Based on the polarization information and the degree of polarization of the target scene surface, the azimuth angle and incident angle of the incident light on the target scene surface are obtained; Based on the azimuth and incident angle, the polarization normal vector is obtained; Based on the multi-view images, the spatial position and view orientation of the polarization camera are obtained; Based on the spatial location and view orientation, the symbolic distance function (SDF) and color values ​​of the target scene are obtained using an MLP network; based on the symbolic distance function (SDF), the scene surface normal vector is obtained. Using the scene surface normal vector as a guiding vector, the polarization normal vector is corrected to obtain the disambiguated surface normal vector, including: The polarization normal vector is corrected by minimizing the energy loss function to obtain the disambiguated surface normal vector; The energy loss minimization function is as follows: ; in, , and Represents pixels, This represents the polarization energy loss function, which is derived based on the scene surface normal vector and the polarization normal vector. This represents the energy loss function for the smoothing term. Represents pixels The polarization normal vector at that point, Represents pixels The polarization normal vector at that point, Represents pixels The polarization normal vector at that point, Indicates the preset weight; The disambiguated surface normal vector is used as the prior normal vector. A pre-trained improved NeuS network is used to perform volume rendering on the position and orientation of the polarization camera, the signed distance function SDF, the color value and the prior normal vector to obtain the predicted color value and normal vector. Based on the difference between the predicted normal vector and the prior normal vector, as well as the difference between the predicted color value and the color value of the target scene, the three-dimensional information of the target scene is obtained.

2. The high-precision stereo reconstruction method for complex scenes according to claim 1, characterized in that, The method of acquiring multi-view images and multi-angle polarized images of the target scene using a polarization camera includes: The target scene was captured using a FLIR polarization camera to obtain video frames; Extract multi-view images from video frames and obtain multi-angle image data of the scene. Combine this with a polarizer on a polarizing camera to obtain multi-angle polarized images. , , , .

3. The high-precision stereo reconstruction method for complex scenes according to claim 1, characterized in that, The process of processing multi-angle polarization images to obtain polarization information and the degree of polarization of the target scene surface includes: The multi-angle polarization image is processed to obtain polarization information; the polarization information is composed of Stokes vectors; the expression of the Stokes vectors is as follows: ; in, This indicates the light intensity passing through a 0° polarizer. This indicates the light intensity passing through a 90° polarizer. This indicates the light intensity passing through a polarizer at a 45° angle to the positive horizontal direction. This represents the light intensity obtained through a linear polarizer at a 135° angle to the horizontal. This represents the right-hand circularly polarized light component. Indicates the left-handed circularly polarized light component; Based on the Stokes vector, the degree of polarization of the target scene surface is obtained using the first formula; the expression of the first formula is as follows: 。 4. The high-precision stereo reconstruction method for complex scenes according to claim 1, characterized in that, Based on the polarization information and the degree of polarization of the target scene surface, the azimuth and incident angle of the incident light on the target scene surface are obtained, including: Based on the polarization information and the degree of polarization of the target scene surface, the azimuth angle of the incident light on the target scene surface is obtained using the second formula. The second formula is as follows: ; in, Indicates azimuth. , Represents the Stokes vector; The angle of incidence is obtained based on the relationship between the degree of polarization and the angle of incidence. The relationship between the degree of polarization and the angle of incidence is as follows: For specular reflection: ; For diffuse reflection: ; in, Indicates degree of polarization. Represents the refractive index of an object's surface. Indicates the angle of incidence.

5. The high-precision stereo reconstruction method for complex scenes according to claim 1, characterized in that, Based on the azimuth and incident angle, the polarization normal vector is obtained, including: Based on the azimuth and incident angle, the polarization normal vector is obtained using the third formula. ; The third formula is as follows: ; in, express exist Components in the axial direction, Indicates in Components in the axial direction, Indicates in Components in the axial direction, Indicates azimuth. Indicates the angle of incidence.

6. The high-precision stereo reconstruction method for complex scenes according to claim 1, characterized in that, Based on the multi-view images, the spatial position and view orientation of the polarization camera are obtained, including: The COLMAP algorithm is used to process the multi-view images to obtain the positions of the polarization cameras corresponding to the multi-view images. and orientation ; The position of the polarization camera was determined using the fourth formula. and orientation Perform position encoding to obtain the corresponding spatial location. and view direction ; The fourth formula is as follows: ; in, or , This indicates preset parameters used when processing spatial positions. When processing the view orientation, .

7. The high-precision stereo reconstruction method for complex scenes according to claim 1, characterized in that, Based on the spatial location and view orientation, the symbolic distance function (SDF) and color values ​​of the target scene are obtained using an MLP network, including: The spatial location is mapped to the signed distance on the surface to obtain the signed distance function SDF of the target scene; The colors of the spatial location and view direction are encoded to obtain the color values ​​of the target scene.

8. The high-precision stereo reconstruction method for complex scenes according to claim 1, characterized in that, Based on the symbolic distance function SDF, the scene surface normal vector is obtained, including: The gradient of the symbolic distance function SDF is calculated using the fifth formula to obtain the scene surface normal vector. ; The fifth formula is as follows: ; in, This indicates the gradient calculation operation. This represents the symbolic distance function SDF.

9. The high-precision stereo reconstruction method for complex scenes according to claim 1, characterized in that, The loss function of the improved NeuS network is as follows: ; in, Represents the color loss function. Represents the normal vector loss function. Indicates the loss of Cheng Han. and This indicates the corresponding weight.

Citation Information

Patent Citations

  • Large-scene polarization three-dimensional imaging method based on monocular depth estimation

    CN116363301A