A polarization 3D reconstruction method and system based on prior-guided fusion network
By using a feature fusion network guided by physical priors, the problems of π-ambiguity and unknown reflection properties in polarization 3D reconstruction are solved, achieving high-precision surface normal estimation and 3D reconstruction, especially with improvements under complex light source conditions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- WUHAN UNIV
- Filing Date
- 2023-11-28
- Publication Date
- 2026-05-26
AI Technical Summary
Existing polarization 3D reconstruction methods suffer from large estimation errors in the surface normal estimation of π-fuzzy domains, and the unknown surface reflection characteristics of objects make it difficult to calculate the zenith, especially in complex scenes where the reconstruction effect is poor.
A physical prior-guided feature fusion network is adopted. The polarization and shadow prior information are corrected in the channel and spatial dimensions through the feature correction module, and the cross-attention mechanism is used for feature fusion. A dual-branch feature fusion network is constructed to reconstruct the surface normals with high accuracy.
It significantly improves the reconstruction quality of surface normals, especially reducing the angular error of specular reflection areas in objects or scenes illuminated by complex light sources, and achieves high-quality 3D target reconstruction.
Smart Images

Figure CN117671142B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision and deep learning technology, and relates to a polarization 3D reconstruction method and system based on a physical prior guided feature fusion network, which is suitable for 3D reconstruction application scenarios with high precision requirements. Background Technology
[0002] Polarization-based 3D reconstruction is an important method for target 3D reconstruction and has been extensively studied over the past few decades. This technique estimates the surface normal vector of a target based on its reflection characteristics of polarized light, thus reconstructing the 3D surface. As a passive 3D reconstruction method, polarization-based 3D reconstruction excels at reconstructing dense texture details and can also be applied to the reconstruction of specular reflections and transparent objects. While some active 3D reconstruction techniques offer advantages such as high resolution, like LiDAR and structured light, LiDAR only provides sparse 3D data and is susceptible to noise, while structured light has a very limited imaging range, restricting its application. Therefore, compared to other 3D imaging techniques, polarization-based 3D reconstruction is superior in target detection and segmentation. [1] It has greater potential in applications. However, in order to achieve higher accuracy in surface normal estimation, the following two problems need to be solved: (1) In the physical model of polarization 3D reconstruction, the azimuth angle has two solutions with a difference of π radians, rather than a unique solution. The surface normal estimation error in the π-fuzzy domain is large, which leads to inaccurate 3D shape estimation of the corresponding region. (2) Calculating the zenith requires knowing the reflection characteristics of the object surface, because the physical model of the object surface in the diffuse reflection region and the specular reflection region are different. However, in most applications, the reflection characteristics of the object surface are usually unknown. In addition, most of the reflected light from the real object surface contains both diffuse light and specular light, which makes calculating the zenith more difficult.
[0003] To address the two aforementioned issues, numerous physical model-based polarization 3D reconstruction methods have been proposed, which can be broadly categorized into two types: depth data-based polarization 3D reconstruction methods (SfP-D) and illumination or shadow constraint-based polarization 3D reconstruction methods (SfP-IS). Kadambi et al. designed a framework that utilizes depth maps obtained from a structured light-based Kinect camera to resolve azimuth ambiguity and correct zenith uncertainty. Since radar can acquire sparse depth data from distant targets, Yoshida et al. proposed a framework that combines the absolute depth obtained from radar with the surface data from polarization 3D reconstruction to reconstruct specular depth. [2]Furthermore, while binocular stereo vision can provide high spatial resolution depth data, its depth accuracy is limited. Zhu et al. used depth maps obtained from binocular cameras as guiding surfaces for initial surface normal disambiguation and employed a high-order graphical model to estimate accurate surface normals. Tian et al. further proposed a joint 3D reconstruction model for polarization 3D reconstruction and active binocular vision depth fusion. [3] This approach aims to address the issues of missing data and outliers in binocular depth images. However, accurate registration between depth and polarization images remains a significant challenge in achieving high-quality 3D reconstruction. To avoid image registration issues, Liu et al. proposed a method that fuses polarization-based 3D reconstruction and polarization modulation ranging to obtain high-quality reconstruction results. [4] However, this method requires multiple modulation imaging steps to obtain depth data, thus limiting the imaging speed. On the other hand, variations in image brightness also reflect its three-dimensional surface geometry. [5]-[8] Based on this characteristic, Drbohlav et al. introduced constraints derived from illumination information to solve the π-ambiguity problem in polarization 3D reconstruction. [9] In addition, Olivier et al. designed an active lighting system consisting of multiple LEDs.
[10] To acquire images under four specific illumination conditions, these images were then combined with polarization information to reconstruct the surface normal. To simplify the illumination setup, Atkinson et al. proposed a novel shape estimation method using polarization and shadow information from two views obtained by rotating the target. To further simplify the imaging system and computation, Mahmoud et al. proposed a method for directly and quickly calculating the zenith using only polarization information, then addressing the π-ambiguity problem of the azimuth angle using complementary shadow and polarization information. Ngo et al. simultaneously estimated the target surface normal and refractive index by combining illumination, shadow, and polarization information.
[11] .
[0004] In recent years, deep learning has been introduced into the field of polarization 3D reconstruction, providing new solutions to the problem. In 2020, Kondo et al. proposed a novel polarization bidirectional reflection distribution function model.
[12] This was used to create a synthetic dataset of polarization images and to estimate surface normals based on a convolutional neural network. Ba et al. proposed a U-Net-based approach.
[13] The neural network. Subsequently, Lei et al. constructed a new complex scene dataset.
[14] Furthermore, a deep learning-based polarization 3D reconstruction method was proposed. The three methods mentioned above only utilize prior polarization information, which leads to significantly poor reconstruction results in complex scenes (especially when the target surface has a large specular reflection area). In addition, Shao et al. specifically proposed a learning-based polarization 3D reconstruction method for transparent objects.
[15] They conducted experiments on their transparent object dataset, which included both real-world and synthetic ensembles. Muglikar et al. used an event camera to record continuous polarization information caused by rotation.
[16] Instead of inputting polarization images of multiple polarization angles, the relative intensities of multiple polarization angles are reconstructed, thereby obtaining polarization information for all polarization angles.
[0005] While introducing reliable physical prior information (such as shading prior) into deep learning-based polarization 3D reconstruction methods is valuable, two challenging problems remain: (1) Existing polarization 3D reconstruction datasets for deep learning only contain polarization images with unknown illumination directions and target surface reflectivity, making it difficult to directly obtain shadow priors. Constructing new datasets by capturing images with different illumination directions can achieve shadow prior calculation, but this increases imaging time and computational complexity. (2) Effective fusion of polarization and shadow priors in deep networks is worth studying. Polarization priors suffer from π-ambiguity, while shadow priors often have estimation errors due to uncertainties in illumination direction and surface reflectivity. Extracting reliable information from these two priors and effectively fusing them through deep learning is key to achieving high-quality target surface normal reconstruction. Summary of the Invention
[0006] This invention addresses the shortcomings of existing technologies by proposing a novel polarization 3D reconstruction method and system based on a physical prior-guided feature fusion network. A feature correction module automatically corrects defects in the input polarization and shadow prior bi-branch information at both the channel and spatial dimensions. A feature fusion module based on a cross-attention mechanism effectively fuses the corrected polarization and shadow prior feature maps, achieving high-precision surface normal vector estimation and thus reconstructing a high-quality 3D target.
[0007] The technical solution adopted in this invention is a polarization 3D reconstruction method based on a physical prior-guided feature fusion network. First, a polarization prior map set is obtained from polarization images at different polarization angles, including polarization degree, unpolarized image, re-represented azimuth angle, and field-of-view encoding, which serves as the input to the first branch of the feature fusion network. Next, a shadow prior map set is obtained from polarization images at different polarization angles, unpolarized image, and the polarization prior surface normal reconstruction results, including specular confidence and shadow prior surface normal reconstruction results, which serves as the input to the second branch of the feature fusion network. Then, a feature correction module is constructed to correct defects in the input information of the two branches of the network in the channel dimension and spatial dimension, respectively. For the extraction of feature maps at each layer, a feature extraction model based on an efficient cross-attention mechanism is constructed to enhance information interaction and fuse the features of the two branches into a single feature map. Using ConvNeXt as the backbone of the encoder and decoder, and incorporating feature correction and feature fusion modules, a feature fusion network for polarization 3D reconstruction is constructed. A loss function suitable for this network is designed, and the network is trained on the object dataset DeepSfP and the scene dataset SPW, respectively. 3D reconstruction is then performed based on the trained polarization 3D reconstruction model. This method includes the following steps:
[0008] Step 1, Obtain the polarization prior input map set: Obtain polarization images of the target object and target scene at different polarization angles. Calculate a polarization representation map set containing the unpolarized image, degree of polarization, re-represented phase angle, and field-of-view encoding based on the obtained polarization images at different polarization angles. This map set serves as the polarization prior input map set. Simultaneously, obtain the polarization prior target surface reconstruction result.
[0009] Step 2, Obtain the shadow prior input map set: Calculate the shadow prior input map set containing the specular confidence and shadow prior surface normal reconstruction results based on polarized images at different polarization angles, unpolarized images, and polarized prior surface target normal reconstruction results;
[0010] Step 3: Construct a bi-branch feature fusion network for polarization 3D reconstruction. The bi-branch feature fusion network includes a bi-branch encoder and decoder, as well as a feature correction module and a feature fusion module. The polarization prior input map group and the shadow prior input map group are input into the bi-branch encoder to extract prior features. The extracted features are corrected by the feature correction module, and the extracted features are fused by the feature fusion module. Then, the fused features are input into the decoder to obtain the final target surface normal reconstruction result.
[0011] Step 4: Train a dual-branch feature fusion network using a loss function, and use the trained network to achieve polarization 3D reconstruction of the target.
[0012] Furthermore, the dual-branch feature fusion network consists of two encoders for extracting two different prior features. The two branches have the same structure and parameters except for the input channels. Each branch has several layers, and each layer includes a downsampling module Down for downsampling the input image to an appropriate feature map size, as well as a Block composed of multiple ConvNeXt blocks for feature extraction.
[0013] A feature correction module is added between the downsampling modules Down and Block in the first layer, and a feature fusion module is used to fuse the features of each layer output of the two branches to obtain fused features at different scales. Then, a skip connection is used to introduce the fused features at different scales into the decoder to output the final target surface normal reconstruction result.
[0014] Final reconstructed target surface normal Represented as:
[0015]
[0016] Where f(·) represents the prediction model, i.e., the dual-branch feature fusion network constructed in step 3, Pol prior For the polarization prior map set, Sha prior This is the shadow prior group.
[0017] Furthermore, the feature correction module is used to correct the first layer features in the two encoders, including channel correction and spatial correction, and then the results of channel correction and spatial correction are weighted to obtain the result of comprehensive correction.
[0018] Furthermore, the specific processing procedure of the feature correction module is as follows;
[0019] Given a polarization prior Feature maps and shadow priors The feature maps are first concatenated, and then max pooling and average pooling are performed along the channel dimension. Then, the two vectors are concatenated, and MLP is used to obtain the two channel weight mappings. and
[0020] CA pol =CW pil *P pol
[0021] CA sha =CW sha *P sha
[0022] Where * denotes channel multiplication, B represents the batch size, C represents the number of channels, H and W are the pixel size of each channel image, and CApol and CA sha Feature correction results for both channels;
[0023] The structure of spatial feature correction focuses on feature correction of important regions in an image. Therefore, P... pol and P sha Connect them together, and then obtain two spatial weight maps through MLP. and The spatial correction operation is as follows:
[0024]
[0025]
[0026] in SA represents space multiplication. pol and SA sha This indicates the spatial correction results for the two channels;
[0027] The overall characteristics after polarization and shadow prior corrections are as follows:
[0028] FC pol =P pol +α1CA sha +α2SA sha
[0029] FC sha =P sha +β1CA pol +β2SA pol
[0030] Where α1, α2, β1, and β2 are four hyperparameters, FC pol FC sha It is the corrected feature map after comprehensive correction.
[0031] Furthermore, the feature fusion module is used to fuse the features of each layer in the encoder, and the specific processing procedure is as follows:
[0032] Two input branches and The output is then subjected to a 1×1 convolution operation and added to an LN, combined using the GELU function, and finally an MLP is applied to the result to obtain a fused mapping. Furthermore, the information interaction operation is as follows:
[0033] FA pol =O pol +FW※O sha
[0034] FA sha =O sha+FW※O pol
[0035] Finally, the fused feature FA is processed through a ConvNeXt block. pol and FA sha Combined into a single feature B represents the batch size, C represents the number of channels, and H and W represent the pixel size of each channel image.
[0036] Furthermore, the following loss function is used as a guide for network optimization:
[0037]
[0038] in, The total loss is represented by γ1, γ2, and γ3, which are three hyperparameters used to adjust different losses. For the final output of the network The cosine similarity loss, used for normal estimation, is expressed as follows:
[0039]
[0040] Where <·,·> denotes the dot product. The surface normal estimated at pixel position (i, j), N ij This is the true value of the surface normal at that location;
[0041] Furthermore, the features generated by the feature correction module are supervised, therefore cross-entropy loss is used separately. and The output FC of the feature correction module pol and FC sha Loss calculation is performed, and simultaneously, the feature FF generated by the last feature fusion module is processed. out Supervise and use cross-entropy loss
[0042] Furthermore, The expression is as follows:
[0043]
[0044]
[0045]
[0046] Furthermore, the specific implementation of step 1 includes the following sub-steps:
[0047] Step 1.1, the intensity of each pixel in the polarized image at different polarization angles is represented as follows:
[0048]
[0049] Where I un The polarization intensity is given by φ ∈ [0, π), the phase angle is given by ρ, and the degree of polarization (DoP) is given by φ. p ∈I0,I 45 ,I 90 ,I 135 Four polarization images (I0, I 45 ,I 90 ,I 135 I was calculated separately in ) un , φ and ρ;
[0050] Step 1.2, Surface normal azimuth angle The zenith angle θ is calculated from the phase angle φ and the degree of polarization ρ, respectively, and is expressed as:
[0051] or
[0052] When specular reflection is dominant, the relationship between ρ and θ is:
[0053]
[0054] Where n is the reflectivity of the target surface;
[0055] When diffuse reflection is dominant, the relationship between ρ and θ is:
[0056]
[0057] Further, the target surface normal is obtained:
[0058]
[0059] Where, N p For the polarization prior target surface normal reconstruction results, N x N y N z These are the component values of the target surface normal vector in three directions;
[0060] Step 1.3, Obtain the redefinition of the phase angle φ. e And the field-of-view code V, where the field-of-view code V can be obtained from the camera's intrinsic parameters, further yielding the polarization prior map group Pol. prior :
[0061] φ e =(cos2φ,sin2φ)
[0062] Pol prior =(Iun ,φ e ,ρ,V).
[0063] Furthermore, the specific implementation of step 2 includes the following sub-steps:
[0064] Step 2.1, Target Unpolarized Image I un The following relationship exists between the light source direction l and the light source direction:
[0065] I un =TN p ·l
[0066] Where T is a generalized shallow undulation transformation, corresponding to the negative value of the surface normal. The above equation is solved using the least squares method, and the value with the smallest residual is selected. As a solution for the direction of the light source;
[0067] Furthermore, the reflectivity of the target surface is related to the following relationship. Make an estimate:
[0068]
[0069] As a nonlinear optimization problem, the bounded constraint solution for reflectivity is obtained using the trust region reflection algorithm;
[0070] Furthermore, the shadow constrains the target surface normal N. s Estimation from shadow cues using the Lambertian model:
[0071] I un =ηN s ·l=η(l x cosφsinθ+l y sinφsinθ+l z cosθ)
[0072] Where η represents the albedo of the target surface, l x , l y , l z These are the three directional component values of the light source direction vector l;
[0073] Step 2.2: Calculate the mirror confidence S based on the intensity changes of the four original polarization images.
[0074]
[0075] y is a sliding pixel block, I0, I 45 ,I 90 ,I 135 These are the intensity images at the corresponding polarization angles;
[0076] The final prior for the shadow is Sha.prior =(S,N s ).
[0077] This invention provides a polarization 3D reconstruction system based on a priori guided fusion network, comprising the following modules:
[0078] The polarization prior acquisition module is used to acquire a polarization prior input map set: it acquires polarization images of the target object and the target scene at different polarization angles, and calculates a polarization representation map set containing the unpolarized image, degree of polarization, re-represented phase angle, and field-of-view encoding based on the obtained polarization images at different polarization angles, which serves as the polarization prior input map set. Simultaneously, it obtains the polarization prior target surface reconstruction result.
[0079] The shadow prior acquisition module acquires the shadow prior input map set: based on polarized images at different polarization angles, unpolarized images, and the reconstruction results of the surface normals of the polarized prior surface, the shadow prior input map set containing the specular confidence and the reconstruction results of the surface normals of the shadow prior is calculated.
[0080] The network model construction module constructs a bi-branch feature fusion network for polarization 3D reconstruction. The bi-branch feature fusion network includes a bi-branch encoder and decoder, as well as a feature correction module and a feature fusion module. The polarization prior input map group and the shadow prior input map group are input into the bi-branch encoder to extract prior features. The extracted features are corrected by the feature correction module, and the extracted features are fused by the feature fusion module. Then, the fused features are input into the decoder to obtain the final target surface normal reconstruction result.
[0081] The reconstruction module is used to train a dual-branch feature fusion network by combining a loss function, and to use the trained network to achieve polarization 3D reconstruction of the target.
[0082] Compared with existing technologies, the advantages and beneficial effects of this invention are as follows: This invention proposes a 3D polarization reconstruction method based on a fusion network guided by physical priors. It effectively combines polarization and shadow priors, proposing a novel deep fusion network for surface normal information extraction and construction. The network is implemented based on a dual-branch architecture, and a feature correction module is designed to mutually correct defects in the channel and spatial dimensions. Furthermore, a feature fusion module based on an effective cross-attention mechanism is proposed to fuse polarization and shadow prior features. Experimental results show that the fusion of polarization and shadow priors significantly improves the reconstruction quality of surface normals, especially for objects or scenes illuminated by complex light sources. In addition, by introducing specular confidence, the angular error of specular reflection areas can be reduced. Finally, because the network can effectively extract and fuse information from different priors, our proposed method outperforms existing deep learning-based 3D polarization reconstruction methods. Attached Figure Description
[0083] Figure 1 This is a comparison image of the 3D reconstruction results of the target after the fusion of shadow priors.
[0084] Figure 2 This is a diagram showing the reconstruction result of the target surface normal of the shadow prior in the embodiment.
[0085] Figure 3 This is a structural diagram of the feature verification module.
[0086] Figure 4 This is a structural diagram of the feature fusion module.
[0087] Figure 5 It is a diagram of a fusion network structure guided by physical priors.
[0088] Figure 6 This is the result of 3D reconstruction of the target surface using the DeepSfP object-level dataset from the embodiment.
[0089] Figure 7 This is the 3D reconstruction result of the target surface of the SPW scene-level dataset in the example. Detailed Implementation
[0090] To facilitate understanding and implementation of the present invention by those skilled in the art, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the embodiments described herein are for illustrative purposes only and are not intended to limit the present invention.
[0091] This invention primarily addresses the high-precision 3D reconstruction application requirements. We propose a novel deep learning-based polarization 3D reconstruction method. A feature correction module automatically corrects defects in the channel and spatial dimensions of the input polarization and shadow prior information. A feature fusion module based on a cross-attention mechanism effectively fuses the corrected polarization and shadow prior feature maps, achieving high-precision surface normal vector estimation and thus reconstructing a high-quality 3D target. A comparison of the 3D reconstruction results after fusing the shadow prior is provided. Figure 1 As shown. The method specifically includes the following steps:
[0092] Step 1: Obtain the polarization prior input map set. Acquire polarization images of the target object and target scene at different polarization angles. Calculate a polarization representation map set containing the unpolarized image, degree of polarization, re-represented phase angle, and field-of-view encoding based on the obtained polarization images at different polarization angles. This map set serves as the polarization prior input map set. Simultaneously, obtain the polarization prior target surface reconstruction result.
[0093] Step 2: Obtain the shadow prior input map set. Based on polarized images at different polarization angles, unpolarized images, and the surface target normal reconstruction results from the polarization prior, a shadow prior input map set containing the specular confidence and the shadow prior surface normal reconstruction results is calculated.
[0094] Step 3: Construct a feature fusion network for polarization 3D reconstruction. The feature fusion network includes a feature correction module and a feature fusion module. It takes two sets of prior maps (polarization prior input map set and shadow prior input map set) as input and outputs the final target surface normal reconstruction result.
[0095] Step 4: Train the feature fusion network using the loss function, and use the trained network to achieve polarization 3D reconstruction of the target.
[0096] Furthermore, the specific implementation of step 1 includes the following sub-steps:
[0097] Step 1.1, the intensity of each pixel in the polarized image at different polarization angles is represented as follows:
[0098]
[0099] Where I un Let φ be the unpolarized intensity, φ ∈ [0, π) be the phase angle, and ρ be the degree of polarization (DoP). We start from the polarization angle φ p ∈I0,I 45 ,I 90 ,I 135 Four polarization images (I0, I 45 ,I 90 ,I 135 I was calculated separately in ) un , φ and ρ.
[0100] Step 1.2, Surface normal azimuth angle The zenith angle θ can be calculated from the phase angle φ and the degree of polarization ρ, respectively. It is expressed as:
[0101] or
[0102] When specular reflection is dominant, the relationship between ρ and θ is:
[0103]
[0104] Where n is the reflectivity of the target surface.
[0105] When diffuse reflection is dominant, the relationship between ρ and θ is:
[0106]
[0107] Further, the target surface normal is obtained:
[0108]
[0109] Where, N p For the polarization prior target surface normal reconstruction results, N x N y n z These are the component values of the target surface normal vector in the three directions.
[0110] Step 1.3, Obtain the redefinition of the phase angle φ. e And the field-of-view code V, where the field-of-view code V can be obtained from the camera's intrinsic parameters, further yielding the polarization prior map group Pol. prior :
[0111] φ e =(cos2φ,sin2φ)
[0112] Pol prior =(I un ,φ e ,ρ,V)
[0113] Furthermore, the specific implementation of step 2 includes the following sub-steps:
[0114] Step 2.1, Target Unpolarized Image I un The following relationship exists between the light source direction l and the light source direction:
[0115] I un =TN p ·l
[0116] Where T is a generalized shallow undulation transformation, corresponding to the negative value of the surface normal. The above equation is solved using the least squares method, taking the value with the smallest residual. As a solution for the direction of the light source.
[0117] Furthermore, we use the following relationship to determine the reflectivity of the target surface. Make an estimate:
[0118]
[0119] As a nonlinear optimization problem, the bounded constraint solution of reflectivity is obtained using the trust region reflection algorithm.
[0120] Furthermore, the shadow constrains the target surface normal N. s The shadow cues can be estimated using the Lambertian model:
[0121] I un =ηN s ·l=η(lx cosφsinθ+l y sinφsinθ+l z cosθ)
[0122] Where η represents the albedo of the target surface, l x , l y , l z These are the three directional component values of the light source direction vector l.
[0123] Step 2.2: The mirror confidence S can be calculated based on the intensity changes of the four original polarization images.
[0124]
[0125] y is a sliding pixel block, 3×3 in this example. I0,I 45 ,I 90 ,I 135 These are the intensity images at the corresponding polarization angles. S focuses only on the regions with high intensity and large variations in the four original polarization images to avoid interference from different surface materials in the imaging scene.
[0126] Step 2 ultimately yields the shadow prior set as Sha. prior =(S,N s The results of the target surface normal reconstruction under shadow constraints are shown in the attached figure. Figure 2 As shown.
[0127] Furthermore, the specific implementation of step 3 includes the following sub-steps:
[0128] Step 3.1: Construct the Feature Correction Module. Polarization priors suffer from π-ambiguity, while shadow priors often have estimation errors due to uncertainties in illumination direction and surface reflectivity. To effectively integrate these two priors, we propose a Feature Correction Module (FCM) to perform feature correction before feature extraction. The module structure is shown in the attached figure. Figure 3 As shown. For a given polarization prior... Feature maps and shadow priors The feature maps are first concatenated, and then max pooling and average pooling are performed along the channel dimension. Then, the two vectors are concatenated, and MLP is used to obtain the two channel weight mappings. and
[0129] CA pol =CW pol *P pol
[0130] CA sha =CW sha *P sha
[0131] Where * denotes channel multiplication, B represents the batch size, C represents the number of channels, H and W are the pixel size of each channel image, and CA pol and CA sha Feature correction results for the two channels.
[0132] Spatial feature correction focuses on feature correction of important regions in an image. Therefore, we will use P... pol and P sha Connect them together, and then obtain two spatial weight maps through MLP. and The spatial correction operation is as follows:
[0133]
[0134]
[0135] in SA represents space multiplication. pol and SA sha This indicates the spatial correction results for the two channels.
[0136] The overall characteristics after polarization and shadow prior corrections are as follows:
[0137] FC pol =P pol +α1CA sha +α2SA sha
[0138] FC sha =P sha +β1CA pol +β2SA pol
[0139] α1, α2, β1, and β2 are four hyperparameters, with a default value of 0.5 in this example. FC pol FC sha It is the corrected feature map after comprehensive correction.
[0140] Step 3.2: Constructing the Feature Fusion Module. For each layer's extracted feature map, we constructed a feature fusion model based on an efficient cross-attention mechanism to enhance information interaction and merge the features from the two branches into a single feature map. The module structure is shown in the attached figure. Figure 4 As shown. Each layer has two input branches. and The output is then subjected to a 1×1 convolution operation and added to an LN, and combined using the GELU function. Finally, an MLP is applied to the result to obtain a fused mapping. In addition, we will define the information exchange operation as follows:
[0141] FA pol =O pol +FW*O sha
[0142] FA sha =O sha +FW*O pol
[0143] Finally, the fused feature FA is processed through a ConvNeXt block. pol and FA sha Combined into a single feature
[0144] Step 3.3: Construct a dual-branch feature fusion network architecture. Our goal is to estimate surface normals based on the fusion of polarization and shadow priors. To this end, we propose a deep fusion network, the structure of which is attached. Figure 5 As shown. We use the two prior graph sets obtained in steps 1 and 2 as the input information for the two input channels of the dual-branch network.
[0145] The network consists of two encoders for extracting two different prior features. The two branches are identical in structure and parameters, except for the input channels. We chose ConvNeXt (miniature) as the backbone for the two encoders and decoders with the same structure. Each branch of the network has four layers, each including a downsampling module to downsample the input image to an appropriate feature map size, and multiple ConvNeXt blocks for feature extraction. The downsampling module in the first layer (Down1) performs a 4×4 non-overlapping convolution, while the other downsampling modules (Down2, Down3, and Down4) perform standard 2×2 non-overlapping convolutions. The number of output channels for Down1, Down2, Down3, and Down4 are (96, 192, 384, and 768), respectively. The number of ConvNeXt blocks in the four layers is (b1, b2, b3, b4) = (3, 3, 9, 3). A 7×7 kernel convolution is used within the ConvNeXt blocks to obtain the global receptive field. Next, layer normalization (LN) and Gaussian error linear unit (GELU) activation functions are applied, and each ConvNeXt block yields intermediate feature extraction results at the current scale. A feature correction module constructed in step 3.1 is added between Down1 and Block1, and the feature fusion module constructed in step 3.2 is used to fuse features from each layer output of the two branches. Then, skip connections are used to introduce the fused features from different scales into the decoder, outputting the final target surface normal reconstruction result.
[0146] Final reconstructed target surface normal It can be represented as:
[0147]
[0148] Where f(·) represents the prediction model, i.e. the feature fusion network constructed in step 3.
[0149] Furthermore, in step 4, we use the following loss function as a guide for network optimization:
[0150]
[0151] in For the final output The cosine similarity loss is commonly used in normal estimation, and its expression is:
[0152]
[0153] Where <·,·> denotes the dot product. The surface normal estimated at pixel position (i, j), N ij This is the true value of the surface normal at that location.
[0154] Furthermore, the features generated by FCM are supervised, therefore cross-entropy loss is used separately. and For FC pol and FC sha Loss calculation is performed. Simultaneously, the feature FF generated by the last feature fusion network is analyzed. out Supervise and use cross-entropy loss
[0155] in, The expression is as follows:
[0156]
[0157]
[0158]
[0159] γ1, γ2, and γ3 are three hyperparameters that adjust different losses, and their default value in this example is 0.5.
[0160] Based on the above steps, we obtained 3D reconstruction results of the target surface at both the object and scene levels. To compare with other methods, we used the methods of Kondo (Kondo Y, Ono T, Sun L, et al. Accurate Polarimetric BRDF for Real Polarization Scene Rendering[C] / / European Conference on Computer Vision. Glasgow, UK: Springer, 2020:220-236.) and Lei (Lei C, Qi C, Xie J, et al. Shape from Polarization for Complex Scenes in the Wild[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. New Orleans, LA, USA: IEEE Computer Society, 2022:12632-12641.) to compare our method on both object and scene level data. The results are shown in the appendix. Figure 6 - Appendix Figure 7 As shown.
[0161] To quantitatively evaluate the 3D reconstruction results, we introduced Mean Angle Error (MEE), Median Angle Error (MEE), Root Mean Square Error (RMSE), and the proportion of pixels with angle errors less than 11.25°, 22.5°, and 30°, respectively, as evaluation metrics. Smaller values for the first three metrics indicate better reconstruction results, while larger values for the latter three metrics indicate better reconstruction results. The quantitative comparison results on the object-level dataset are as follows:
[0162] Table 1. Quantitative analysis of different reconstruction methods on object-level datasets.
[0163]
[0164] The quantitative comparison results on the scene-level dataset are as follows:
[0165] Table 2 Quantitative analysis of different reconstruction methods on scene-level datasets
[0166]
[0167] Quantitative results indicate that the reconstruction results obtained by the method proposed in this invention are superior to existing methods in both scene-level and object-level data. It can reconstruct target surface information with high quality, has stronger ability to reconstruct detailed information, and has better generalization ability.
[0168] On the other hand, embodiments of the present invention provide a polarization 3D reconstruction system based on a priori guided fusion network, comprising the following modules:
[0169] The polarization prior acquisition module is used to acquire a polarization prior input map set: it acquires polarization images of the target object and the target scene at different polarization angles, and calculates a polarization representation map set containing the unpolarized image, degree of polarization, re-represented phase angle, and field-of-view encoding based on the obtained polarization images at different polarization angles, which serves as the polarization prior input map set. Simultaneously, it obtains the polarization prior target surface reconstruction result.
[0170] The shadow prior acquisition module acquires the shadow prior input map set: based on polarized images at different polarization angles, unpolarized images, and the reconstruction results of the surface normals of the polarized prior surface, the shadow prior input map set containing the specular confidence and the reconstruction results of the surface normals of the shadow prior is calculated.
[0171] The network model construction module constructs a bi-branch feature fusion network for polarization 3D reconstruction. The bi-branch feature fusion network includes a bi-branch encoder and decoder, as well as a feature correction module and a feature fusion module. The polarization prior input map group and the shadow prior input map group are input into the bi-branch encoder to extract prior features. The extracted features are corrected by the feature correction module, and the extracted features are fused by the feature fusion module. Then, the fused features are input into the decoder to obtain the final target surface normal reconstruction result.
[0172] The reconstruction module is used to train a dual-branch feature fusion network by combining a loss function, and to use the trained network to achieve polarization 3D reconstruction of the target.
[0173] The specific implementation methods of each module are the same as those of each step, and will not be described in this invention.
[0174] It should be understood that any parts not described in detail in this specification belong to the prior art.
[0175] It should be understood that the above description of the embodiments is quite detailed, but it should not be considered as a limitation on the scope of protection of this invention. Those skilled in the art can make substitutions or modifications under the guidance of this invention without departing from the scope of protection of the claims of this invention, and all such substitutions or modifications fall within the scope of protection of this invention. The scope of protection of this invention should be determined by the appended claims.
[0176] References
[0177] [1]S.Zou,X.Zuo,S.Wang,Y.Qian,C.Guo,and L.Cheng,“Humanpose and shapeestimation from single polarization images,”IEEE Trans.Multimedia,2022.
[0178] [2]T.Yoshida,V.Golyanik,O.Wasenmuller,and D.Stricker,“Improving¨time-of-flight sensor for specular surfaces with shape from polarization,”inProc.IEEE Int.Conf.Image Process.(ICIP),2018,pp.1558-1562.
[0179] [3]X.Tian,R.Liu,Z.Wang,and J.Ma,“High quality 3D reconstructionbasedon fusion of polarization imaging and binocular stereo vision,”Inf.Fusion,vol.77,pp.19-28,2022.
[0180] [4]R.Liu,H.Liang,Z.Wang,J.Ma,and X.Tian,“Fusion-based highqualitypolarization 3D reconstruction,”Opt.Las.in Eng.,vol.162,p.107397,2023.
[0181] [5]M.Castelan,W.A.Smith,and E.R.Hancock,“A coupled statisticalmodelfor face shape recovery from brightness images,”IEEE Tran.Image Process.,vol.16,no.4,pp.1139-1151,2007.
[0182] [6]C.Li,S.Su,Y.Matsushita,K.Zhou,and S.Lin,“Bayesian depth-fromdefocus with shading constraints,”in Proc.IEEE / CVFConf.Comput.Vis.Pattern Recognit.(CVPR),2013,pp.217-224.
[0183] [7]L.Jiang,J.Zhang,B.Deng,H.Li,and L.Liu,“3D face reconstructionwithgeometry details from a single image,”IEEE Tran.Image Process.,vol.27,no.10,pp.4756-4770,2018.
[0184] [8]J.M.Di Martino,Q.Qiu,and G.Sapiro,“Rethinking shape fromshadingfor spoofing detection,”IEEE Tran.Image Process.,vol.30,pp.1086-1099,2020.
[0185] [9]Drbohlav and R.Sara,“Unambiguous determination of shapefromphotometric stereo with unknown light sources,”in Proc.IEEEInt.Conf.Comput.Vis.(ICCV),vol.1,2001,pp.581-586.
[0186]
[10] M.Olivier,M.Fabrice,S.Christophe,and G.Patrick,“Activelightingapplied to 3D reconstruction of specular metallic surfaces bypolarizationimaging,”Le2i UMR CNRS,vol.5158,p.12,2005.
[0187]
[11] T.Ngo Thanh,H.Nagahara,and R.-i.Taniguchi,“Shape andlightdirections from shading and polarization,”in Proc.IEEE / CVFConf.Comput.Vis.Pattern Recognit.(CVPR),2015,pp.2310-2318.
[0188]
[12] Y.Kondo,T.Ono,L.Sun,Y.Hirasawa,and J.Murayama,“Accuratepolarimetric BRDF for real polarization scene rendering,”inProc.Eur.Conf.Comput.Vis.(ECCV),2020,pp.220-236.
[0189]
[13] Ronneberger,P.Fischer,and T.Brox,“U-net:Convolutional networksforbiomedical image segmentation,”in Proc.MICCAI,2015,pp.234-241.
[0190]
[14] C.Lei,C.Qi,J.Xie,N.Fan,V.Koltun,and Q.Chen,“Shapefrompolarization for complex scenes in the wild,”in Proc.IEEE / CVFConf.Comput.Vis.Pattern Recognit.(CVPR),2022,pp.12 632-12 641.
[0191]
[15] S.Mingqi,X.Chongkun,Y.Zhendong,H.Junnan,and W.Xueqian,“Transparent shape from single polarization images,”arXiv:2204.06331,2022.
[0192]
[16] M.Muglikar,L.Bauersfeld,D.P.Moeys,and D.Scaramuzza,“Eventbasedshape from polarization,”arXiv:2301.06855,2023.
[0193]
[17] C.Lei,C.Qi,J.Xie,N.Fan,V.Koltun,and Q.Chen,“Shapefrompolarization for complex scenes in the wild,”in Proc.IEEE / CVFConf.Comput.Vis.Pattern Recognit.(CVPR),2022,pp.12 632-12 641.
Claims
1. A polarization 3D reconstruction method based on a priori guided fusion network, characterized in that, Includes the following steps: Step 1, Obtain the polarization prior input map set: Obtain polarization images of the target object and the target scene at different polarization angles. Calculate a polarization representation map set containing the unpolarized image, degree of polarization, re-represented phase angle, and field-of-view encoding based on the obtained polarization images at different polarization angles, and use it as the polarization prior input map set; simultaneously obtain the polarization prior target surface reconstruction result. Step 2, Obtain the shadow prior input map set: Calculate the shadow prior input map set containing the specular confidence and shadow prior surface normal reconstruction results based on polarized images at different polarization angles, unpolarized images, and polarized prior surface target normal reconstruction results; Step 3: Construct a bi-branch feature fusion network for polarization 3D reconstruction. The bi-branch feature fusion network includes a bi-branch encoder and decoder, as well as a feature correction module and a feature fusion module. The polarization prior input map group and the shadow prior input map group are input into the bi-branch encoder to extract prior features. The extracted features are corrected by the feature correction module, and the extracted features are fused by the feature fusion module. Then, the fused features are input into the decoder to obtain the final target surface normal reconstruction result. The feature correction module is used to correct the first layer features in the two encoders, including channel correction and spatial correction, and then the results of channel correction and spatial correction are weighted to obtain the result of comprehensive correction. Step 4: Train a dual-branch feature fusion network using a loss function, and use the trained network to achieve polarization 3D reconstruction of the target.
2. The polarization 3D reconstruction method based on a priori guided fusion network as described in claim 1, characterized in that: The dual-branch feature fusion network consists of two encoders for extracting two different prior features. The two branches have the same structure and parameters except for the input channels. Each branch has several layers, and each layer includes a downsampling module Down, which is used to downsample the input image to an appropriate feature map size, and a Block composed of multiple ConvNeXt blocks for feature extraction. A feature correction module is added between the downsampling modules Down and Block in the first layer, and a feature fusion module is used to fuse the features of each layer output of the two branches to obtain fused features at different scales. Then, a skip connection is used to introduce the fused features at different scales into the decoder to output the final target surface normal reconstruction result. Final reconstructed target surface normal Represented as: in This represents the prediction model, namely the dual-branch feature fusion network constructed in step 3. This is a set of polarization prior maps. This is the shadow prior group.
3. A polarization 3D reconstruction method based on a priori guided fusion network as described in claim 1 or 2, characterized in that: The specific processing procedure of the feature correction module is as follows; Given a polarization prior Feature maps and shadow priors The feature maps are first concatenated, and then max pooling and average pooling are performed along the channel dimension. Then, the two vectors are concatenated, and MLP is used to obtain the two channel weight mappings. : in This indicates channel multiplication, where B represents the batch size, C represents the number of channels, and H and W represent the pixel size of each channel image. and Feature correction results for both channels; The structure of spatial feature correction focuses on feature correction of important regions in an image. Therefore, it will... and Connect them together, and then obtain two spatial weight maps through MLP. and The spatial correction operation is as follows: in Represents space multiplication. and This indicates the spatial correction results for the two channels; The overall characteristics after polarization and shadow prior corrections are as follows: in , , , There are 4 hyperparameters. , It is the corrected feature map after comprehensive correction.
4. The polarization 3D reconstruction method based on a priori guided fusion network as described in claim 1, characterized in that: The feature fusion module is used to fuse the features of each layer in the encoder. The specific processing procedure is as follows: Two input branches and The output is then subjected to a 1×1 convolution operation and added to an LN, combined using the GELU function, and finally an MLP is applied to the result to obtain a fused mapping. Furthermore, the information interaction operation is as follows: Finally, the fused features are combined using a ConvNeXt block. and Combined into a single feature B represents the batch size, C represents the number of channels, and H and W represent the pixel size of each channel image.
5. The polarization 3D reconstruction method based on a priori guided fusion network as described in claim 1, characterized in that: The following loss function is used as a guide for network optimization: in, For the total loss, , , These are three hyperparameters that adjust for different losses. For the final output of the network The cosine similarity loss, used for normal estimation, is expressed as follows: in Represents the dot product. It is the surface normal estimated at pixel position (i, j). H represents the ground truth value of the surface normal at that location; H and W are the pixel sizes of each channel image. Furthermore, the features generated by the feature correction module are supervised, therefore cross-entropy loss is used separately. and Output of the feature correction module and Loss calculation is performed, and simultaneously, the features generated by the last feature fusion module are processed. Supervise and use cross-entropy loss .
6. The polarization 3D reconstruction method based on a priori guided fusion network as described in claim 5, characterized in that: , , The expression is as follows: 。 7. The polarization 3D reconstruction method based on a priori guided fusion network as described in claim 1, characterized in that: Step 1 includes the following sub-steps: Step 1.1, the intensity of each pixel in the polarized image at different polarization angles is represented as follows: in For unpolarized intensity, The phase angle, The degree of polarization (DoP) is the polarization angle. Four polarization images ( ) were calculated separately , and ; Step 1.2, Surface normal azimuth angle and zenith From the phase angle respectively and polarization degree The calculation yields the following result, which is expressed as: When specular reflection is dominant and The relationship is: in The reflectivity of the target surface; When diffuse reflection is dominant and The relationship is: Further, the target surface normal is obtained: in, The result is the reconstruction of the surface normal of the target based on polarization prior. , , These are the component values of the target surface normal vector in three directions; Step 1.3, Obtain the phase angle Representation and field of view coding field of view coding This can be obtained from the camera's inherent parameters, and further, a set of polarization prior maps can be obtained. : 。 8. The polarization 3D reconstruction method based on a priori guided fusion network as described in claim 1, characterized in that: Step 2 includes the following sub-steps: Step 2.1, Unpolarized image of the target and the direction of the light source The following relationship exists: in This is a generalized shallow undulation transformation, corresponding to the negative value of the surface normal. The above equation is solved using the least squares method, taking the value with the smallest residual. As a solution for the direction of the light source; Furthermore, the reflectivity of the target surface is related to the following relationship. Make an estimate: As a nonlinear optimization problem, the bounded constraint solution for reflectivity is obtained using the trust region reflection algorithm; Furthermore, the shadow constrains the target surface normal. Estimation from shadow cues using the Lambertian model: in Target surface albedo The direction vector of the light source The three directional component values; Step 2.2: Calculate the mirror confidence level based on the intensity changes of the four original polarization images. S : y For a sliding pixel block, These are the intensity images at the corresponding polarization angles; The final shadow prior is obtained as .
9. A polarization 3D reconstruction system based on a priori guided fusion network, characterized in that, Includes the following modules: The polarization prior acquisition module is used to acquire polarization prior input map sets: acquire polarization images of the target object and the target scene at different polarization angles, calculate a polarization representation map set containing the unpolarized image, degree of polarization, re-represented phase angle, and field-of-view encoding based on the obtained polarization images at different polarization angles, and use it as the polarization prior input map set; at the same time, obtain the polarization prior target surface reconstruction results; The shadow prior acquisition module acquires the shadow prior input map set: based on polarized images at different polarization angles, unpolarized images, and the reconstruction results of the surface normals of the polarized prior surface, the shadow prior input map set containing the specular confidence and the reconstruction results of the surface normals of the shadow prior is calculated. The network model construction module constructs a bi-branch feature fusion network for polarization 3D reconstruction. The bi-branch feature fusion network includes a bi-branch encoder and decoder, as well as a feature correction module and a feature fusion module. The polarization prior input map group and the shadow prior input map group are input into the bi-branch encoder to extract prior features. The extracted features are corrected by the feature correction module, and the extracted features are fused by the feature fusion module. Then, the fused features are input into the decoder to obtain the final target surface normal reconstruction result. The feature correction module is used to correct the first layer features in the two encoders, including channel correction and spatial correction, and then the results of channel correction and spatial correction are weighted to obtain the result of comprehensive correction. The reconstruction module is used to train a dual-branch feature fusion network by combining a loss function, and to use the trained network to achieve polarization 3D reconstruction of the target.